Internet Archive
Billions of archived pages, books, software, and media; the universal fallback when a source disappears.
Thirteen free archives, leaked-document databases, and official-statistics portals for people who have to get sourcing right — journalists, fact-checkers, researchers, and nonfiction writers. Each entry says what it's good for, what it costs, and where it lies.
Selection criteria: primary material or official data (not commentary), a usable free tier, and stability worth citing. Every link verified live on .
Because sources move, delete, and die. These give you a citable version of what was actually there.
Billions of archived pages, books, software, and media; the universal fallback when a source disappears.
Time-travel a specific URL: cite the snapshot for the version of the page you actually read.
Cross-border document sets and company records that no single newsroom holds — and collaborators who already indexed them.
The Panama/Pandora Papers class of leaks, plus the Offshore Leases database for free company-record searches.
Eastern Europe/Central Asia depth plus Project Search: cross-linked company, flight, and court records.
The record itself, before anyone interprets it — budgets, spending, people, and the datasets agencies publish.
The U.S. federal open-data catalog: search by agency, topic, or format and go straight to the source file.
Every federal contract, grant, and loan by recipient and location — the follow-the-money workhorse.
Population, income, housing, business counts — the denominators every "X% of Americans" claim needs.
Harmonized EU statistics — the only way to compare member-state numbers on the same basis.
UK population, economy, and society with unusually candid methodology and quality pages.
Cross-country development indicators (GDP, health, education) in clean downloadable series.
Research-grade long-form data essays with every chart downloadable and sources footnoted.
U.S. public-attitude and social research with methodology-first releases.
A cross-web catalog of published datasets — the starting map when you know the topic but not the publisher.
Finding the source is half the job; proving later which version you read is the other half. Pair this kit with the free source-log template (guide): one row per source, with a verification status that only moves when you opened the original. For the full workflow, see Marqly for journalists and the dead-link checker that keeps a filed story's citations from rotting.
Three criteria: the resource publishes primary material or official statistics (not aggregated commentary), it has a free tier usable without institutional access, and it is stable enough to cite. Every link on this page returned HTTP 200 to an automated check on 2026-09-26. Two well-known tools (DocumentCloud, Perma.cc) were left out of this edition because bot protection made them unverifiable by the same automated method — they are excellent resources; we simply only list what we could check.
The check date is printed on the page. When a link breaks or a service changes its access terms, the entry is corrected and the date bumped — a kit that never shows its age is a kit nobody maintains.
These are find-and-verify tools; the missing discipline is recording what you verified and where. The free source-log template (CSV or guide) is built for exactly that: claim, quote, locator, and a verification status that only moves when you opened the original.