Curated resource · link-checked 2026-09-26

The primary-source kit

Thirteen free archives, leaked-document databases, and official-statistics portals for people who have to get sourcing right — journalists, fact-checkers, researchers, and nonfiction writers. Each entry says what it's good for, what it costs, and where it lies.

Selection criteria: primary material or official data (not commentary), a usable free tier, and stability worth citing. Every link verified live on .

Archives & preservation

Because sources move, delete, and die. These give you a citable version of what was actually there.

Internet Archive

Billions of archived pages, books, software, and media; the universal fallback when a source disappears.

Access: Free; no account needed to read.

Watch out: Coverage is crawl-shaped, not complete — quiet corners of the web may never have been captured.

Wayback Machine

Time-travel a specific URL: cite the snapshot for the version of the page you actually read.

Access: Free.

Watch out: Snapshots can lag days or years; robots-blocked pages are missing; always note the snapshot date, not just the URL.

Investigations & leaked documents

Cross-border document sets and company records that no single newsroom holds — and collaborators who already indexed them.

ICIJ — International Consortium of Investigative Journalists

The Panama/Pandora Papers class of leaks, plus the Offshore Leases database for free company-record searches.

Access: Site and Offshore Leases database free; full leak corpora are member/newsroom-gated.

Watch out: The databases cover what has been leaked, not what exists — absence of a match is not innocence.

OCCRP — Organized Crime and Corruption Reporting Project

Eastern Europe/Central Asia depth plus Project Search: cross-linked company, flight, and court records.

Access: Free tools and reporting.

Watch out: Corrections exist; treat any single hit as a lead to verify against primary registries.

United States: official data

The record itself, before anyone interprets it — budgets, spending, people, and the datasets agencies publish.

Data.gov

The U.S. federal open-data catalog: search by agency, topic, or format and go straight to the source file.

Access: Free.

Watch out: Quality varies wildly by agency; freshness is not guaranteed — check the dataset’s own update notes.

USAspending

Every federal contract, grant, and loan by recipient and location — the follow-the-money workhorse.

Access: Free; bulk downloads available.

Watch out: Agency-reported data with reporting lag; subawards are under-reported — read their data-quality notes before quoting.

U.S. Census Bureau

Population, income, housing, business counts — the denominators every "X% of Americans" claim needs.

Access: Free.

Watch out: Different surveys measure different things on different schedules; mixing ACS with Current Population Survey numbers is the classic error.

Europe & UK: official statistics

Eurostat

Harmonized EU statistics — the only way to compare member-state numbers on the same basis.

Access: Free; downloadable tables.

Watch out: Methodology changes silently between releases; pin the version when you cite.

Office for National Statistics (UK)

UK population, economy, and society with unusually candid methodology and quality pages.

Access: Free.

Watch out: Post-Brexit trade series have known breaks — read the caveats before comparing eras.

Global data & research

World Bank Open Data

Cross-country development indicators (GDP, health, education) in clean downloadable series.

Access: Free.

Watch out: Modeled estimates for many low-income countries — the number is an estimate with error bars, even when the chart hides them.

Our World in Data

Research-grade long-form data essays with every chart downloadable and sources footnoted.

Access: Free.

Watch out: An interpretation layer on top of primary data — cite their essay and the underlying source both.

Pew Research Center

U.S. public-attitude and social research with methodology-first releases.

Access: Free.

Watch out: U.S.-centric unless stated; sampling frames shape what "Americans" means in each study.

Google Dataset Search

A cross-web catalog of published datasets — the starting map when you know the topic but not the publisher.

Access: Free.

Watch out: It indexes what publishers expose; dead links happen — verify any dataset before relying on it.

Working the kit into a story

Finding the source is half the job; proving later which version you read is the other half. Pair this kit with the free source-log template (guide): one row per source, with a verification status that only moves when you opened the original. For the full workflow, see Marqly for journalists and the dead-link checker that keeps a filed story's citations from rotting.

Kit questions

Why these and not others?

Three criteria: the resource publishes primary material or official statistics (not aggregated commentary), it has a free tier usable without institutional access, and it is stable enough to cite. Every link on this page returned HTTP 200 to an automated check on 2026-09-26. Two well-known tools (DocumentCloud, Perma.cc) were left out of this edition because bot protection made them unverifiable by the same automated method — they are excellent resources; we simply only list what we could check.

How often is this list re-checked?

The check date is printed on the page. When a link breaks or a service changes its access terms, the entry is corrected and the date bumped — a kit that never shows its age is a kit nobody maintains.

How does this fit a research workflow?

These are find-and-verify tools; the missing discipline is recording what you verified and where. The free source-log template (CSV or guide) is built for exactly that: claim, quote, locator, and a verification status that only moves when you opened the original.