← Systems & Methods

Systems & Methods

Reading the Files:
AI is now the interface

The Epstein Files Transparency Act requires the documents be published in a “searchable and downloadable” format. What arrived was 341,000 files and 246 gigabytes, largely as scans — one release had text extractable from about 2,000 of its 23,124 documents. Journalists, engineers and volunteers filled the gap with AI tools, and those tools are now, for most people, the only practical way into the record. None of them is audited. The New York Times, describing its own system, warns that AI is prone to hallucination “particularly around sensitive issues like redactions” — which in this case means re-identifying people the government already exposed once.

Released
3.5m pages · 246 GB
One batch searchable
~2,000 of 23,124
Public AI tools
A dozen or more
Audited
None
Built with AI
This site included
Read this first — including about this site

This archive was built with AI assistance, and several of its sources are AI-generated document indexes. Everything on this site should be checked against the primary documents it links to. A page about the risks of AI-mediated access to these files would be dishonest if it did not begin by saying that.

The Finding
The law required the files to be searchable. The government released scans. Almost everything the public knows about these documents now passes through AI tools that nobody audits.
The Epstein Files Transparency Act requires the material be published in a “searchable and downloadable” format. In practice the DOJ’s search has been described as “crude at best,” with users reporting major malfunctions at scale. Of 23,124 documents in one House Oversight release, only about 2,000 had extractable text — and that subset is what news outlets analysed.

Three and a half million pages, 180,000 images, 2,000 videos, 246 gigabytes. No person can read that. The question of who gets to know what is in these files is therefore a question about software.

Into the gap stepped journalists, engineers and volunteers. The New York Times built its own tooling — scraping DOJ results into spreadsheets, semantic search, AI tagging, text extraction from images, audio and video, and a system to scan all three million pages. Its own account is careful: AI “cannot determine newsworthiness,” reporters treat outputs as tips, and humans make the final editorial calls.

Outside newsrooms, a dozen public tools now offer to answer questions about the files — EpsteinGPT, Jmail, askepstein, epsteinunboxed, Sifter Labs and others. Most are small volunteer projects. None is audited. All of them are, for most people, the only practical way in.

The most dangerous failure mode is specific, and the Times names it. AI is “prone to error and hallucination, particularly around sensitive issues like redactions.”

This site already documents what happens when redaction goes wrong: the January 2026 release exposed at least 31 people who had been victimised as children, and the harm was permanent within hours. A model that fills in a blacked-out name is not producing a formatting error. It is producing a re-identification.

And the failures are not hypothetical. One volunteer database returned “gibberish” transcripts from handwriting. A widely repeated “23,000 emails” figure turned out to be wrong — only about 3,000 were emails. A CJR study found AI search tools inventing citations to articles that do not exist, with premium tools performing worse because they are tuned to sound authoritative.

Which is the honest summary. These tools are doing genuinely valuable work that the government failed to do, and the good ones link every claim to a source document. But the public record of this case is now substantially mediated by unaudited software — and the only real safeguard is that every claim remains traceable to a page.

The Gap the Law Left

What the Act requires: the files be published in a searchable and downloadable format.

What was delivered: 341,000 files, 246 GB, largely as scans. DOJ search described as “crude at best.”

One release: 23,124 documents, ~2,000 with extractable text.

The estate flight batch: 5,233 records, 213 searchable.

The transparency obligation was met in form and failed in substance — and private parties filled the gap with no mandate and no oversight.

Section 01

What Goes Wrong

Six documented failure modes — and one thing the responsible tools all do right.

Hallucination around redactions

The New York Times, describing its own tooling, warns that AI is “prone to error and hallucination, particularly around sensitive issues like redactions.” Given that the January 2026 release already exposed at least 31 people victimised as children, a model that guesses at a blacked-out name is not a technical bug. It is a re-identification risk.

Gibberish from handwriting

404 Media on one volunteer database: “Some of the transcripts are gibberish, presumably caused by blurry or illegible type and handwriting on the source documents.” The flight logs are handwritten — which is why published flight counts differ by thousands.

Counting the wrong things

Sifter Labs had to publish a correction: “Media reports claiming ‘23,000 emails’ are incorrect — only ~3,000 are emails; the rest are various documents.” A widely repeated figure, wrong because the underlying set was never properly characterised.

Confidence without accuracy

A Columbia Journalism Review study found AI search tools generating citations to articles that do not exist and misattributing real ones. It also found premium tools performed worse — because they are optimised to sound authoritative.

No verification layer

Jmail’s Ilan Igel, on a small volunteer team: “It’s impossible for us to make sure that every single email is correctly verified.” His mitigation is the right one — a button on every email linking to the original document in the DOJ release.

What the good ones do

The responsible tools all converge on the same practice: every claim links back to the source document. Sifter Labs states it plainly — AI summaries are for initial research only; for serious work, verify against the original. That is the standard, and it is achievable.

Section 02

The Other AI Question

There is a second, backward-looking question people reasonably ask: did Epstein fund artificial intelligence, and did any of it serve his aims?

He funded some. Marvin Minsky, a founder of the field, and Joscha Bach, whose work is on computational models of cognition, both received support through the MIT orbit. The Media Lab took roughly $800,000 directly and $7.5 million arranged.

This archive audited all ten of his funded research programmes against his stated aims. AI is filed as “adjacent, not instrumental.” Mind-uploading is a transhumanist objective and cognitive architecture is its notional substrate — but the funded work was basic research decades from any such application.

Four of the ten programmes had no possible connection to anything he wanted. Cosmology cannot breed anybody. Quantum information theory has no route to heritable modification. The honest finding was that he was not buying a capability — he was buying a room.

Where AI does connect substantively is downstream, and it is not about him. The surveillance thread runs from cameras in a townhouse to Palantir, Carbyne, facial recognition sold into Nigeria, and Paragon’s Graphite now under a $2 million ICE contract. Those are machine-learning systems, and they are operating now.

And Minsky’s own position should be stated. He was named in a deposition, denied it, and died in 2016. No finding was ever made.

How this site was built
Research — web search against news reporting, court filings and primary DOJ documents, with AI assistance in locating and summarising.
Sources — include AI-generated document indexes. Where a figure comes from one of those, this site says so.
Verification — claims are checked against named primary sources, and every page links to them.
Uncertainty — where a fact is contested, disputed or unestablished, the page says so rather than smoothing it over.
Errors — this site has made and corrected several. Corrections are made on the page, not quietly.
Check anything here against the primary documents. That is what the links are for.

Section 03

Open Questions

?
Why were the files released unsearchable?
The Act requires a searchable format. No explanation has been given for releasing largely un-OCR’d scans, and no remediation has been announced.
?
Has any AI tool re-identified a survivor?
The Times warns of hallucination specifically around redactions. No audit of the public tools for re-identification risk has been conducted or published.
?
How much of the corpus is still unreadable?
One release had ~2,000 of 23,124 documents with text; one flight batch had 213 of 5,233. No overall figure for machine-readable coverage has been published.
?
Which claims in circulation came from AI error?
The “23,000 emails” figure was wrong and widely repeated. No systematic review of AI-derived errors in Epstein coverage exists.
?
Who is accountable for a public tool’s errors?
Most are volunteer projects with no institutional backing. There is no standard, no certification and no liability framework for AI tools built over government document releases.
?
What happens to the 3.3 million unreleased pages?
Roughly 3.3 million pages identified as responsive have never been published. If they arrive in the same format, the same gap reopens.

Section 04

Sources

Nieman Journalism Lab

AI-Powered Search Is Fueling a Wave of Transparency Projects

Mar 2026. The DOJ’s “crude at best” searchability, the volunteer tools, and Jmail on the limits of verification.

niemanlab.org →
Reuters Institute

How Newsrooms Are Digging Into the Files

The BBC, NYT and Guardian on their AI tooling — and the Times’ own warning about hallucination around redactions.

reutersinstitute.politics.ox.ac.uk →
404 Media

A Data Hoarder Builds a Searchable Database

Oct 2025. The volunteer OCR project, and the gibberish transcripts produced by handwriting.

404media.co →
Sifter Labs

The “23,000 Emails” Correction

The OCR coverage figures, and the published correction to a widely repeated media claim.

epstein-files.org →
Cross-reference

Elon Musk

Who was asking whom — and the fabricated email that now travels alongside the real ones.

Read the profile →
Cross-reference

The Redactions

What happened when redaction failed — and why hallucination around it is the acute risk.

Read the report →
Cross-reference

The Research Ledger

The ten programmes he funded, audited against his aims — including the AI work.

Read the report →
Cross-reference

Surveillance

Where machine learning actually connects — Palantir, Carbyne, Graphite and the ICE contract.

Open the hub →
Cross-reference

The Files

What has been released, what is withheld, and the 3.3 million pages still unpublished.

Read the report →