This archive was built with AI assistance, and several of its sources are AI-generated document indexes. Everything on this site should be checked against the primary documents it links to. A page about the risks of AI-mediated access to these files would be dishonest if it did not begin by saying that.
Three and a half million pages, 180,000 images, 2,000 videos, 246 gigabytes. No person can read that. The question of who gets to know what is in these files is therefore a question about software.
Into the gap stepped journalists, engineers and volunteers. The New York Times built its own tooling — scraping DOJ results into spreadsheets, semantic search, AI tagging, text extraction from images, audio and video, and a system to scan all three million pages. Its own account is careful: AI “cannot determine newsworthiness,” reporters treat outputs as tips, and humans make the final editorial calls.
Outside newsrooms, a dozen public tools now offer to answer questions about the files — EpsteinGPT, Jmail, askepstein, epsteinunboxed, Sifter Labs and others. Most are small volunteer projects. None is audited. All of them are, for most people, the only practical way in.
The most dangerous failure mode is specific, and the Times names it. AI is “prone to error and hallucination, particularly around sensitive issues like redactions.”
This site already documents what happens when redaction goes wrong: the January 2026 release exposed at least 31 people who had been victimised as children, and the harm was permanent within hours. A model that fills in a blacked-out name is not producing a formatting error. It is producing a re-identification.
And the failures are not hypothetical. One volunteer database returned “gibberish” transcripts from handwriting. A widely repeated “23,000 emails” figure turned out to be wrong — only about 3,000 were emails. A CJR study found AI search tools inventing citations to articles that do not exist, with premium tools performing worse because they are tuned to sound authoritative.
Which is the honest summary. These tools are doing genuinely valuable work that the government failed to do, and the good ones link every claim to a source document. But the public record of this case is now substantially mediated by unaudited software — and the only real safeguard is that every claim remains traceable to a page.
What the Act requires: the files be published in a searchable and downloadable format.
What was delivered: 341,000 files, 246 GB, largely as scans. DOJ search described as “crude at best.”
One release: 23,124 documents, ~2,000 with extractable text.
The estate flight batch: 5,233 records, 213 searchable.
The transparency obligation was met in form and failed in substance — and private parties filled the gap with no mandate and no oversight.
Section 01
What Goes Wrong
Six documented failure modes — and one thing the responsible tools all do right.
The New York Times, describing its own tooling, warns that AI is “prone to error and hallucination, particularly around sensitive issues like redactions.” Given that the January 2026 release already exposed at least 31 people victimised as children, a model that guesses at a blacked-out name is not a technical bug. It is a re-identification risk.
404 Media on one volunteer database: “Some of the transcripts are gibberish, presumably caused by blurry or illegible type and handwriting on the source documents.” The flight logs are handwritten — which is why published flight counts differ by thousands.
Sifter Labs had to publish a correction: “Media reports claiming ‘23,000 emails’ are incorrect — only ~3,000 are emails; the rest are various documents.” A widely repeated figure, wrong because the underlying set was never properly characterised.
A Columbia Journalism Review study found AI search tools generating citations to articles that do not exist and misattributing real ones. It also found premium tools performed worse — because they are optimised to sound authoritative.
Jmail’s Ilan Igel, on a small volunteer team: “It’s impossible for us to make sure that every single email is correctly verified.” His mitigation is the right one — a button on every email linking to the original document in the DOJ release.
The responsible tools all converge on the same practice: every claim links back to the source document. Sifter Labs states it plainly — AI summaries are for initial research only; for serious work, verify against the original. That is the standard, and it is achievable.
Section 02
The Other AI Question
There is a second, backward-looking question people reasonably ask: did Epstein fund artificial intelligence, and did any of it serve his aims?
He funded some. Marvin Minsky, a founder of the field, and Joscha Bach, whose work is on computational models of cognition, both received support through the MIT orbit. The Media Lab took roughly $800,000 directly and $7.5 million arranged.
This archive audited all ten of his funded research programmes against his stated aims. AI is filed as “adjacent, not instrumental.” Mind-uploading is a transhumanist objective and cognitive architecture is its notional substrate — but the funded work was basic research decades from any such application.
Four of the ten programmes had no possible connection to anything he wanted. Cosmology cannot breed anybody. Quantum information theory has no route to heritable modification. The honest finding was that he was not buying a capability — he was buying a room.
Where AI does connect substantively is downstream, and it is not about him. The surveillance thread runs from cameras in a townhouse to Palantir, Carbyne, facial recognition sold into Nigeria, and Paragon’s Graphite now under a $2 million ICE contract. Those are machine-learning systems, and they are operating now.
And Minsky’s own position should be stated. He was named in a deposition, denied it, and died in 2016. No finding was ever made.
Section 03
Open Questions
Section 04
Sources
AI-Powered Search Is Fueling a Wave of Transparency Projects
Mar 2026. The DOJ’s “crude at best” searchability, the volunteer tools, and Jmail on the limits of verification.
niemanlab.org →How Newsrooms Are Digging Into the Files
The BBC, NYT and Guardian on their AI tooling — and the Times’ own warning about hallucination around redactions.
reutersinstitute.politics.ox.ac.uk →A Data Hoarder Builds a Searchable Database
Oct 2025. The volunteer OCR project, and the gibberish transcripts produced by handwriting.
404media.co →The “23,000 Emails” Correction
The OCR coverage figures, and the published correction to a widely repeated media claim.
epstein-files.org →Elon Musk
Who was asking whom — and the fabricated email that now travels alongside the real ones.
Read the profile →The Redactions
What happened when redaction failed — and why hallucination around it is the acute risk.
Read the report →The Research Ledger
The ten programmes he funded, audited against his aims — including the AI work.
Read the report →Surveillance
Where machine learning actually connects — Palantir, Carbyne, Graphite and the ICE contract.
Open the hub →The Files
What has been released, what is withheld, and the 3.3 million pages still unpublished.
Read the report →