Frontier models are already producing real results in codebreaking and textual provenance, but the author finds the bottleneck is not model intelligence, it is whether archives will open their doors.

The author is a historian of science, and his usual headache is deciphering bad handwriting in old manuscripts. This time he pointed GPT-6 and Opus 5.5 at something more ambitious: using the models to actually solve long-standing historical problems.
What kind of history problem can an AI bite into
He lists the conditions. Experts have to have already agreed it needs solving. The source material has to be digitised and freely available. The work should play to a model's strengths: multilingual reasoning, maths, and grinding through large document sets on its own. And most importantly, the answer has to be clearly provable or disprovable.
Maths fell to these models precisely because theorems can be verified. Most historical debate is interpretation and standpoint, with no single right answer, so progress is far slower. What is left tends to cluster in three places: codebreaking, tracing texts across languages, and connecting findings that sit scattered in narrow subfields.
It legitimately produced new findings
One case: a German military message from 1941. The breakthrough was not cryptography itself but the model stumbling onto a note in the German Federal Archives about a newly catalogued message collection, then assembling the pieces. The author says even leading human experts do not fully understand how it got there.
Another case is more concrete. Opus 5.5 downloaded over five thousand primary files from the Hartlib archive and spotted that Newton and Hartlib each used a different anagram for the same alchemical ingredient, Hungarian vitriol. Newton's is a true anagram, letters rearranged; Hartlib's is closer to backwards writing. This may be a new finding, publishable even.
In a third case the model partly deciphered two 16th-century Spanish letters. On checking, one had been decrypted in the 1530s and the other in 1916. The author says this shows how easily someone without expertise redoes work a scholar finished a century ago.
The real blockage isn't compute
The author keeps returning to one point: these models work because huge volunteer efforts transcribed, scanned and posted these old documents for free. The wall he hits most often in testing is that material is either undigitised, or digitised but available only to people with privileged access.
So he calls on AI labs and foundations to fund three things: get libraries and archives to digitise undigitised manuscripts and put them online for free; give historians free compute and API access; and have historians band together and publish a list of open problems, the way mathematicians have.
He is informally surveying colleagues on which unsolved historical mysteries would be worth throwing AI at.
He is not naive about it. Using AI to write or to replace original thought does encourage cognitive offloading. Using it to ask questions no single person could ever read their way to is something else.
Why it matters
The same model performs wildly differently across fields, and the difference is not intelligence but whether the field has a referee who can call an answer right or wrong. The point worth taking away is that whether archives open up may matter more to where history research goes than whether the next model is smarter.



Curated from high-quality sources, with concise summaries and key takeaways.