Internal figures from a frontier lab show AI completing research tasks that take humans about 15 minutes, while benchmarks and forecasts put the same number at 4 to 11 hours.

「Recursive self-improvement」 means an AI smart enough to build a smarter version of itself, which then builds the next one, until after a few rounds you have something far beyond human. Science fiction calls it takeoff.
The futurist Ramez Naam has written a long post throwing cold water on the excitement. His conclusion: on the best data available, the self-improvement loop would need to be 5 to 10 times stronger before it could sustain itself.
Paper capability versus the lab
METR publishes a curve often called the most important graph in AI. It measures how long a coding task a model can finish on its own. By that chart, the best models handle work a human would need three hours to do.
OpenAI's own internal report says that on real research tasks, the work a model finishes unaided is roughly 15 minutes of human time. And that counts as a good result.
In other words, the public curve and what happens inside the lab differ by a factor of ten or more. The author argues for discounting outside forecasts.
The internal input numbers are stranger still. OpenAI researchers used 124 times more tokens per person than before, and engineers shipped about 7 times as many lines of code.
But experiments run per person rose only 1.6 times. And experiments are the thing that might actually produce a discovery.
Anthropic's estimate is blunter: to double the company's overall research progress, AI would have to raise employee productivity by roughly 40 times, not the 4 times observed now.
Good ideas get scarcer
Why doesn't piling on more help? The author uses an apple-picking image: AI strips the low-hanging fruit fast, but once one agent has picked it, sending a hundred identical agents just re-searches empty branches.
In fields where correctness can be verified automatically, like math, superhuman performance really has arrived; a model recently produced a proof of the Navier-Stokes existence problem.
Open-ended research is different. Anthropic itself admits its models prefer small incremental steps and avoid ambitious hypotheses; on a biology exercise the model deferred to published papers and could not develop new directions.
The author's closing line: takeoff may come, but somebody has to show evidence that the loop is genuinely strengthening. So far, it hasn't.
Why it matters
Whether AI can accelerate itself decides not just the technical roadmap but where the money goes: if the loop needs to be 5 to 10 times stronger to hold, then today's compute spending and valuations priced for an imminent takeoff are borrowing against a timetable the data does not yet support.



Curated from high-quality sources, with concise summaries and key takeaways.