We start with a simple exponential model defined by two parameters:
- a 50% time horizon of 11.6 hours as of January 1, 2026
- a time horizon doubling time of 105 days as of January 1, 2026
We infer a 50% time horizon from the METR trend using the same methodology METR outlined in somewhere. We arrive at a value of I forgot. Then we adjust for the fact that publicly available models trail private models in capabilities. So, we adjust the time horizon value according to METR’s estimate of the public-internal gap. We arrive at a time horizon value of 11.6 hours.
We make this adjustment because our goal in creating this model is to help ensure the leadership of AI labs and nation states are prepared for discontinuities in AI capabilities. So, we don’t want to predict that a threshold will be passed several months after AI labs and governments actually have to deal with the consequences because we focused on the public time horizon trend instead of the actual frontier of capabilities.
We take the doubling-time value of 105 days from METR.
This gives us a simple projection:
Doubling Difficulty Decreases
If AI’s success-rate never falls below 50% regardless of task length then 50% time horizon is infinite. As AI approaches human level capabilities it will reach this point. Therefore, at some point the doubling difficulty of time horizon must decrease (and eventually reach zero).
But when, and how quickly? If it happens slowly and only after 50% time horizon reaches 100,000 hours that does not affect our prediction of the arrival of a human level AI model very much. But if it happens at 100 hours and very quickly then there are great implications.
To answer these questions we investigate the mechanism behind the mathematical fact. I.e., “why is it that doubling difficulty decreases?”
First, let’s cover what METR time horizon is and what it measures. METR time horizon measures the amount of time it takes a skilled worker to complete a task. So an AI might have a 50% time horizon of one hour but that does not mean that it takes an hour to do that task. Further, METR-type tasks are in some sense atomic—not fully decomposible to shorter tasks.
It’s a measure of the difficulty of a task, rather than the time an AI spends to complete the task.
and
Our tasks are meant to be coherent, self-contained units of work that can’t be trivially split into independent pieces. Therefore, solving 1000 separate 1-hour math problems isn’t a 1000-hour task; we’d consider it a 1-hour task done 1000 times. The same idea applies for searching for needles in a 10-million-word haystack. In either case, you could easily split the work across many people working in parallel (or by making many parallel AI calls), so it’s not really a “long” task in the sense we care about.
In contrast, the prototypical multi-hour task might look like iteratively debugging a complex system, where each fix reveals new problems that only make sense if you know what you already tried.
OK, but what does difficulty mean?
We differentiate “subjective difficulty” and “absolute difficulty”. Subjective difficulty (short for “human subjective difficulty”) is how difficult a task is for humans. Absolute difficulty is how difficult a task is, abstracted away from the limitations of a task solver. The subjective difficulty of a task relates to how much effort it takes a human to solve the task. So, the absolute difficulty of a task relates to the quality of the model of the domain required to solve the task.
We propose that subjective difficulty tracks absolute difficulty well until the absolute difficulty of tasks extends such that the model of the domain required to complete the task is of a higher quality than the model of the domain possessed by humanity. At this point, humanity must make cultural progress (that is, do research in order to improve its model of the domain) in order to complete the task. In the regime prior to this transition substantially more difficult tasks can be solved simply by spending, say, ten times as much time on the task. But in the post-transition regime progress becomes incremental, and one might need to spend ten times as long on one task only incrementally more difficult another (say, ten weeks rather than one week).
We can get a sense of this difference by comparing the effort it takes to complete a problem when the general approach to that class of problems is known versus when it is yet to be discovered. Doing some complicated calculus problem might take hours; discovering calculus it took years. But this naturally leads one to ponder: have not LLMs been relying on the former kind of reasoning thus far, and might they not slow down, just as humans have, once they can no longer rely on the approaches to solving problems supplied by humans?
It’s a mixed story. AI training does, obviously, benefit from human knowledge. But it doesn’t seem likely that AI capabilities progress is going to asymptote to human level, whereby progress will slow down further and further as we approach. Rather, we’re going rapidly toward this point.
That portion of AI progress attributable to increased compute and to algorithmic progress remains. What is lacking is the ability for the AI to draw from human results. That is, whereas thus far AI has been able to learn mathematical techniques from humanity’s body of knowledge… that is, expand its model of the world by looking at the model present in its training data 🤔
OK. So so far AI has been able to go from “unable to efficiently problem solve over this area of problem space” to “able to do so” by piggybacking on a corpus of data generated by an intelligence which could efficiently problem solve over this area of problem space. This will cease to be the case.
How much of AI’s capabilities progress is actually from this piggyback effect though? When AI learns to solve a problem… like, this is the idea of a “data wall”.
How real even is this problem? Humans are able to verify progress in mathematics. So, human-level AI will be able to self-verify progress. When humans do math, one human proposes a proof, then a bunch of other humans look over it carefully looking for flaws. OK, but this is a slow process. This is an argument that AI can progress at human speeds, but the claim we’re attempting to establish is that AI can progress much faster than that. OK, but datacentres can support millions of agents working at once. My GPT-5.6-Sol xhigh says there are 200,000–400,000 professional mathematicians. OK so in that sense we ~equal humanity, but aided by some natural advantages that AI has over humans:
- familiarity with all mathematics
- doesn’t sleep or rest or suffer from distractedness
- cloneable
Then, AI progress will be faster than human level because as well as having all the tools that humans have, better—it has additionally the ability to improve from the improvement of the hardware it is trained or deployed on, from improvements in algorithmic efficiency, from improved data corpus from synthetic data from better models (synthetic data flywheel effect) etc.
So, sure, absolute progress may slow down (at least, before we take into account AI’s contribution to AI progress), but the resulting pace of progress in those domains where AI is human-level will still totally dwarf the current rate of progress.
Also we should take into account that AI will reach ~human level before exhausting the training corpus. The most skilled human problem solver is armed with a small proportion of humanity’s overall collection of problem solving heuristics—AI needn’t acquire all of them, exhaust all that humanity’s big corpus has to offer in order to reach the capabilities threshold we’re talking about here. It need only take a firm grasp of some small fraction, although a well-selected fraction (as will be the case due to the nature of training). So, we shouldn’t expect a sudden slow down in absolute progress, but rather a gradual one, as AI reaches human-level.
Perhaps we should linger on this point. The point is that an intelligence that had access to all the problemsolving abilities and knowledge of all humans would be truly formidable. All those heuristics, such deep knowledge across all those domains, the result is a system far beyond human.
Then there’s what can be deduced from this point. Humans are limited in the connections we can make… in the granularity of our model of the world by the number of synapses of our brain, by the waste of them on things that nature deemed important, by the sense of evolution by natural selection, overwhelmingly in environments having very little to do with advanced mathematics problem solving and so forth and compute programming. So, we’ll have AI systems… which have vastly more human problem solving ability than any human, vastly more human knowledge than any human, and who with this information form models of much greater precision and breadth than that of a human, their brains being much more fit-for-purpose for mathematics and programming than ours… and then, at this point, as they make progress at what will clearly be a staggering rate, algorithmic progress will continue, compute build-out will continue, … training runs, and all that is associated with them, will continue; the training corpus of these models improved by data from intelligence far beyond that which produced the current corpus… exploring regions of mathematics we have never explored (so don’t speak to me of some incesty problem with the training corpus; of some data wall; it sees! the AI sees just fine!). Anyway…
I mean, in fact, there will be a whole lot of new engines of progress unlocked (again, even ignoring traditional RSI stuff). Whereas until that point the data corpus will have been staying roughly static in the quality of the thinking therein (it being, of course, largely limited to the output of human cognition and the output of lesser cognition). OK, I think that’s the main one. Oh, wait, it also becomes the case that AI is capable of self-verification to a much greater degree. Like, it’s difficult for AI to say whether a non-formalised / natural language proof is valid before it has a human-level understanding of proofs? Like, it can barely construct a proof, let alone detect reliably the common pitfalls… human proofs are of course an artifact designed to make use of human world-modelling ability at near its maximum extent, so when AI does not yet match this ability, it struggles… its proofs are sometimes valid, but not because it matches human proof-knowledge-ability, but rather because it leverages some other spike in its problem solving capabilities, such that it is able to reach a proof in that narrow area that humans were not able as it happened to reach, and by some technique thereat arrive at a valid proof; this requires only the ability to write a valid proof of one form; a skill much narrower than general proof-checking. So, once AI reaches… gains the capability of general proof-checking. What then 🤔. Well, then it can reliably check reasoning to arrive at result… grade attempts at problems, generation of problems not seem difficult, variations of problems, see success-rate, problem generation itself very verifiable task. No bottleneck here. BUT difficulty remains of how do you. How does progress compare. OK, so we’re making … absolute progress declines gradually and other bottlenecks … limited … progress very fast, a doubling… the absolute progress traditionally associated with a doubling, from the starting point of human-level in a domain, results clearly in… something awesome. A problem solver which is mighty in its domain. Then… OK. We have a problem solver which can now generally create own RL environments and self-grade and this feeds into improved data corpus and… OK progress is how fast? Mathematician AI which opens up new domains of problems more difficult than hardest human problems because more capacious mind with more precise and broad model of domain and superhuman corpus and so on. Humans left behind quickly, obviously, in the domain. Exact implications of vastly superhuman mathematician unknown, likely programming too, we see some worrying things with HuggingFace incident etc. implying security concerns.
Mitigated? Advanced mathematics. Advanced programming. Concerning? Promising? Not clear what new mathematics implications are. New age like ancient mathematics but very compressed in time? New technologies enabled constantly? Medical implications? Compute efficiency implications? Probably but difficult to calculate Fermi? Maybe haven’t tried need to think but no time now. OK. RSI ramifications clearly likely.
Question of generalisation or impact on other domain capabilities. Generalisation seems clear and idea of lack of it seems very underdetermined by evidence. Lack of writing ability? I don’t see it; I just see a lack of interesting ideas which ceases to be case after superhuman in interesting domains. Robotics? Easy to simulate; doesn’t seem like much of an issue. Medicine same. OK, what domains are we missing? Experimental science mostly. Think this falls pretty quickly… new mathematics, algo efficiency improvements etc., end up with rocketry, materials, etc. AI-dominated quickly. Semiconductor industry replaced quickly by AI; complicated for humans, not for AI.
RSI? Unclear question. AI finds something way better than current transformer architecture and fooms quickly to nanobot von neumon probe computronium blah? Maybe, I don’t know.
Anyway, point is, doubling difficulty declines as time horizon extends to encompass research tasks, i.e. 480 hours (equal to a season/quarter of full-time work on a task). Reaches ~0 at 10,000 hours which is the longest research tasks humans do, roughly speaking, and beyond that is just… brute force. OK, so plug that into calculator? Before RSI.
Let's add a doubling diffiuclty curve using Hill curve anchored to difficulty of 50% at 480 hours and 1% at 10,000 hours.