A Map of AI Development Trajectories After Recursive Self-Improvement
SummaryAfter AI automates capabilities research, increasingly better models create increasingly better research, which creates even better models. This feedback loop may cause capabilities to grow explosively, or it could get stuck, or fizzle out. I map out the conditions for each trajectory to occur, compare how they differ to our current rate of improvement, and synthesize research by the safety community to predict which regime we are most likely in.
What this essay is and why I’m writing it
Hi, I’m Xiaoyao! At the time of this writing, I am a rising Sophomore at Stanford focused on alignment research.
This is the first in a series of four essays reviewing chapters of the Redwood AI Futurism Reading List. In each essay, I synthesize what I’ve learned to produce my own answer to the chapter’s relevant question.
I am doing this because what I can create is likely bounded by what I can imagine. By learning what possible trajectories AI development can take, what threats can emerge from them, and what solutions are already being tried can give me better intuitions for what project to embark upon.
I also hope to make these essays friendly to anyone encountering AI safety for the first time. If this is you, I’ve paid special attention to make this piece a good starting point for your explorations of the questions and contentions of this field, so read on! If anyone, new or experienced, has feedback on how I can do a better job of this, please do not hesitate to reach out to me via email! (marcus.lu.sg@gmail.com)
Overview and Introduction
Chapter 1 is about takeoff modeling, i.e.: What trajectory the development of AI capabilities will likely take.
My approach to this essay is as follows:
- I will flesh out the particularly dangerous trajectory of the software intelligence explosion, characterizing how it is different from our current trajectory.
- I will map out the possible alternative trajectories, and the variables which determine which trajectory we end up on.
- I will evaluate the arguments about what the values of those variables are.
I focus on analyzing an intelligence explosion based in recursive self-improvement in software, without considering changes in chip design and production. This is because, if possible, a software intelligence explosion is likely to be immensely impactful on its own due to the possible speed at which it can arrive.1
This is not a new subject of analysis. Forethought’s Will AI R&D Automation Cause a Software Intelligence Explosion and AIFP’s AI 2027 both center on the software intelligence explosion, and much of my analysis is based on Forethought’s work. For instance, they identified the crucial conditions at which trajectories branch I’m analyzing in the aforementioned article. Where this essay differs is that it lays the trajectories out side by side, tracking what happens to doubling times along each, rather than concentrating on the likelihood of the most worrying one.
To summarize my conclusions:
- I believe it is quite likely that automated software research will arrive in the coming 1-2 years and will shorten doubling times on the order of months.
- I believe it is moderately likely that this results in a strong self-improving feedback loop which shortens doubling times further, taking previous estimates of at 1.0-1.5.
- I believe it is plausible that the superexponential gets stuck at time lags produced by training runs, which I doubt can be pushed below one month; but it’s equally plausible that there are high ceilings to post-training enhancements which arrive on the order of days, through which the superexponential trend can continue.
- Consequently, I think continued superexponential growth is a serious possibility and can run until it approaches limits on the order of ~12 OOMs of effective compute away, as estimated by Forethought.
See “Limitations & Conclusion” for my critique of my model and its predictions. It goes without saying that no one can predict the future, but in my view, trying can give me the opportunity to learn.
The Problem
The problem articulated by Forethought, AIFP, and other parts of the AI Safety community is the prospect of an Intelligence Explosion caused by the recursive self-improvement of AI models. The premise is that eventually, models will be capable enough to conduct research & development in domains that improve their own capabilities, which will result in more capable models which do even better research. Ultimately, capabilities increase at an increasing rate, which threatens loss of control over the models.
To better understand how this trajectory is different to the present, we need a good metric for measuring capability. METR measures this with time horizons, where we evaluate models by the time duration of tasks they can succeed at with a set level of reliability (e.g. 50%).
METR’s graph of task time horizon that LLMs can complete with a 50% reliability.
Measured in this way, model capabilities have a doubling time of 4-7 months.2 In other words, the status quo trajectory of AI development, with its existing inputs, is exponential growth, whereby capability increases at an increasing rate.
Throughout this essay, “doubling” refers to a doubling of capability in this sense. I will also refer to other units later on—algorithmic efficiency, and orders of magnitude of effective compute—but these are proxies for capability, and I convert them back to this metric for coherence.
This creates a clear standard to compare trajectories of further speed-ups against. A trajectory can only become more extreme than the status quo in two ways: either it causes a one-time shrink to the doubling time, resulting in a steeper exponential, or it causes the doubling time to shrink continuously, resulting in a superexponential trajectory of development.
Intuitions for how much time we have in each trajectory
To build clearer intuitions for the on-the-ground impact of these trajectories, consider a toy example. Suppose the threshold we care about is a model that can carry out a month-long research project unsupervised: a 50% time horizon of roughly 167 working hours. From today’s frontier horizon of a few hours, that is about six doublings, or ~50x. This threshold is picked for illustration and nothing below depends on the specific number—models well short of it are already capable of consequential misaligned autonomous action.
Suppose the current exponential trajectory has a doubling time of 6 months. Then we reach the threshold in:
With the one-time speedup from automated research shortening doubling time to 2 months:
And with a superexponential trajectory, where each doubling arrives quicker than the last. At —the value I defend below—each doubling is 1.26x quicker than the one before it, starting from the post-automation 2 months:
Identifying the conditions at which these three figures occur is the subject of this essay. A one-time shortening of doubling time proportionately shortens the time available to make models safe, and a superexponential dynamic keeps taking bites out of that time, until the gap between doublings is shorter than the time humans need to respond to the last one. When, then, can these different trajectories unfold?
A Map of Possible Trajectories
In this section, I outline possible trajectories of doubling times due to improvements in software, and identify the conditions which result in different branches of this trajectory.
My graph of doubling-times vs. time for different development trajectories. Y-axis value are qualitative ballparks. The main spine is the intelligence explosion superexponential, where doubling times continue to shrink. When not all conditions are fulfilled, we land on other branches. is the return on R&D, where smaller values mean larger diminishing returns. The shape of these trajectories assume no significant changes in the current trajectories of other relevant factors to capabilities development, including compute and regulation.
Overview of Trajectories
There are broadly five trajectories based on different possibilities in software development:
- ASR does not arrive: doubling-time gradually lengthens.
- ASR arrives but feedback loop is weak: there is a one-time speedup to research due to automation. However, the return to software R&D , so the intelligence feedback loop is too weak to result in super-exponential growth and doubling times gradually lengthen.
- Feedback loop is strong but gets stuck due to time lags: , so the trajectory is super-exponential and doubling times shrink until getting stuck at the irreducible time lag between converting research insights into the next, more intelligent model.
- Time lag reducible, but the ceiling of software capabilities is low: the physical limits of how far software can progress are close, so doubling times shrink dramatically until we near this low ceiling, at which point progress quickly drops off because we reach the physical limits of software-based advancements.
- Intelligence explosion: If there is a high ceiling and all the previous conditions are satisfied, we reach an intelligence explosion which bursts through many orders of magnitude of improvement within a short amount of time.
How each condition qualitatively affects doubling time
In my graph and the overview above, I make claims about the trend or ultimate value of doubling-times given whether each condition is satisfied. Below I explain for each condition how I arrive at these claims.
-
An automated software researcher results in a one-time shortening of doubling times. An ASR will increase the cognitive inputs to software research because fumbling humans who have to sleep are replaced by much quicker non-stop agents. Greenblatt also notes that these agents will be better at leveraging compute than humans, further increasing effective compute. Thus, doubling rates should see a one-time uplift, which has been estimated to have a magnitude of 3x..
-
Returns to R&D results in a superexponential trajectory: is the number of times capability doubles for each doubling of the cumulative research effort spent on software. So each capability doubling makes the next one times more expensive in research effort, while the doubled capability itself supplies twice the research effort. The next doubling therefore arrives times quicker than the last. If , then doubling times continuously shorten, resulting in a superexponential. For instance, at , each doubling arrives 1.26x quicker than the one before it.
- If , then the trajectory stays exponential and has a faster doubling-time than our current rate, because it retains the one-time uplift from automation.
- If , then the trajectory is sub-exponential; every doubling takes longer because difficulty grew more than capability did. Doubling-time gradually lengthens.
- If this is not intuitive, I highly recommend reading Forethought’s more rigorous and concrete formulation of these three cases through a toy example here.
-
The superexponential gets stuck if there is an irreducible time lag between research and deployment. For example, a researcher may discover an especially capable model architecture, but actually training the new model with this architecture to deploy the improved capabilities can take months (e.g. GPT-4 took 3-4 months to train). If this time lag turns out to be irreducible, then even if the time taken to do research shrinks to zero, it will take at least as long as this lag.
-
How far the superexponential grows depends on the limits of software capabilities. There is probably a theoretical limit to how much model capabilities can be improved by optimizing software alone, at which R&D should reap zero returns. Since the main danger of a superexponential growth trajectory is how much it grows in a small amount of time, if this growth quickly hits a ceiling, then growth will quickly drop back down.
Which one is our likely trajectory?
This depends on the answers to the four questions below. For each, I will attempt to bring arguments from both sides and come to a best guess. The people I am quoting are far more knowledgeable and have thought far deeper than me on this, so this is more an exercise for myself to form takes on difficult questions than an attempt to make a new contribution. I also imagine this would be a helpful synthesis of views for readers who are new to the field.
1. Automated software researcher seems likely
Dwarkesh Patel makes two main arguments why this is unlikely in his conversation with Ryan Greenblatt.
- No Verifiability: Models became good at the domains they’re good at (math, coding) through RLVR, which requires problems which are verifiable and grindable (many problems to do rollouts on). It is hard to verify how good a research insight is.
- Models have not yet displayed ingenuity: Even in domains that models are particularly good at, e.g. math, we have not seen it produce new theory. It seems that research taste in ML would require this type of theoretical reasoning.
Greenblatt argues that both of these can be overcome:
- Machine learning is pretty verifiable: There are many grindable tasks for training ML taste. As a simple example, whether a model can reduce the loss of a training pipeline is very verifiable as an RL task. Moreover, if models can predict the results of ML experiments, then this is somewhat transferrable to research taste. This is yet another verifiable and grindable task.
- Machine learning is more empirical than theoretical: ML progress does not usually seem to emerge from deep theoretical insights, but rather trying many things and producing clever ideas in response to empirical attempts. It is thus okay if models do not have deep ingenuity.
Given this, it seems that the theoretical component of ML is not a sufficient roadblock for models, and they have a good source of RL tasks to grind research taste from. Hence, it is pretty plausible that models will get significantly better at ML research over time.
Perhaps it cannot replace all human researchers at first, but, if it reaches adequacy, supremacy is probably not far (as defined by Cotra) since its capabilities follows the doubling-time of AI capabilities while human capabilities stay mostly constant.
2. Returns to R&D seem to be greater than 1
To estimate , we need measurements for:
- How much AI capabilities have grown due to increased cognitive (as opposed to computational) inputs.
- How much cognitive inputs needed to scale to result in this increase.
Epoch has conducted this estimation, although with some caveats. They measured how many times algorithmic efficiency in various domains doubled over respective time periods (e.g. Computer Vision from 2012-22). This proxies the growth of capabilities caused purely by research and not compute, since efficiency measures how much the same piece of hardware can do. They then proxied the increase in cognitive inputs by the growth in the number of researchers in that field.

Median and 5-95% confidence bounds for estimates of in “Estimating Idea Production”
On average, they found for all domains that their median estimates of , with vision and RL–two fields particularly relevant to modern language models–both at near . Separately, Forethought estimated for algorithms writ large (not restricted to ML) to provide an outer value to check existing estimates against. From 1970-2014, they found that .
My main critique of the reliability of these estimates is that the growth in specifically the machine learning fields that were evaluated could be attributed to the scaling of compute.
While the researchers were careful to evaluate algorithmic efficiency, which is more independent from compute than capability is, the scaling of compute does allow for increasingly large numbers of experiments and more empirical exploration. Hence, the growth of AI capabilities across these time periods may not be attributable to the increase in cognitive inputs.
However, this critique applies less to fields which are not as compute-hungry as deep learning, e.g. SAT solvers, linear programming, and algorithms in general, most of which have large time windows which precede the deep learning era. For these domains, is at least one, if not significantly higher than one, which suggests that returns to R&D are likely to be high.
Given that deep learning is also a rather new field with a parent field with , it is quite plausible that , and a strong feedback loop is possible.
3. Time lag of training new models may be difficult to reduce
Forethought discusses the time lags in the software feedback loop in “Three Types of Intelligence Explosion”. They note that it takes about three months to train a new State-of-the-Art model, but point out that post-training advancements like scaffolding and finetuning on curated data can be implemented within days, and have achieved improvements on the order of 5-30x effective compute. To tie this unit back to capability: GPT-4 was trained with roughly 3 more orders of magnitude of effective compute than GPT-3, and is about 20x as capable on METR’s time horizons. Thus, very roughly, three OOMs of effective compute buys around four doublings of capability.
To me, this seems the most likely place at which shrinking doubling times get stuck. While automated researchers can probably optimize the speed of training as it is the job of many engineers to do today, there seem to be close physical limits for training, e.g. in how much more kernels can be optimized and how many steps new models have to be trained for. My high-uncertainty guess is that the lag cannot be reduced to less than 1 month from the 3-4 months it is now.
However, it seems pretty plausible that automated researchers would be able to continue finding post-training enhancements which don’t suffer from the time lag of full training. In which case, the super-exponential is not bottlenecked by the one-month time lag of full training.
4. The ceiling of software capabilities development is probably high
Forethought estimates, with high uncertainty, that there are perhaps still 12 orders of magnitude of effective compute’s worth of efficiency improvements in software. This means that they believe these improvements are equivalent to scaling compute by a factor of one trillion. They provide the following rationale.
“If top-human-level AI is initially trained with 1e29 FLOP, that would be ~5 OOMs less efficient than human learning (which takes ~1e24 FLOP). Then we estimate a further ~7 OOMs of software progress might be possible above the human brain, with very wide uncertainty. (This only includes training efficiency, omitting other sources of software progress.)”
At the conversion rate above, this implies about four more leaps of the magnitude from GPT-3 to GPT-4 remaining from software research alone, which feels like an incomprehensible level of improvement given the current intelligence of the models.
Having this level of potential remaining seems plausible. Relative to other fields, deep learning is incredibly new (circa 2010s), and transformers have existed for less than a decade. After all, algorithmic efficiency in fields of computing older by decades still see doubling times of single-digit years.3 Hence, I lean towards higher ceilings for software capabilities development.
Limitations & Conclusion
No one can predict the future; every expert doing takeoff modeling I’ve read acknowledges the vulnerability of their work to unknown unknowns, and, having built this model over five days of research and writing, I am far from an expert. This model in particular is also flawed in that its base doubling rate of 7 months is an empirical observation likely dependent on the current exponential scaling of compute, which is an uneven and moving floor. I am unsure how fluctuations in this base rate due to factors like compute compose with impacts of software-based changes I analyze. Additionally, since METR’s doubling rate estimate is an empirical one, it is unclear whether the factors I analyze here are in fact incorporated by instead of additive to the predictions of that empirical trend (although I am reasonably confident that it is the latter, see footnote for reasoning).4
With these uncertainties, how then should I treat the predictions I’ve currently produced? My current answer takes inspiration from the Bayesian tradition: I will stand by the best answer I can muster with the limited knowledge I have, and keep learning and updating my beliefs as I go. I already have evidence that this approach is beneficial: as I forced myself to commit to answers over the course of writing this essay, I have become more viscerally aware of the assumptions I am making and what I do not know. In this spirit, I forward the (likely wrong) conclusions I have derived and will continue to derive, in hopes that over time, with your feedback, I will learn and and approach an accurate understanding of the world over time.
Footnotes
-
There are many reasons for this, chief among which is that improving software is completely virtual and frontier labs own most of the data on this already, so it is easiest for them to train models to get very good at this. As for other intelligence explosions, three inputs have been identified as the main factors for the development of AI capabilities: algorithms, data, and compute. More recently, in “Three Types of Intelligence Explosion”, Forethought proposed a new method of categorizing inputs to AI progress: software, chip technology, and chip production. In this taxonomy, software absorbs both algorithms and data, while compute is partitioned into chip technology and chip production. Due to the more diffused nature of information and longer time lags in chip technology and production, they are increasingly difficult to automate, result in more tenuous feedback loops when automated, and are amenable to greater diffusions of power. Read Forethought’s article to understand why this is. ↩
-
METR’s estimate (7 months) for doubling time is based on progress in the past 7 years. AIFP observes that METR’s graph seems to indicate a speedup since 2024, after which doubling time has been closer to 4 months. ↩
-
Epoch’s “Estimating Idea Production”, section 5.3.1. ↩
-
To this last point, I am relatively confident that it is additive rather than incorporative because 1) the recent growth trends are probably more attributable to compute than increases in talent for reasons in the second footnote of this article by Greenblatt, so the trend has no information on which to anticipate changes to cognitive inputs; 2) the relatively large deviation of Opus 4.5, 4.6, and Fable from the METR trendline potentially suggests an unanticipated recent increase in growth rates which may be explainable by partial automation of research due to Claude Code (internal use starting early 2025). If this is true, then this would be an example of how the METR trendline does not predict increases to cognitive inputs. ↩