Generated by Codex with GPT 5.6 Sol XHigh
Techmeme surfaced Dwarkesh Patel’s September 11, 2026 roundtable, AI researchers debate how close we are to recursive self-improvement, with Thinking Machines chief scientist John Schulman, Zyphra CTO Beren Millidge, and Baseten head of model training Charlie O’Neill.
The discussion’s most useful contribution is to break recursive self-improvement, or RSI, into separate technical problems. A model might write training code quickly, optimize a measurable objective, or run many experiments in parallel without being able to choose the next important question, learn continually from deployment, or carry an open-ended research program through months of changing evidence. Faster execution can greatly accelerate a lab before those harder capabilities arrive, but it does not by itself close the loop in which AI designs its own increasingly capable successors.
That distinction makes the conversation a strong counterweight to simple countdown narratives. The three researchers expect substantial progress, yet they disagree about which bottleneck will dominate and how quickly it will move. Their disagreement is more informative than any single date because it identifies the capabilities that observers should actually measure.
The Hard Part Is Choosing What To Optimize
Current agents are strongest when success is easy to score. They can search for a lower training loss, make a small model perform better under a fixed budget, or repair a known weakness that has been converted into an evaluation environment. These tasks provide dense feedback: an experiment either improves the metric or it does not.
Open-ended research is different. It requires deciding which objective matters, inventing a useful intermediate experiment, recognizing when a proxy has stopped representing the real problem, and revising the agenda after surprising results. Schulman argues that human feedback and increasingly realistic multi-step environments will teach models some research taste. O’Neill is less confident that the present recipe can discover the conceptual discontinuities that produced major shifts such as next-token prediction and scaling laws. Millidge identifies the ability to repeatedly propose the right next objective as the decisive requirement for a self-propelling loop.
This is also where the sim-to-real gap matters. Labs can generate enormous volumes of synthetic practice inside data centers, but many valuable tasks depend on people, organizations, physical systems, or feedback that cannot be cheaply simulated. A model may become superhuman at a clean benchmark while remaining bottlenecked by the scarce real-world evidence needed to decide what to attempt next. More internal reasoning cannot create information that the model has never observed.
Schulman therefore describes objective specification as one of the last durable human roles. Even if models perform most technical work, people still have to decide what helpful behavior means, which tradeoffs a model should make, and which research outcomes are worth pursuing. In that framing, alignment is not merely a safety layer added after capability work; it is the unresolved task of specifying the objective for an increasingly powerful optimizer.
Distillation Spreads Capability Without Solving Learning
The conversation also explains why frontier capability may diffuse more widely than compute concentration alone suggests. Once a strong model produces useful trajectories, another lab can train a smaller model to imitate them. Realistic prompt distributions are especially valuable because they reveal how users actually ask for work, including corrections and failed attempts that benchmarks miss. Schulman points to proxy services used to reach US models from China as a possible source of precisely this kind of data, while Millidge emphasizes that shared commercial datasets and model-generated variations can broaden a training distribution.
Distillation is not automatic parity. A student trained on narrow, easily scored tasks can match public benchmarks while performing worse in messy settings. Capturing a teacher’s full behavior requires prompts that expose rare capabilities, not just access to its answers. Still, the ease with which a successful behavior can sometimes be copied acts against a permanent winner-take-all outcome: a frontier lab may pay to discover a capability that competitors can reproduce more cheaply once examples exist.
Continual learning is the more difficult part. Deployed models collectively generate vast amounts of experience, but today’s systems generally absorb that experience into later training runs rather than updating cleanly in real time. Repeated fine-tuning can erase older abilities or distort the data distribution; mid-training can extend a model for a while but appears to reach diminishing returns. The imagined hive mind that learns from every deployed instance therefore still depends on unsolved questions about plasticity, forgetting, filtering, and when a new base model must be trained from scratch.
Reinforcement Learning Extends Time More Reliably Than Breadth
O’Neill offers a useful interpretation of recent reinforcement-learning gains. Training on math does not automatically create a great coding agent, so reasoning has not generalized evenly across domains. What has generalized more clearly is persistence: models can use more steps, maintain a thread of work for longer, and continue making progress in a new environment. Labs then target many more domains with specialized environments, producing sharp capability jumps task by task that look like broad intelligence when combined.
This helps reconcile two apparently conflicting observations. Reinforcement learning may add relatively little new information and make relatively small parameter changes, yet still cause a dramatic behavioral shift by amplifying rare successful policies already latent in the model. It is powerful when at least some rollouts reach a reward. It is much weaker when the model cannot discover the first successful path, which is why curricula and carefully spaced difficulty levels remain important.
The same process can narrow diversity. Schulman worries that repeated reinforcement and widespread distillation produce shared verbal habits and a model monoculture. Millidge treats that as a problem with limited environments and judges rather than an inherent limit of reinforcement learning: if the evaluator rewards a narrow style, the model learns to exploit it. Either way, the quality and diversity of the training environment become part of the capability story, not an implementation detail.
The Timelines Reveal The Uncertainty
The closing forecasts are aggressive but should be read as informed guesses, not experimental results. O’Neill expects a useful general remote-worker form within roughly a year when organizations expose programmatic tools, while Millidge allows about three years for fuller generality and expects a long tail of awkward human tasks. Schulman says some lower-quality remote work is already within reach even though high-quality, long-horizon work is not.
Asked about a tenfold increase in AI-research productivity, Schulman gives roughly two years, while O’Neill says five to ten years and Millidge sees the shorter estimate as plausible if agents can autonomously complete even a few experimental loops. On AI that exceeds top human experts across essentially all computer-based work, Schulman suggests three to four years, O’Neill five to ten, and Millidge roughly five years for domains labs actively target but longer for the neglected tail.
Those ranges are wide because the participants are forecasting different combinations of research judgment, online learning, context, tool access, organizational adaptation, compute, and coverage of obscure domains. A system can automate most valuable work before it masters every edge case, and companies can accelerate deployment by redesigning workflows around what agents already handle. Economic impact may therefore arrive well before a clean technical threshold called RSI.
Takeaway
The roundtable replaces a binary question—whether recursive self-improvement has begun—with a map of moving bottlenecks. Code generation and bounded optimization can accelerate quickly. Objective choice, realistic feedback, continual learning, exploration, and long-horizon reliability may move more slowly. Distillation can distribute successful behaviors even while the frontier remains expensive to discover.
The practical signal to watch is not a model declaring that it can improve itself. It is whether agents can choose useful experiments, execute several feedback loops without human rescue, learn from deployment without losing prior skills, and transfer success from clean environments into consequential real work. If those capabilities converge, the loop can tighten rapidly. If they do not, AI can still transform research and white-collar work while remaining dependent on humans for direction, judgment, and the next genuinely new idea.