Bargo
AI

Recursive Self-Improvement Is Already Here — Just Not How You Think

Dwarkesh, All-In and a16z debate whether AI building AI is autocatalytic help or true takeoff — the mechanism, the math, and the timeline

Bargo · 2026-08-24

The most important question in AI right now is whether AI doing AI research creates a feedback loop. Across Dwarkesh Podcast, All-In and a16z, the consensus is narrow: AI helping build AI is already happening, true recursive self-improvement where AI fully automates its successor is plausible but not proven, and the next two years of lab results will decide it.

What RSI means, and why definitions matter

The classic definition goes back to I.J. Good (1965) and Yudkowsky (2008): an ultraintelligent machine designs a better version of itself, creating a feedback loop. Lilian Weng traces this history and argues harness design is the lever.

Today's podcasts split the term in two:

A useful ladder from Don't Worry About the Vase, July 17 frames it as L0 delegation, L1 net positive, L2 ignition where it compounds, L3 inflection where it takes off. The first experimental hint cited is Jung Yao Jang's 8-day auto-research agent beating a 2-year hand-tuned harness.

This distinction matters for investors. Autocatalytic progress lifts productivity steadily, true RSI would compress years of progress into months and change compute demand abruptly.

The concrete mechanism bulls point to

Ryan Greenblatt, chief scientist at Redwood Research, laid out the most concrete plan on Dwarkesh, Aug 11:

So now we want to train GPT-7.5, and we come up with a bunch of different environments. There is already this repo that is the descendant of Andrej Karpathy's nanoGPT speedrun, where you just try to change everything about the model, from the optimizer to the hyperparameters to the architecture, to get it to a fixed training loss as fast as possible.

You could have other kinds of environments where you could say, Hey, GPT-7.5, I want you to train a really good video game playing model. I want you to train a model that actually improves as it plays the same video game again and again.

Then you basically put GPT-7.5 through a bunch of this kind of training, you build GPT-8. GPT-8 is now an amazing ML researcher.

In plain terms: 100s of small, verifiable, containerized tasks with clear scores, train GPT-7.5 with RL on all of them, and let intuition transfer to building the next frontier model. Greenblatt's intuition pump is mathematics, where once a domain became fully verifiable, progress came like a flood. ML is even better, he argues, because you see intermediate progress and innovations stack additively.

Why bulls think it compounds fast

Greenblatt's quantified claim: once AI matches top human researchers, expect about 4 to 5 years of normal progress compressed into one year. For scale, GPT-4 to Claude 5 was about 3 years. That much in one year, he says, is wildly superhuman, drop it in TSMC or Texas politics in the 1940s and it beats top humans.

Why skeptics push back

Casado's middle ground is that using AI to build software is not new, we have always done autocatalytic loops, the question is whether it becomes self-improving research.

Timelines where they land

What to watch

Sources: Dwarkesh — Ryan Greenblatt, Aug 11 / YouTube; All-In, Aug 21; a16z — Casado, Aug 22; Don't Worry About the Vase, Aug 15 and July 17; Lilian Weng — Harness Engineering, July 4.

More research at bargo.ai/research.