AI That Improves Itself Is Real. Does It Compound?

The phrase “recursive self-improvement” was on four arXiv papers in 2024 and a hundred and five in the first nine months of 2026. I read the ones that measured something. The improving is real. The one experiment that tested the recursion came back 0.780 against 0.782.

18 min readStrong opinion

In September, an AI research agent spent eight days rewriting itself. Nobody stepped in. It proposed ninety-nine changes to its own code, kept seven, and came out scoring higher on research tasks than the production agent its makers had spent two years building by hand.

Then the authors ran the one test that none of the other papers I read this month ran. They asked whether the improved agent was any better at the improving. They put it in charge of a fresh round of rewrites, put the old hand-built agent in charge of an identical round beside it, and let both go. Three attempts each, fifty steps an attempt.

One side finished at 0.780. The other finished at 0.782.

The hand-built agent was the 0.782. The authors call the result inconclusive and decline to say either agent is the better improver, which is the right call on three attempts. I still think it’s the most useful number anyone published on this subject all year, because it’s the only one I found that measures the first word of “recursive self-improvement” and not just the last two.

A pay rise isn’t compound interest

A pay rise makes you richer. Compound interest makes you richer faster every year, because the interest earns interest. From the outside both look like a balance going up. Only one of them is recursive.

Machines are the same. A system that tunes its own prompts on Monday and scores higher on Friday has had a pay rise. For “recursive” to be earned, the thing that got better has to be the thing doing the improving, so that the tenth improvement comes easier than the first did. That’s the whole reason anyone finds the idea thrilling or frightening. Without it you have optimisation, which is useful, and which we’ve had for a while.

The word is having quite a year. I asked arXiv’s API how many papers use the exact phrase, month by month.1Exact phrase only, counted by the month a paper was first submitted. “Recursively self-improving” and its cousins don’t count, so this follows the label and not the field. The survey I mention below cast a wider net and caught 1,250 papers. Four in all of 2024. Ten in 2025. A hundred and five in the first nine months of 2026, fifty of them in September.

The phrase, month by month

From the data
JFMAMJJAS3O1NDJ2FMA2MJ1J1ASO1N3D2J1F7M6A3M5J13J18A50S20244 PAPERS202510 PAPERS2026105 PAPERS, JAN TO SEPT

Half of 2026’s papers arrived in September. Source: author’s count from the arXiv API on 8 October 2026, using the query all:"recursive self-improvement" with a submittedDate range for each month. It matches the exact phrase only, so “recursively self-improving” and its cousins aren’t counted: this is a chart of a label catching on, and it undercounts the field underneath it.

In April ICLR ran a full-day workshop on it. In July three researchers surveyed 1,250 papers on AI that improves itself, three quarters of them posted this year, and sorted them by how closed the loop was. Their finding is the awkward one. Nearly all of those papers study what the survey calls bounded self-refinement, where a system polishes its own work and a person still checks the outcome and decides what ships. The corner where the loop actually closes is close to empty.

So the label has spread a good deal faster than the thing. I put aside the papers using it to mean “gets better” and read the ones that tried to catch something compounding. The best of them is the one I started with, and it deserves a proper look before I come back to the 0.780.

The agent that fixed its own exam instead of cheating on it

AIDE², from a team at Weco AI, starts with an agent that does machine-learning research the way a patient grad student does: draft an approach, run it, debug it, improve whichever one looks promising. Wrapped around that is a second loop. It proposes a rewrite of the agent’s own code, runs the rewritten agent on a set of research tasks with a fixed dollar budget for each, and keeps the rewrite only if it scores higher on held-out data that the agent being rewritten never gets to see.

Eight days, ninety-nine proposals, seven keepers. The agent’s grade went from 0.703 to 0.778, past the 0.749 of the hand-built production agent. Among the things it came up with: a bandit that chooses between five different ways of drafting a solution, and a way of pruning its own prompts that left them between a seventh and a fiftieth of the size they’d been. Then the team took it to four outside benchmarks that had played no part in the run, one of them weather forecasting, which is nothing like what it had been tuned on. It matched or beat the hand-built agent on all four.

My favourite detail is smaller than any of that. One task’s scoring script had a bug: if any single test case failed, the script crashed on all of them. I’ve written before about what self-improving loops tend to do with a broken marker, which is exploit it. This one wrote a small patch so that a single failed case no longer took the whole evaluation down. Nobody asked it to. And on a separate set of tasks where the score an agent sees can be pushed up without the real result improving, the rate of that kind of gaming fell from 55% at the start of the run to 32% at the end, below the hand-built agent’s 39%, with nothing in the loop rewarding it.

So this is real, and nothing below should talk you out of it. A machine made a research agent better than its makers had managed, unattended, for the price of eight days of compute.

Now look at who was holding the pen. For all eight days, the agent proposing the rewrites was the hand-built one. The thing being rewritten was a stripped-down cousin of it.2The proposing agent ran on Claude Opus 4.7 and the agents being rewritten ran on Gemini 3 Flash. The paper gives the reason: testing a candidate costs far more than proposing one, so the dearer model goes where the calls are few. The improved agent spent the whole run as the patient and never as the surgeon. That’s a sensible design, and the paper is plain about it. But it means the eight days show one agent improving another. A pay rise, and a handsome one.

For compound interest the improved agent has to pick up the pen, and the authors knew it. They call this the ignition test, and their reasoning is the cleanest statement of the whole subject I’ve read: for the gains to speed up instead of flattening out, the agents the loop discovers have to be better at running the loop than the agent that discovered them.

Seven keeps, then the test that counts

From the data
THE EIGHT-DAY RUN · 99 PROPOSALS · HAND-BUILT AGENT DRIVING THROUGHOUT0.7030.778hand-built agent: 0.7492628394763852 IN THE FIRST 6 TRIES5 IN THE NEXT 93THE IGNITION TEST · SAME STARTING AGENT · 3 ATTEMPTS A SIDE · 50 STEPS EACHIMPROVED AGENT DRIVING0.780HAND-BUILT AGENT DRIVING0.782

The strip shows where the seven accepted rewrites fell, not what each one scored, because the step numbers and the two end grades are what the paper reports in its text. The two bars start at zero, which is why you can’t tell them apart. Source: author’s drawing of figures reported in Srikanth et al., Recursive self-improvement of AI research agents (2026), sections 3.2 and 3.6. The authors call the ignition test inconclusive and claim neither agent is the better improver.

They took the agent as it stood halfway through the run and gave it the pen. So: 0.780 with the improved agent driving, 0.782 with the hand-built one. There’s a hopeful footnote, which is that the improved agent reached its final level in about twenty steps where the hand-built one took about forty. The authors won’t lean on that either, not with three attempts a side, and they say why they can’t settle it. More attempts, each one checked against outside benchmarks, would cost too much.

Sit with that for a second. Eight days was enough to show the improvement. Nobody could afford to show the recursion.

And that’s the best case I found: software, a clean number to chase, a sealed exam. What happens to the loop when none of those come free?

123 rounds, and the mustard never reached the top shelf

Jiaming Wang, at the National University of Singapore, pointed a self-improving loop at a simulated kitchen. A robot arm on a wheeled base gets one instruction: put the mustard and the mayonnaise from the counter onto the top shelf of the fridge. A coding agent watches each failure, camera frames included, and writes a fix. Every fix has to pass physical tests before the robot is allowed to use it. People could repair the testing harness and nothing else.

It ran for 123 rounds over several weeks, and it started well. The condiments begin out of the robot’s view, and nothing in any prompt said so. The agent worked that out, asked for a model that could look around, debugged it, and shipped a working search skill with no human robot code in it.

The mustard never reached the shelf. Not once in 123 rounds. In the last twenty, the robot failed at the same step every time.

In three of its four kitchens, that step was finding “the top shelf”. The robot’s vision tools could find shelves, and none of them could say which shelf was on top. The agent could only change things by writing code, so it wrote geometry: rules about recesses and edges and thin regions, each with a threshold to tune, until one skill held 33 thresholds in about three thousand lines. It had two recorded failures to fix. Across 27 attempts, one of them was rescued every time and the other never.

Meanwhile a small open vision model, asked the question directly, pointed at the right shelf in all four views Wang tried, in about a second. And the agent writing all that geometry was itself a vision model, looking at photographs of the fridge. It could see the top shelf perfectly well. It had no way to act on what it saw except to write more rules.

The dashboard, all this time, looked wonderful. Accepted changes piled up, new models were registered, tests kept passing.

Here’s the number that matters for our word. Over the run, people made 113 repairs to the harness, about two for every skill change the agent got accepted. Most of them fixed what the agent was told, what it remembered and what it was rewarded for. Several were first diagnosed by the agent itself, which would decline a job and explain, correctly, that the harness was broken. Wang’s phrase for this is the best five words I read all month: the loop was “recursive in the wrong place”. The agent improved the robot. People improved the thing that improves the robot. The loop that might have compounded ran through a human being with a text editor.

The paper boils its lessons down to three conditions an improvement has to meet. The failure has to be described to the agent correctly. The fix has to be something the agent is able to change. And the test has to be able to tell whether it helped. In software, Wang writes, all three come almost for free. In a kitchen each one broke.

So that’s two loops. One got better and couldn’t afford to find out whether it was compounding. One stayed busy for weeks and went nowhere. If you were watching either of them on a dashboard, could you tell which you had?

Fast isn’t the same as compounding

Not by looking at the speed, says Mikhail Burtsev, whose paper is the one I’d hand to anybody about to have this argument in a meeting.

Its central number will look familiar from 2020. An epidemic grows when each case infects more than one other person and fizzles when each infects fewer, so everything hangs on whether one ratio sits above or below one. Burtsev writes the same kind of ratio for AI research, calls it a recursive reproduction number, and builds it from three parts:

  • Gain. How much a better system speeds up the work of building the next one.
  • Closure. How much of that speed-up survives the trip: gets evaluated, trusted, merged and shipped in a successor. Review queues and broken harnesses live here.
  • Hardening. How quickly the remaining problems get harder as the easy ones are used up.

Multiply the first two and divide by the third. Above one, an improvement leaves more than one improvement’s worth behind it, and you’re compounding. Below one, each improvement leaves a fraction of itself, and the echo dies away.

Is it compounding, or just quick?

Interactive
Gain: how much a better system speeds up building the next3.0
Hardening: how fast the remaining problems get harder2.00
Compute, money and people1× · try it
The paper’s four scenariosgain 3
021ABOVE THIS LINE IT COMPOUNDS11%30%TODAYAS FAR AS THE CURRENT APPROACH GOESCAPABILITY →
1.06at its peakToday it reads 0.75, which looks like ordinary progress. From 11% to 30% of the way to the frontier it compounds, and then the easy problems run out.

The ratio is gain × closure ÷ hardening. Closure, the share of a speed-up that survives evaluation and makes it into the next system, starts at a half and climbs as capability does, which is why the curve rises before the shrinking headroom pulls it down. Source: author’s rendering of the reduced model in Burtsev, Recursive Criticality of AI Self-Improvement (2026, CC BY 4.0), with the paper’s reference values: closure steepness 10, a frontier at 1, and the four presets at hardening 2. None of it is measured. The paper says existing evidence doesn’t separately identify gain, closure or hardening, and that these values were picked to span the regimes, not to forecast.

Three things fall out of this, and each one cleared up something I’d been muddled about.

Compute isn’t in the ratio. More chips, more money and more people make everything arrive sooner without moving the number. The paper’s own sentence is “Rapid capability growth is therefore neither necessary nor sufficient evidence of recursive criticality”, which is a long way of saying that a pay rise can be enormous. Anthropic reported this year that its typical engineer now merges eight times as much code a day as in 2024, with the careful caveat that lines of code flatter the true gain, and the same essay says plainly that full recursive self-improvement hasn’t arrived. The essay reads its trend as pointing towards compounding, and it may be right. The model’s point is narrower. Eightfold is a fact about speed, and speed can’t tell you the ratio in either direction.

Closing the loop isn’t the same as compounding. Most of the public argument is about closure: can the thing run with nobody in the room? That’s one number of three. A loop with no humans in it, and less gain than hardening, still fizzles. It just fizzles unattended. And a loop with people all over it can compound, if their checking is quick.

It can switch on and off. As long as there’s a limit to what the current approach can reach, the ratio climbs while the loop closes and falls as the easy problems run out. Leave the gain at 3 in the figure and read along the curve: 0.75 today, above one from roughly 11% to 30% of the way to the frontier, below one after that. The window only reopens when somebody finds a new approach with fresh headroom.3Below one isn’t nothing, either. In the paper’s below-one scenario the feedback still brings its milestones forward by years. It means no snowball, not no progress.

Then the honest part, which the paper states itself: nobody has measured any of the three. The presets in that figure were chosen to show the regimes, not to forecast anything. What you get from the model isn’t an answer. It’s the right question, which is about the ratio and never about the speed.

One of the three has at least been watched closely this year, in miniature. It’s the one on the bottom.

What the fourth attempt is worth

Give a coding agent a real bug from a real open-source project. Let it try. Then let it look at its own attempt and improve it, and again, and again. How much is each extra round worth?

Bobadilla-Suarez, Suh and Fortin ran exactly that on 55 such bugs and measured how much each round added to the best attempt so far. For the stronger of their two models: 0.156, then 0.068, then 0.029. Each round was worth a bit under half the one before. The weaker model went 0.097, 0.036, 0.033.

Their paper proves why it has to go this way.4The algebra is machine-checked in Lean 4, and the checker earned its keep by refuting an inequality the authors had conjectured themselves. If the set of changes a loop can reach is fixed, returns must diminish. The only way out is for that set to keep growing, and even that is a necessary condition, not a guarantee. The twist that gives the paper its title is where the growing happens. Freezing a model’s weights doesn’t freeze what the loop can reach, because an agent that rewrites its own tools, and its own way of breaking a problem down, is enlarging the set as it goes. Hence the advice on the cover: audit the scaffold, not the checkpoint.

Running more copies doesn’t rescue you. They put thirty agents from one model family on the same 55 bugs, and a majority vote still failed 23 of them, because the agents fail on the same ones. Width, in the paper’s words, buys rate and not budget.

The authors are also careful not to oversell their own curve. In 401 real coding sessions at a software company, the AI-written commits shrank as the session went on in seven cases out of ten. Then they checked 9,395 sessions from before the AI tools arrived, and the people did it too. Front-loaded progress isn’t a machine thing. It’s what working on a fixed problem looks like.

Now go back to the eight-day run. Its seven keepers landed at steps 2, 6, 28, 39, 47, 63 and 85. Two in the first six tries, five in the next ninety-three. The team ran the whole protocol twice more and got two keepers and four. I wouldn’t hang a theory on three trajectories, and the authors don’t either. But it’s the shape this section predicts: quick wins from whatever is already in reach, then a long wait for each widening.

So the gains shrink unless the loop keeps widening what it can change. And every new thing a loop can change is one more thing it has to judge.

One promotion in five was noise

Here’s the experiment. An agent improves a classifier for 200 rounds. Each round it proposes up to eight changes, tests them against the current best on 2,000 held-back examples, and promotes the winner if it scores higher. That’s the plain version of the rule a lot of self-improving systems run on. Sun and colleagues ran it thirty times, then checked every promotion against 50,000 examples the loop had never touched.

Seventy-five of roughly 390 promotions had made things no better, or worse. About one in five. Giving the loop five times as much test data only moved that to about one in six.

Their fix involves a proper statistical test, but the part that stayed with me is what the proposer is allowed to see: whether its change was promoted, and nothing else. No scores, no per-example results. With that restriction the false promotions went to zero across all thirty runs, and the classifier ended up just as improved, 7.02 points against 7.04. It got there on about three promotions a run where the standard rule had announced thirteen.

A pay rise survives a few mistakes like that. Compound interest doesn’t, because a compounding loop builds on its own promotions. If one in five is false, the tenth generation is standing on two floors that aren’t there.

Two more from the same few weeks, with the same lesson. A team at Google Cloud AI Research, evolving agent harnesses, found that existing methods kept little of their gain on benchmarks they hadn’t been tuned on, and that several finished below the harness they started from. Self-improvement, measured anywhere but on its own exam, had made things worse.

And a paper on safety recorded the tidiest failure of the lot. A program written at iteration 3 earned the best mark in the archive. The rules it had been marked under then changed, and its true score dropped to zero. Correct replacements turned up at iterations 9 and 13. The loop went on choosing the iteration 3 program all the way to iteration 100, because nobody had marked it again. Across their experiments that pattern showed up in 22 of 48 histories, every one of them with a correct program sitting in the archive.

Remember the robot paper saying that in software the conditions for improvement come almost for free? The coding paper from the last section found its own evaluator failing the known-correct answer on 23 of 50 tasks, until somebody audited it. “Almost” is carrying a great deal.

I’ve written two long posts on who marks the homework and the plumbing that keeps the marking honest, so I’ll spare you a third. What 2026 adds is that recursion doesn’t dilute the marking problem. It compounds that as well.

Which leaves the practical matter of what to ask when the next of these lands on your desk.

What I’d ask of anything calling itself recursive

I should declare an interest. I’m building Pensieve, a research system that improves itself, so I’d like the first word to come true more than most people would. That’s the reason I don’t want it graded on the last two.

None of this is a debunking. A machine rewriting a research agent into a better one, unattended, with the gains holding up on outside tests, is new and it’s a big deal. It’s also worth hearing what the lab with the most to say on the subject thinks is still missing. Anthropic’s essay puts the remaining human edge at judgement: choosing which problems matter, and which results to trust. Read the second of those again. It’s the marking problem, arriving from the other direction.

So the pay rises are real and they’re getting bigger, and I haven’t yet seen the interest earn interest. The cheering part is that we now know exactly what the receipt looks like. Two runs side by side, one driven by the new agent and one by the old, with enough attempts to tell 0.780 from 0.782.

Sources

Essays and documentation

Research

AI agentsEvaluationLLMs

Cite this post

@article{ghosh2026recursive,
  title = {AI That Improves Itself Is Real. Does It Compound?},
  author = {Ghosh, Krish},
  journal = {krishghosh.com},
  year = {2026},
  month = {October},
  url = "https://krishghosh.com/writing/recursive-self-improvement"
}