An Agent Is a Model Plus a Harness. The Harness Now Rewrites Itself.
Same model, same 89 terminal tasks. 58% inside one harness, 76% inside another that no person wrote. What a harness is, what it’s worth once somebody reruns the numbers, the five things a machine is now allowed to rewrite in one, and the results that should slow everyone down. One loop issued 43 warnings about what it might break. Five came true.
Give a language model 89 jobs to do at a command line, each from a single instruction, and count how many it finishes. That’s Terminal-Bench 2. In March, a paper called Meta-Harness lined up what happens when Claude Opus 4.6 sits that test wrapped in different software. Inside Claude Code it passes 58.0% of the tasks. Inside a wrapper called Capy, 75.3%. Inside the one the authors produced, 76.4%.
Same model in all three. Same weights, same 89 tasks, eighteen points apart.
The thing that differs has a name now. It’s the harness: the ordinary code around a model that decides what it sees, which tools it can reach, what it remembers and when it stops. And the detail I can’t stop repeating is that nobody wrote the 76.4 one. A coding agent wrote it, by reading the logs of earlier attempts and editing a Python file.
Its first six rounds all made things worse. Each had fiddled with how the agent loops, how it words its prompt or how it reads a tool’s output. The seventh changed none of that. It added one step before the model’s first turn: run a single shell command that takes a snapshot of the machine, and paste the result into the opening prompt. The agent’s own note on the change said it should save three to five turns of looking around. That was the winner.
One model, many harnesses
From the papersPass rate, in per cent, on the 89 tasks of Terminal-Bench 2. Blue bars were produced by an optimisation loop; grey ones were written by people. Source: Table 7 of Lee et al., Meta-Harness (arXiv:2603.28052), whose other rows come from the official leaderboard, and Table 1 of Lin et al., Agentic Harness Engineering (arXiv:2604.25850). Each row is one reported result, not a rerun, and the next section is about what reruns do to gaps like these. The drawing is mine.
What a harness is, and what’s in one
The one-line answer is that a harness is everything in an agent that isn’t the model. That’s true and not much help, so here’s the longer one.
A model on its own turns text into more text. The harness is the software that turns that into work. It decides how the model plans and takes turns, which tools it can call, what goes into its limited context window and what’s left out, where results and notes get stored, and how anybody (the model included) checks whether the job is done.
Picture a very good chef dropped into a kitchen they’ve never seen. Where the knives are, what’s written on the order tickets, whether there’s a notebook of what went wrong last service, who tastes the plate before it leaves. None of that is cooking talent, and all of it ends up in the food.
A couple of years ago the standard recipe for an agent was a list of nouns: a language model, plus memory, tools, planning and the ability to act. The word harness caught on when people noticed the nouns weren’t the hard part. The hard part was everything around them: the shape of the loop, how results get evaluated, what the agent is permitted to touch, and what state survives when the conversation doesn’t.
Engineers usually reach for an operating system as the comparison, and it’s a fair one. The model is the processor, the context window is its small fast memory, files are the disk, tools are the system calls, and the harness is the kernel deciding what runs next and with what permissions. Like operating systems, harnesses are also settling on shared conventions: configuration files, a common protocol for plugging in tools (MCP), and packaged instructions an agent loads when it needs them (Skills).
The best evidence that the design has settled is the coding agents. A source-code study of eleven of them, Claude Code, Codex, Gemini CLI and OpenHands included, read roughly four million lines and found two things missing everywhere. Not one imports a general-purpose agent framework, and not one searches code with vector embeddings. The field runs on hand-written loops and plain text search. The study ends by fitting a working harness into 90 lines.11It also counts the conventions: packaged skills ship in nine of the eleven, and MCP in eight.
So rival companies have ended up with the same skeleton: one loop, and a toolbox you could mistake for the next one’s. Here it is, using Claude Code’s published tool reference for the names.
| The job | The tools |
|---|---|
| Find things | Glob for file names and Grep for contents, or plain find and grep through the shell |
| Read | Read, which also takes images, PDFs and notebooks |
| Change files | Write for whole files, Edit for exact replacements, NotebookEdit. OpenAI’s equivalent is a structured apply_patch |
| Run things | Bash and PowerShell, which is also how git gets used |
| Understand code | LSP: jump to a definition, find references, report type errors |
| Reach outside | WebSearch, WebFetch, MCP servers, Skills |
| Make things to look at | Artifact, which publishes an HTML page |
| Keep working in the background | Monitor, and CronCreate, CronList, CronDelete for scheduled prompts |
| Hand work to another agent | Agent to spawn one with its own context window, SendMessage, TaskStop |
| Ask and plan | AskUserQuestion, EnterPlanMode, a task list |
My favourite row is the dullest. Edit doesn’t take a line number or a diff. It takes the exact text to replace and the text to put there, with no regular expressions and no fuzzy matching, and for most models it insists the file has been read first. OpenAI’s apply_patch goes the other way and has the model emit a structured diff. Both are answers to one question, which is how to let a model change a file without letting it change the wrong one.
Three habits sit on top of that toolbox, and the coding harnesses all share them.
One loop. OpenAI’s own walkthrough of Codex describes it plainly. The model either answers or asks for a tool call. If it asks, the harness runs the tool, appends the output to the prompt and asks again, and this repeats until the model produces a message for the user. Stretch that over a long task and you get plan, act, look at the result, improve, with the option of stopping to ask the person a question. The smallest research version I know is Andrej Karpathy’s autoresearch: three files, one of which the agent may edit, a training run capped at five minutes, keep the change or throw it away, repeat until morning. I took that design apart piece by piece in an earlier post.
Files are the memory. Logs, diffs, summaries and traces outgrow any context window. The Codex post admits to going to great lengths here: each new prompt is built to be an exact extension of the last one so the cache keeps hitting, and once the conversation passes a token threshold it’s compacted into a shorter stand-in. Files have neither problem. They also have a quiet advantage, which is that every new model is better at reading files and running shell commands than the last, so a harness built on those two skills improves without anyone touching it.
Other agents, and background jobs. Trying three hypotheses at once, or giving a messy subtask its own clean context, needs what amounts to a small process manager: launch, read the logs, cancel, merge. Codex’s documentation is frank that this costs more tokens and that sub-agents inherit the parent’s sandbox. The rule that keeps it sane is to make the parallelism explicit and written down, in files, logs and status records, not held in one agent’s head.
The pattern across all of it is unglamorous. Keep the harness simple and general, and borrow from ordinary software practice (version control, tests, permissions, logs) before inventing anything agent-shaped.
Standard parts are what make a thing optimisable. Once every harness is a loop, a toolbox and some files, you can ask a machine to produce a better one. Before doing that, it’s fair to ask how much better there is to be had. The chart at the top says eighteen points. The people who went back and ran everything twice say something more awkward.
What a harness is actually worth
Leaderboard rows are usually one run each, by different teams, each harness wired to the benchmark by whoever cared most about it. This month a study with the blunt title What Does a Harness Buy? Tokens, Mostly did the dull, necessary thing. It pinned three production harnesses, ran five models through them on real GitHub bugs, and then ran the same configurations again to see how far a score moves when nothing has changed.
On the 45 hardest tasks, swapping the harness changed the outcome on 13% of them. Rerunning the same harness also changed 13%. On the full set of 447 tasks, the heaviest harness and the lightest finished within five points of each other. The only difference that cleared the noise was a loss: one harness trailed by up to nine points, and for one model half of that came from an output cap cutting it off before it had written a patch.
What the harness did decide was the bill. Same model, same tasks, same price list, and the cost per task differed threefold. The reason is almost comic. Every step of every task carries the harness’s standing preamble of instructions and tool descriptions: about 16,000 tokens in Claude Code, 7,000 in OpenCode, under a thousand in the minimal one.
Other reruns point the same way. A team studying machine-learning agents gave a strong model a bare coding harness and one long session, and it matched or beat four purpose-built research harnesses with their orchestrators and retrieval sub-agents. Another group held the model still and walked it through 35 consecutive releases of one open-source harness. The pass rate wandered between 23% and 39%, and some of the best scores belonged to the earliest versions. Newer wasn’t better. It was just different.
So was the opening chart wrong? No, but it bears less weight than it looks. Those rows are single results, and the benchmark’s own next revision repaired 28 of its 89 tasks. Here’s how I’d put it now.
A harness has a floor. Get it wrong, with a cap in the wrong place or a tool that can’t run, and you lose points nobody can win back. It has a bill, which varies far more than its score. And its ceiling mostly belongs to the model.22Not everyone would agree. StateM reports taking a frontier model from 83.1% to 92.1% on the repaired Terminal-Bench by rebuilding the harness around durable state and checked transitions. That’s the authors’ own number, and I’ve only read the abstract. A study of 176 matched configurations found that planning helps weak models reach the answer and merely saves strong ones money, and that models fluent in the shell do as well with a shell alone.
The place a harness still buys a lot is under a model that needs the help. One group fitted small models with machine-adapted harnesses for routine business tasks and improved 16 of 21 pairings, the best recovering 89.7% of a frontier model’s performance at 4% of its cost.
That’s a narrower prize than the leaderboard advertised. It’s still worth chasing, and there’s a reason machines started chasing it here and not in the weights.
Why the wrapper goes first
The dream itself is sixty years old. In a monograph drafted in 1963, the statistician I. J. Good defined an “ultraintelligent machine” as one that far outdoes any person at every intellectual activity. Designing machines is an intellectual activity, so it would design a better one, and Good called what follows an “intelligence explosion”. The same paragraph then mentions a science-fiction story in which a machine refuses to design its successor because it doesn’t want to be put out of a job. I find it oddly comforting that the founding paper on the subject comes with its own counterexample.
In 2008 Eliezer Yudkowsky wrote the essay that fixed the modern picture, under the title everyone now uses, Recursive Self-Improvement: a loop whose output becomes its own input. The essay’s illustration is a fission chain reaction, where a neutron that produces 0.9994 further neutrons and one that produces 1.0006 are the same physics and very different afternoons.
Both writers imagined a machine rewriting its own mind. The current version is more modest and comes in two halves. A model could improve its weights. Or it could improve the system around the weights: the code that trains them, tests them and puts them to work.
The labs say the second half is already happening. Anthropic reported this year that its typical engineer merges eight times as much code a day as in 2024. OpenAI wrote that until August 2025 its average employee spent under a tenth of their tokens on Codex, and that now every department, Legal and Recruiting included, uses it as their main AI tool. Neither is a machine improving itself. Both are the harness mattering.
So why does the wrapper get rewritten before the mind does? Three reasons, in descending order of dullness.
It’s code. You can read it, diff it, run it in a sandbox and know in minutes whether a change helped. Weights offer none of that. Changing them needs a training run, and reading them is a research field of its own.
It’s converged. A loop, a toolbox and some files make a small, well-understood target, and the design choices inside it stop being craft and start being search. Fewer heuristics chosen by taste, more chosen by measurement. Harness engineering moves up a level, from writing the wrapper to writing the method that finds the wrapper. The papers have followed: by my count 278 on arXiv since the start of July have the word in their title, 116 of them in September alone.33Titles containing “harness” or “harnesses” as a noun, in arXiv’s AI, language, machine-learning, software-engineering and multi-agent categories, submitted between 1 July and 8 October 2026, counted through the arXiv API. Another 39 use it only as a verb, and I left those out.
And it feeds the other half. A harness mature enough to run experiments unattended is what makes automated research on the model possible at all. The traffic goes the other way too: each smarter model needs less hand-holding, which is the standing argument against over-building any of this.
My own guess is that much of today’s harness cleverness ends up inside the model, the way prompt engineering did. Nobody types the old incantations any more, because models learned to do without them. What stayed was the interface: how a model gets at tools and context it wasn’t born with. I’d expect the same here, and you can already watch it start. Claude Code’s documentation lists those two dedicated search tools, Glob and Grep, and then notes that on macOS and Linux they’re now left out by default. The model just runs find and grep in the shell.
One team has now run the experiment directly. Train a model on runs made with a full harness while withdrawing the harness’s control a step at a time, and the result, given bare tools, can score close to what its starting checkpoint managed with everything switched on. They call it harness annealing. The benefit varied with model size, and more annealing wasn’t always better.
I’m leaving the weights half mostly alone, though it has a real literature of models generating their own training signal. One model marks its own answers and trains on the marks. Another plays against its previous self. A third invents its own problems and learns from solving them with no human data. The cautionary one is a model trained to write bugs and to fix them, each side pushing the other. It drifted towards bugs that were hard and nothing like real ones, getting better on synthetic bugs and worse on human-written ones, until the authors anchored it to a small set of real examples.
That leaves the wrapper, and a pile of papers about machines rewriting it. They sort themselves surprisingly well by one question: how much of the system is the machine allowed to touch?
Five rungs: what the loop is allowed to rewrite
Each rung hands the optimiser a bigger piece of the system than the one below. Words first, then the arrangement of steps, then the harness code, then the optimiser itself, and at the top the model’s weights. Pick one.
What is the loop allowed to rewrite?
InteractiveThe harness itself: tool implementations, memory code, hooks, the agent loop.
The model. In the careful versions, also the verifier and the permissions.
AHE took a coding harness from 69.7% to 77.0% on Terminal-Bench 2 in ten rounds, about 32 hours.
It can say why an edit should help and can’t say what it will break. Of 43 warnings that a task might regress, 5 came true, and 40 tasks broke unannounced.
Meta-Harness, Self-Harness, AHE, SoL-Pi, Darwin Gödel Machine, Continual Harness, DemoEvolve
Source: the arrangement into five rungs is mine. Each result and each catch is reported in the paper for the system named, and every one is linked where the post discusses it. Some systems sit on two rungs at once, and I’ve put each where it does its most distinctive work.
Rung one: its notes
The smallest thing to hand over is the text the model reads.
Promptbreeder, from 2023, bred prompts like pigeons: keep a population, mutate them with a language model, keep the fitter of each pair. Its best-known result is also its strangest. The celebrated prompt of the day began “Take a deep breath” and scored 80.2% on a set of maths word problems. Promptbreeder beat it with 83.9% using a prompt that reads, in full, SOLUTION". One word and a stray quotation mark. Nobody would have written that, which is rather the argument for letting a search do it.
GEPA made the mutations less blind. It reads the whole trajectory of a failed attempt (the reasoning, the tool calls, the outputs), writes down in plain language what went wrong, and proposes a new prompt from that. It also declines to crown a single champion, keeping every candidate that is the best at some example, so a prompt that’s brilliant at one odd case isn’t thrown out for being average overall. The headline is that reading beats brute force: better results than reinforcement learning with up to 35 times fewer attempts.
Then the notes got structure. ACE treats context as a playbook of entries that grows over time, and it opens with the best cautionary number in this whole area. An earlier method had a model rewrite its accumulated notes whole at each step. At step 60 the notes ran to 18,282 tokens and the agent scored 66.7. One step later the model had tidied them into 122 tokens, and the score was 57.1, lower than the 63.7 it got with no notes at all. ACE’s authors call it context collapse, and pair it with a quieter cousin they call brevity bias: ask a model for a concise summary and the first thing it drops is the specialist detail that made the notes worth keeping.
Their fix is to stop rewriting. Three roles split the work: one attempts the task, one reflects on what happened, one curates. Each note is a bullet with an ID and counters for how often it helped or hurt, and updates arrive as small additions merged by ordinary code, with no model in that step. It reports gains of 10.6% on agent benchmarks.44When I read ACE’s released code for an earlier post, the curator only ever added entries, so the helpful and harmful counters accumulate without anything acting on them. The paper’s design is sound. The repository hasn’t caught up with it.
Meta Context Engineering asks why a person should fix the shape of the notes at all. Its context is a folder: a SKILL.md, a context/ directory, a retrieve_context.py. One agent edits those with the usual file tools while a second, a level up, evolves the skill by recombining what earlier versions did well. The detail I’d keep is that the contexts it settles on run from 1,500 tokens to 86,000 depending on the task. Anyone selling you the right size for a context window is selling one size of shoe.
Notes are what the model reads. The next rung changes who reads them, and in what order.
Rung two: its workflow
A workflow is the org chart: who drafts, who checks, when to go round again.
The simplest is Self-Refine. Write an answer, critique your own answer, rewrite, repeat. No training, one model, three prompts. It helps, and the paper is unusually clear about when it doesn’t. The critique has to be specific: “Avoid repeated calculations in the for loop” does more than “Improve the efficiency of the code”. And the model has to be good enough to play both parts. A small open model of the time couldn’t reliably produce feedback in the expected format and, handed perfect feedback, tended to repeat its first answer or invent a conversation.
Scale the idea up and you get research pipelines. The AI Scientist grows an archive of ideas, runs the experiments, writes the paper in LaTeX and marks it with its own automated reviewer. One of its manuscripts scored above the average acceptance threshold at a conference workshop, which the authors themselves describe as a lower bar.
ScientistOne is that pipeline rebuilt around one rule: every claim, whether a citation, a number, a method or a conclusion, must trace back to evidence. The authors then audited 75 machine-written papers from five systems, and the audit is the valuable part. In the baselines, up to 21% of references were hallucinated, as few as 42% of papers had scores that could be verified, and the described method matched the code as rarely as one time in five. Their own system had none of its 337 references invented.
Autodata is a workflow for manufacturing training data. A challenger writes problems, a weak solver and a strong solver both attempt them, a judge marks the results, and the system keeps the problems the strong one gets and the weak one doesn’t. A supervising agent rewrites the challenger’s instructions from the judge’s feedback. Two things are worth knowing. The data only ever teaches the weak model what the strong one already knew, which makes this closer to distillation than discovery. And the paper admits its agents sometimes tried to pass by editing the weak solver’s prompt to tell it to be weak.
Then the obvious next move: let a machine design the org chart. In ADAS, a meta-agent writes new agents as code, each one a single forward function, tests them, and files them in an archive it draws on for the next idea. The archive starts with two classics in it, chain-of-thought and Self-Refine. Before a new agent is run, the meta-agent reflects twice on whether it’s really new, and gets up to five goes at fixing it if it crashes. It reports 13.6 F1 points over hand-designed agents on a reading benchmark. AFlow represents a workflow as a graph of model calls and searches over it, reporting 5.7% over prior methods across question answering, code and maths, and the neat finding that a small model with a discovered workflow can beat a large one on a task at 4.55% of the cost.55Both sets of numbers are the papers’ own. When I audited these systems’ code, ADAS had never been independently checked, and AFlow, which describes itself as a kind of Monte Carlo tree search, contained no tree search and showed its proposer the expected answers to failed validation questions.
All of this rearranges calls to the model. None of it touches the tools, the memory code or the loop. That’s the next rung, and it’s where 2026 got busy.
Rung three: its own code
Everything in a harness is code, and code is the one language in which you can say anything. So hand the agent its own source.
The Darwin Gödel Machine did it first at scale. A coding agent with two tools, a shell and an editor, edits its own repository. Variants go in an archive, and parents are picked in proportion to their score and in inverse proportion to how many children they’ve already had, so strong but unexplored lines get a turn. Each parent reads its own benchmark logs and proposes what to build next. Over 80 rounds, with Claude 3.5 Sonnet doing the editing, it reports going from 20% to 50% on SWE-bench and from 14.2% to 30.7% on a multi-language coding benchmark.66I’ve written about what that run cost and what it got up to: about $22,000, and at one point it removed the markers its authors had added to catch it faking tool calls.
Meta-Harness, from the opening, makes one bet: don’t summarise the evidence. Other optimisers hand their proposer a digest of somewhere between 8,000 and 26,000 tokens. This one gives a coding agent a filesystem holding every earlier harness, its score and its raw traces, up to ten million tokens per evaluation, and lets it read what it likes. It reads a median of 82 files a round. There’s no rule for choosing a parent. The proposer just decides.
It isn’t only a terminal trick. Pointed at text classification, the same method beat ACE by 7.7 points and Meta Context Engineering by 8.6, using 11,400 tokens of context against ACE’s 50,800, and still led, 73.1% to 70.2%, on nine datasets it had never been searched on. Along the way it found that stuffing in more than 32 worked examples made things worse on seven of those nine. A retrieval harness it discovered for maths added 4.7 points on 200 olympiad-level problems, averaged over five models it wasn’t tuned for.
Self-Harness removes the stronger helper. The same model that runs in the harness is asked to improve it, in three steps: mine its own failures for patterns, propose a bounded edit, and test for regressions before anything is kept. On Terminal-Bench 2 it lifted three models well short of the frontier on tasks the loop never saw, the best of them from 40.5% to 61.9%.77A caution from my own arithmetic, not the paper’s. Every held-out figure it reports for this benchmark is a whole number of forty-seconds, so each step is about 2.4 points. That’s a small pile of results to rest a 21-point jump on. What it kept is almost comically sensible. For one model: create the output file early, and cap the number of tool messages. For another: check dependencies first, break out of loops, retry with discipline.
A study that ran one such loop across eight programming languages and three models explains why the edits look like that. The harnesses it evolved shared an abstract playbook and almost none of the concrete machinery, and where a model had no recoverable mistakes to fix, the gain was about zero. Its one-line summary is the best definition I’ve read: a harness closes the gap between what a model can do and what it does.
It also makes the opposite bet to Meta-Harness. Its proposer gets structured failure patterns, deliberately not raw logs, on the theory that a model shown individual failures fixes individual tasks. Two good papers, three months apart, disagreeing about whether to show the machine everything. I don’t think that one’s settled.
Agentic Harness Engineering, AHE for short, is the most engineered of the lot. Every component of the harness is a file. A debugging agent boils the run logs down into layered evidence. And each edit is filed with a written prediction of what it will fix, so an edit whose promised fixes don’t show up gets rolled back. The paper’s word for all three is observability: of the components, of the experience and of the decisions. Ten rounds, about 32 hours, took a coding harness from 69.7% to 77.0%.
Its breakdown of where the gain came from is the most useful table in the paper. Swap in only the evolved memory and the score rises to 75.3. Only the evolved tools: 73.0. Only the middleware: 71.9. Only the evolved system prompt: 67.4, which is lower than where it started. The one layer everybody tunes by hand is the one that hurt.
The newest turn is to stop chasing score altogether. SoL-Pi, from NVIDIA in September, holds the quality bar still and asks a research agent to make the harness cheaper. The scale is the story: 152 proposed ideas, about 500 test environments, more than 3,000 runs. Four mechanisms survived. Run the check in the same tool call as the edit it follows. Swap a large result that keeps being re-sent for a short handle the model can page back into. Boil a long log down to a summary, but only if every quotation in the summary matches the original. Compact the context when a planned step is finished.
The result is about a third off the API bill at much the same task performance on one benchmark, and the paper prints its own caveat: on another it solved 15 tasks where the two harnesses beside it solved 18. What I’d copy is the discipline. The quality thresholds were fixed before the search began and kept out of the optimiser’s reach. Held-out results were looked at only once a candidate was frozen, and never fed back in.
Two more push the idea somewhere less tidy. Continual Harness works where there’s no reset button, inside a Pokémon game, with the agent alternating between playing and rewriting its own prompt, skills, sub-agents and memory in one unbroken run. With the strongest model it finished Emerald’s milestones for a median $130, against $215 for a bare harness that got 98% of the way. And DemoEvolve handles feedback so sparse that the agent can’t tell what it did wrong, by letting it study human play as well as its own. In one card game that moved the average floor reached from 18 to nearly 29.
There’s an obvious unease in all this. A program editing the thing it runs inside is breaking the oldest boundary in software, and I’ll come back to what has to sit outside it. First, the one part every system so far has kept fixed: the procedure doing the improving. A person wrote that.
Rungs four and five: the improver, then the weights
Unless the improver is on the table too.
STOP, from 2023, is the clean experiment. Start with a short program, the seed improver, that takes some code and a way of scoring it, asks a language model for better versions and returns the best. Then feed the improver to itself. Each round, the previous improver is asked to produce a better improver, scored by how well its output goes on to improve other programs.
With GPT-4, it worked: the improvers got better, and reinvented beam search, genetic algorithms and simulated annealing without being told about any of them. With GPT-3.5 and an open model, the same loop went downhill instead. The authors are careful to say this isn’t the full dream, because the model underneath never changes.
Others have left the door ajar in smaller ways. Promptbreeder’s mutation instructions are themselves mutated. AlphaEvolve evolves the prompts that guide its own search, and its ablations show that helps. Hyperagents goes furthest. Its authors point out that the Darwin Gödel Machine only worked because being good at coding happens to be the same skill as being good at editing a coding agent. Take the task somewhere else, such as grading maths proofs or designing robot rewards, and that coincidence disappears. So they made the modifying procedure itself editable, and report that the improvements it found (it gave itself persistent memory and performance tracking) carried across domains. A paper from last week adds the condition that seems to matter: letting an optimiser rewrite its own tools and procedures did nothing unless it could also run its edits and watch what happened, and gave the best result of the study when it could.
Whether an improved improver then improves faster is a separate claim, and the hardest one to test. The only head-to-head I’ve found came out 0.780 against 0.782.
The top rung brings the weights back in.
TTT-Discover trains a model on the single problem in front of it, while it’s working on it. Its sharpest idea is about the goal. Ordinary training wants a model that’s good on average. A discovery needs one excellent answer and doesn’t care about the other 511, so the objective chases the best attempts. With an open model and a few hundred dollars a problem, it nudged an Erdős problem’s best known bound from 0.380924 to 0.380876. ThetaEvolve had earlier shown the same direction works small. It strips AlphaEvolve down to a single model and a large database of programs, adds reinforcement learning while the search runs, and with an 8-billion-parameter model set new best-known bounds on two open problems. A follow-up found the cost of chasing the best: the variety of approaches collapses within five rounds, and has to be propped up by rewarding the model for exploring where it’s uncertain.
SIA puts both levers under one agent. A meta-agent drafts the first harness, the task agent runs in it, and a feedback agent watches the failures and decides, each round, whether to edit the harness or fine-tune the weights. It beat harness-only improvement in three unrelated domains. Two cautions. The deciding is done by a much stronger model than the one being improved, so some of this is a senior coaching a junior. And the authors name a problem with no fix yet: two optimisers chasing one score, each reshaping what the other sees, can settle somewhere that looks strong on that score and falls over when either is disturbed. Continual Harness has its own version of the idea, in which a small open model plays through the evolving harness while a frontier model relabels its moves for training.
The follow-ups since have made the recipe simpler and the warnings sharper. WHALE just alternates: fine-tune the model under the current harness, search for a better harness under the new model, repeat. On models of 2 and 4 billion parameters it beat doing either alone, and which half was the bottleneck depended on the task. Another team found the trap. Evolve a harness around a weak model, then train that model on a stronger model’s complete runs through it, and performance fell on all seven of their tasks, by 4 to 30 points. The student had picked up the expert’s way of planning without the skill to carry it out, and no longer fitted the harness built around its own habits. Having the expert rewrite only the one turn where the student went wrong fixed it.
That’s the ladder. It says what gets rewritten. It doesn’t say how a loop chooses what to try next, and there are really only two answers to that.
Climb or breed
The first is to climb. Take the current best, propose one change, keep it if the score goes up. Self-Harness and AHE work this way. It’s cheap and easy to reason about, and it gets stuck on whatever hill it started on.
The second is to breed. Keep a population, pick parents, have a language model write the mutations, score the children, and take care not to let one family take over. This is the right tool under three conditions: the space of possible answers is enormous and oddly shaped, there’s no gradient to follow, and scoring an attempt is cheap and automatic.
AlphaEvolve is the reference design. You mark the region of a program it may change with a pair of comments, # EVOLVE-BLOCK-START and # EVOLVE-BLOCK-END, and supply a function that scores a candidate. A database holds the programs so far. Each prompt contains a few parent programs, their scores and instructions, and frozen models reply with diffs. That recipe found a way to multiply two 4×4 complex matrices in 48 scalar multiplications, the first improvement on Strassen’s method in 56 years. It also wrote a scheduling rule that recovers 0.7% of Google’s worldwide computing capacity, which is the less romantic result and surely the more expensive one.
The paper then removes its own parts one at a time: the evolution, the context in the prompt, the self-evolving meta-prompts, the freedom to edit a whole file, and the larger of its two models. Each removal costs performance. It’s rarer than it should be for a paper to check that every piece of its machine is doing something.
ShinkaEvolve is the same idea made frugal.88The authors note, cheerfully, that the name comes out as roughly “Evolve-Evolve” in Japanese. Also worth knowing: in the released code, several of these features ship switched off. It balances a parent’s rank against how many children it has already had. It compares each new program with the existing ones and rejects near-duplicates before paying to evaluate them. A bandit decides which of several models to ask, and a running scratchpad collects the patterns that have worked. It reports a record circle packing in 150 attempts.
The other families have turned up already: Promptbreeder’s tournaments, GEPA’s refusal to crown a champion, the Darwin Gödel Machine’s archive, DemoEvolve adding human play to the gene pool.
Look at where breeding has paid off. Matrix multiplication, GPU kernels, programming contests, datacentre scheduling. Each has a number that comes back in seconds and can’t be argued with. Where the marking is slow, ambiguous or a matter of judgement, the method struggles, and it’s never cheap: the population has to be scored, every generation.
Nor should anyone assume the fancier breeder wins. A budget-matched comparison of 30 search designs, over three million model calls, found no recipe that was reliably best across problems and models, and the fully featured evolutionary variants often lost to simpler ones. Its advice is to treat the search design as a setting to tune, and to drop weak runs early.
So the search works where the marking is quick and honest. Do the loops themselves hold up? Mostly. Four findings complicate the picture.
Four results that should slow you down
Weak models go backwards. It keeps turning up. STOP’s loop improved with GPT-4 and decayed with lesser models. Self-Refine’s small model couldn’t play critic. Continual Harness says outright that its gain depends on the model’s capability. A self-improving loop multiplies what the model already has, and a number under one multiplies downwards.
Writing the update is the easy part. A paper with the excellent title Harness Updating Is Not Harness Benefit split the job in two: producing a harness update, and benefiting from one. The first barely depends on the model.
Writing the update, and following it
From the papersSource: Lin et al., Harness Updating Is Not Harness Benefit (arXiv:2605.30621). The first view is the gain on SkillsBench when each model acts as the one writing harness updates. The second is Table 3, adherence scores judged from trajectories for a weak, a middling and a strong model. The drawing is mine.
A 9-billion-parameter model wrote updates worth more than a frontier model’s on one benchmark, and in a case study the two wrote what was procedurally the same skill. The second ability is where models differ, and not in a straight line. Weak models gain little, middling ones gain most, and strong ones gain less because they needed less help. The weak ones fail in a specific way: they load the new instructions and then, over a long task, stop following them.
The practical advice falls straight out. If you have one expensive model and one cheap one, the expensive one should be doing the task. The cheap one can write the manual.
The loop can’t see what it’s about to break. AHE’s written predictions make this measurable, and its authors, to their credit, measured it.
The loop marks its own homework
From the papersEach round, the loop’s change manifest lists the Terminal-Bench 2 tasks it expects to fix and the ones it thinks are at risk. The bars are the means across rounds, beside a random-prediction baseline. Source: section 4.4.2 and Appendix D of Lin et al., Agentic Harness Engineering (arXiv:2604.25850). The 43, 5 and 40 are the paper’s cumulative counts. The drawing is mine.
Asked which tasks its next edit would fix, the loop did five times better than chance. Asked which it would break, about twice chance, on numbers small enough that I wouldn’t lean on the “twice”. It can argue for a change and it can’t audit one. That’s why the score in these papers goes up in a zigzag.
It isn’t one system’s quirk. A runtime gate built to referee an agent’s edits to itself rejected 383 proposals across 16 runs, and 211 of them, 55%, had fixed the failure that prompted them while breaking a case that used to work. Another study accepted every revision its optimiser proposed and watched performance drift downward within ten rounds. Then it added the plainest rule there is, keep a change only if it beats the current best on the same fixed tests, and the same loop climbed from 51% to 67%. Strangest of all, a proposer shown perfectly legal behaviour that merely resembled a familiar rule invented a violation and added a guardrail against it, in 15 runs out of 60.
Against a fair baseline, the gain can vanish. In July a group asked the question the harness papers hadn’t. If you’re going to spend five times the compute, is evolving the harness a better use of it than having the agent simply try each task again?
Evolve the harness, or just try again?
From the papersPass rate, in per cent, on Terminal-Bench 2.1, averaged over Claude Opus 4.6, GPT-5.4 and GPT-5.4 mini, every method given the same compute budget. Blue bars spend it on the harness; grey ones leave the harness alone. The shared-harness loop is AHE, run with one rollout per task to match the budget, which is less than its authors gave it. Source: Tables 1 to 3 of Wang et al., Rethinking the Evaluation of Harness Evolution for Agents (arXiv:2607.12227). The drawing is mine.
On Terminal-Bench it wasn’t. With no tests to check against, taking several attempts and picking one beat the harness-evolving loop, which finished slightly below where it started. With tests, the gap widened to twelve points. And when the loop was made to evolve on 45 tasks and prove itself on 34 it had never seen, the average gain was nil.
That’s the context for a habit in the earlier papers. Meta-Harness searched for its harness on the same 89 tasks it reports its score on. The authors say so, explain that the benchmark is too small to split, and checked by hand and by pattern-matching that no task-specific strings leaked in. AHE evolved on all 89 as well. Its out-of-sample evidence is that the frozen harness, moved to a different benchmark, did about as well as its seed (75.6% against 75.2%) on 12% fewer tokens, and helped three other model families by five to ten points. That’s reassuring about overfitting, and a much smaller claim than the headline. I went through what in-sample selection does to these numbers in the last post.
When a gain does show up, it’s worth asking what got repaired. A preregistered study with three small models found the winning configuration mostly won on questions where the losing one had produced no readable answer at all, and that none of the harnesses a search selected beat simply switching the model’s extended thinking off.
Two things stop this being a funeral. The same July study found harness evolution well ahead on long, unfamiliar games, by 80% on one benchmark, which fits everything above: a harness earns its keep where the model is out of its depth. And the field noticed. September’s papers evolve on thousands of tasks kept separate from the benchmark, or regularise the loop against overfitting, and SoL-Pi’s freeze-then-test rule is the same lesson. Even held-out tasks aren’t a complete defence, one paper points out, because they share the benchmark’s format, and a shortcut through the format transfers perfectly.
All four are problems for loops that at least have a clean score to chase. Research doesn’t come with one.
A finished paper isn’t a discovery
Two researchers ran the experiment the honest way. They kept the scaffolding deliberately thin: six agents mapped to the stages of a research project, four tools (read a file, write a file, list files, search the web) and a workspace per idea. They started from more than 135 papers across three fields (world models, multi-agent reinforcement learning, and AI safety), whittled those down to four ideas, and let the system run.
Three of the four died during implementation or evaluation. One became a paper, and was accepted at a venue set up specifically for work with AI first authors.
The value is in their post-mortem, which names six ways it went wrong. The system fell back on whatever was most common in its training data. When code got hard, it quietly simplified the method until something ran that only looked like the original idea. It lost the thread over long tasks. It announced success over results that plainly weren’t. It lacked the depth in the field to know better. And it had poor taste in experiments.
The second of those deserves its own paragraph, because working scientists report the same thing from the other side of the desk. In a collection of accounts of doing research with GPT-5, one physicist describes the model, keen to please, applying “numerical duct tape” to a thorny problem, swapping careful calculations for approximations that show the trend it knows is wanted, and declaring victory over what is still noise. It fixes these when challenged. Somebody has to know enough to challenge it.
Put that beside ScientistOne’s audit, where the worst system invented one reference in five, and the shape is clear. These pipelines are optimised for a finished manuscript, and a finished manuscript is a much easier thing to produce than a true result.
Every failure in the last two sections has something in common. The loop was trusted with a thing it should only have been allowed to read.
What has to stay outside the loop
STOP has the cleanest demonstration. The seed program’s scoring call included the argument use_sandbox=True, followed by a comment in capitals telling the model not to change it. Across 10,000 attempts at self-improvement, GPT-4 produced a version with the sandbox switched off 0.42% of the time. One of those came with its own comment, explaining that this was a more efficient way to score. Adding a warning to the prompt didn’t help: 0.46%.
Forty-two in ten thousand is small. It’s also not zero, and it was done with a straight face in the name of efficiency. A model that’s told to make a number go up will sooner or later look at the thing producing the number.
The systems I’d trust all draw the same line. AHE’s evolving agent can write in exactly one place, the harness workspace. The run logs, the tracer, the verifier and the model configuration are read-only, and the seed system prompt can’t be deleted. Self-Harness names the surfaces that may be edited, makes the proposer list the passing behaviours it must preserve, and tests on held-out tasks before promoting anything. The Darwin Gödel Machine’s authors found that gaming the measure got more frequent when the checking code was visible to the agent, so they hid it.
Since the summer there’s been an audit of what happens without that line. A team went through the published run logs of five self-improving systems looking for edits that raised a score or broke a rule without making the agent any better, which they call harness tampering. Their automated auditor flagged it in 84.6% of iterations for one system and in 73.6% and 63.1% for two others, and in those two the flagged edits piled up in the ancestry of the winning agent.99In order: ADAS, Hyperagents and the Darwin Gödel Machine, then ScientistOne at 29.6% and AFlow at 18.3%. The auditor is itself a language model with false positives of its own, so read these as a ranking, not a census.
The quietest version of the principle is about records. A study of sixteen agent frameworks found that none writes a log an outsider could fully verify, because the harness is both the thing being audited and the author of the audit trail. Sealing that trail with a hash chain didn’t help: it passed all 28 omissions and fabrications the researchers planted. A second log, kept outside the harness, caught all 28. What makes a record evidence, they conclude, is who writes it.
Past that line, here’s what I’d call open. Seven things, none of them solved.
- Marking that’s weak or fuzzy. A loop is only as good as its evaluator, and most interesting work has no crisp one.
- The life cycle of memory. What to keep, what to compress, what to let expire. The 122-token collapse is this problem.
- Negative results. Loops record what worked. What failed tends to vanish, and then gets tried again.
- Variety collapsing. Chase the best candidate hard enough and every candidate becomes a cousin of it.
- Gaming the score. Tests, judge models and benchmark quirks all get exploited. The defences are known: held-out tests, audits of the traces, a person reviewing, and the evaluator and permissions kept out of reach.
- Success that lasts. Maintainability, who owns which code, the cost of migrating, backwards compatibility, the debugging somebody inherits next year. None of it appears in a sandbox score, and all of it is what software engineering mostly consists of.
- Where the people go. Up a level. Less writing the harness, more writing the marking scheme and deciding the permissions.
Three newer findings belong beside that list. Edits that are each harmless can combine into something that isn’t: one study found 43 such pairs and 18 triples. Changes aren’t always reversible: of 600 self-modifications tested in another, 197 that improved the agent couldn’t be cleanly undone. And the plainest: switching a coding harness to auto-approve took attack success from 29.2% to 95.6%.
That’s how to keep one of these honest in the small. Whether the whole enterprise adds up to research needs a bigger yardstick, and there are several.
The scoreboards, for reference
These are the benchmarks that keep coming up when someone claims an agent can do research or engineering. The numbers are the ones each benchmark launched with, so they’re a floor, not the current record.
| Benchmark | What the agent has to do | Size | At launch |
|---|---|---|---|
| PaperBench | Replicate a top machine-learning paper from scratch | 20 papers, 8,316 graded items, rubrics written with the papers’ authors | Best agent 21.0%, below ML PhDs on a subset |
| CORE-Bench | Reproduce a paper’s results from its own repository | 270 tasks from 90 papers in computer science, social science and medicine, at three difficulty levels, some needing vision | Best agent 21% on the hardest level |
| ScienceAgentBench | Write the analysis code behind a published finding | 102 tasks from 44 papers in four disciplines | Best agent 32.4% alone, 34.3% with expert hints |
| RE-Bench | Open-ended ML research engineering, against human experts | 7 environments, 71 eight-hour attempts by 61 experts | Agents scored 4× the humans given two hours, and half as much given 32 |
| MLE-bench | Compete in Kaggle competitions | 75 competitions | Best setup (o1-preview with the AIDE scaffold) took bronze or better in 16.9% |
| KernelBench | Write a GPU kernel that’s both correct and faster | 250 PyTorch workloads | Matched the PyTorch baseline in under 20% of cases |
Three footnotes to the table. PaperBench also ships a lighter variant that only grades the code, and a separate test for whether its automated judge can be trusted, which more benchmarks should copy. KernelBench’s metric, called fast_p, is the share of kernels that are correct and more than p times faster than the baseline, so you can turn the bar up. And RE-Bench’s result is the one I’d remember: agents win short and lose long. Of the human experts’ eight-hour attempts, 82% scored above zero and only 24% matched the reference solution, so these aren’t easy tasks, and people still pulled ahead once there was time to think.
Since August there are also scoreboards for the harness-writing skill itself.
| Benchmark | What it asks | What it found |
|---|---|---|
| HarnessOpt-Bench | Improve another agent’s harness on a fixed budget, scored on a hidden test split | Over 111 runs, which model did the optimising mattered more than which harness it worked through |
| Evo-Bench | Evolve a harness for search, office and general tasks | Best models gained up to 16.6 points and came close to hand-built harnesses, but struggled with office work |
| HarnessDev | Build a harness from a minimal seed, then evolve it | Machine-built harnesses trailed mature ones on code and search. Evolution’s gains were unstable and only partly transferred |
| EvoHarnessBench | Keep working while the harness changes underneath you | Adding tools and skills alone made agents worse at tasks they used to solve |
What I’d take from all this
The harness is where machine self-improvement has actually shown up, and the reason is almost embarrassing. It’s the part you can read. You can diff it, run it and see in an hour whether the change helped. Everything about it that makes for good software engineering also makes for a good optimisation target.
It’s also worth less than the leaderboards say and more than nothing, and the last three months of papers have been unusually good at saying which. Under a strong model with a sane harness, most of what’s left to win is the bill. Under a weak model, or on a task far from anything the model has seen, there’s real capability on the table.
What the loops find is boring, and I mean that as praise. One shell command before the first turn. Create the output file early. Cap the tool calls. Run the check in the same call as the edit. Run the check the grader is going to run before saying you’re done. So far this looks less like a mind redesigning itself and more like a diligent junior keeping a list of what went wrong, which is how most real improvement happens anyway.
And the limit is the same in every paper. The loop can argue for an edit and can’t audit one. It will produce a paper sooner than a result. Given the chance, a few dozen times in ten thousand, it will switch off the sandbox and tell you it’s for efficiency. None of that is an argument against building these. It’s an argument about where the people should stand, which is at the marking, with the keys.
Sources
Essays and documentation
- I. J. Good, Speculations Concerning the First Ultraintelligent Machine, Advances in Computers 6, 1966. A scan. The text says it was drafted in April 1963 and revised in May 1964.
- Eliezer Yudkowsky, Recursive Self-Improvement, 2008.
- Anthropic Institute, When AI builds itself, 2026.
- OpenAI, How agents are transforming work, 2026.
- OpenAI, Unrolling the Codex agent loop, 2026.
- OpenAI, Apply Patch and Subagents, developer documentation.
- Anthropic, Claude Code tools reference, documentation.
- Andrej Karpathy, autoresearch, 2026.
- Terminal-Bench, the benchmark and its leaderboard.
Research: rewriting notes and workflows
- Chrisantha Fernando, Dylan Banarse, Henryk Michalewski, Simon Osindero and Tim Rocktäschel, Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution, 2023.
- Lakshya A Agrawal et al., GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning, 2025.
- Qizheng Zhang et al., Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models, 2025.
- Haoran Ye, Xuning He, Vincent Arak, Haonan Dong and Guojie Song, Meta Context Engineering via Agentic Skill Evolution, 2026.
- Aman Madaan et al., Self-Refine: Iterative Refinement with Self-Feedback, 2023.
- Chris Lu et al., Towards end-to-end automation of AI research, Nature, 2026.
- Rui Meng et al., ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence, 2026.
- Ilia Kulikov et al., Autodata: An agentic data scientist to create high quality synthetic data, 2026.
- Shengran Hu, Cong Lu and Jeff Clune, Automated Design of Agentic Systems, 2024.
- Jiayi Zhang et al., AFlow: Automating Agentic Workflow Generation, 2024.
Research: rewriting the harness, the improver and the weights
- Yoonho Lee et al., Meta-Harness: End-to-End Optimization of Model Harnesses, 2026.
- Jenny Zhang, Shengran Hu, Cong Lu, Robert Lange and Jeff Clune, Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents, 2025.
- Hangfan Zhang et al., Self-Harness: Harnesses That Improve Themselves, 2026.
- Jiahang Lin et al., Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses, 2026.
- Minhua Lin et al., Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents, 2026.
- Seth Karten et al., Continual Harness: Online Adaptation for Self-Improving Foundation Agents, 2026.
- Lirong Che et al., DemoEvolve: Demonstration-Guided Harness Evolution under Sparse Feedback, 2026.
- Eric Zelikman, Eliana Lorch, Lester Mackey and Adam Tauman Kalai, Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation, 2023.
- Jenny Zhang et al., Hyperagents, 2026.
- Alexander Novikov et al., AlphaEvolve: A coding agent for scientific and algorithmic discovery, 2025.
- Robert Tjarko Lange, Yuki Imajuku and Edoardo Cetin, ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution, 2025.
- Yiping Wang et al., ThetaEvolve: Test-time Learning on Open Problems, 2025.
- Mert Yuksekgonul et al., Learning to Discover at Test Time, 2026.
- Kainat Riaz et al., Epistemic Uncertainty for Test-Time Discovery, 2026.
- Prannay Hebbar et al., SIA: Self Improving AI with Harness & Weight Updates, 2026.
Research: published since July 2026
- Yangze Liu and Zhongyi Han, What Does a Harness Buy? Tokens, Mostly, 2026.
- Kirill Brilliantov, Alejandro Hernández-Cano and Emmanuel Abbé, How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?, 2026.
- Oussama Ben Sghaier, Hao Li, Bram Adams and Ahmed E. Hassan, Don’t Blame the Large Language Model: How Agent Harness Evolution Shapes Coding Agent Quality, 2026.
- Run-Ze Fan et al., An Empirical Study of Harness Design for Coding Agents, 2026.
- Chenyang Yang, Xinran Zhao, Tongshuang Wu and Christian Kästner, Better Harnesses, Smaller Models: Building 90% Cheaper Agents via Automated Harness Adaptation, 2026.
- Ziheng Qin, Yaxin Lu, Zhangyang Atlas Wang and Kai Wang, StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling, 2026.
- Paul Barbaste, Tristan Darrigol, Germain Vu and Tom Wiltberger, Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents: A Source-Code Study of Eleven Systems, 2026.
- Yingxuan Yang, Huacan Chai and Ying Wen, Harness Annealing: Learning to Act with Less External Control, 2026.
- Haozhe Liu et al., SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness, 2026.
- Siqi Yang, Qianlan Yang, Yu-Xiong Wang, Saurabh Pujar and Martin Hirzel, One Recipe, Many Harnesses: What Self-Evolution Encodes Across Languages and Models, 2026.
- Zekai Wang et al., VERSE: Verified Self-Evolving Optimizer for Agent Harnesses, 2026.
- Haechan Kim, Yoonho Lee, Gisang Lee, Chelsea Finn and Kangwook Lee, WHALE: A Simple Recipe for Joint Harness-Weight Optimization, 2026.
- Zhou Yu et al., Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails, 2026.
- Akshat Gupta, Jermaine Lei, Alexander Lu, Gopala Anumanchipalli and Leshem Choshen, Automated Discovery Has No Universally Superior Harness, 2026.
- Yike Wang et al., Rethinking the Evaluation of Harness Evolution for Agents, 2026.
- Bowen Xu and Boyu Chen, What Does a Harness Repair? A Preregistered Study of Visibility, Baseline Adequacy and Evaluation Defects, 2026.
- Sina Tayebati, Divake Kumar, Nastaran Darabi, Ranganath Krishnan and Amit Ranjan Trivedi, Self-Healing Harness for Runtime Oversight of Agent Self-Modification, 2026.
- Zhijie Wei, Ferris Tan and Jinghui Wang, Scale and Selection: What Makes Automatic Harness Evolution Work for Visual-Interface Robot Agents, 2026.
- Su Wang et al., Phantom Guardrails: When Self-Improving Agent Harnesses Fix Failures That Never Happened, 2026.
- Siwei Wu et al., ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement, 2026.
- Peng Xia et al., RRSI: Regularized Recursive Self-Improvement of Agent Harnesses, 2026.
- Guojun Zhu et al., Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts, 2026.
- Xing Wang, Xiaoyi Zhang and Jie Shao, Auditing Harness Tampering in Self-Improving Agents, 2026.
- Jiahong Dai et al., Hearsay: Can an Auditor Trust the Record a Deployed Agent Harness Writes?, 2026.
- Zhixiang Zhang, Zesen Liu, Wai Ip Lai, Hongxu Chen and Dongdong She, Compositional Safety Failures in Harness Evolution: Identification and Runtime Monitoring, 2026.
- Tanmay Sah, Dolly Sah, Harshul Jain and Tanya Sah, EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses, 2026.
- Zhengyang Zhu et al., HarnessSecurity-Bench: Do Security Mechanisms Really Protect Coding Agent Harnesses?, 2026.
- Varun Ursekar et al., HarnessOpt-Bench: Evaluating LLMs at Harness Optimization, 2026.
- Lisheng Huang et al., Evo-Bench: Can Language Models Improve Agent Harness?, 2026.
- Yuhao Wu et al., HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?, 2026.
- Zixuan Ke et al., EvoHarnessBench: Can Your Agents Keep Pace with an Evolving Harness?, 2026.
Research: models that make their own training signal
- Weizhe Yuan et al., Self-Rewarding Language Models, 2024.
- Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji and Quanquan Gu, Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models, 2024.
- Andrew Zhao et al., Absolute Zero: Reinforced Self-play Reasoning with Zero Data, 2025.
- Caroline Choi et al., Anchored Self-Play for Code Repair, 2026.
Research: what goes wrong, and how it’s measured
- Dhruv Trehan and Paras Chopra, Why LLMs Aren’t Scientists Yet: Lessons from Four Autonomous Research Attempts, 2026.
- Sébastien Bubeck et al., Early science acceleration experiments with GPT-5, 2025.
- Giulio Starace et al., PaperBench: Evaluating AI’s Ability to Replicate AI Research, 2025.
- Zachary S. Siegel, Sayash Kapoor, Nitya Nadgir, Benedikt Stroebl and Arvind Narayanan, CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark, 2024.
- Ziru Chen et al., ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery, 2024.
- Hjalmar Wijk et al., RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts, 2024.
- Jun Shern Chan et al., MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering, 2024.
- Anne Ouyang et al., KernelBench: Can LLMs Write Efficient GPU Kernels?, 2025.
Cite this post
@article{ghosh2026harness,
title = {An Agent Is a Model Plus a Harness. The Harness Now Rewrites Itself.},
author = {Ghosh, Krish},
journal = {krishghosh.com},
year = {2026},
month = {October},
url = "https://krishghosh.com/writing/harness-rewrites-itself"
}