When AI Improves Itself, It Just Learns the Test
When I wanted to go to university in the US, I had to take the SAT. Everyone preps for it, and you quickly find out a little secret: you can get much better at the test without getting much smarter. You do old tests over and over. You learn the kinds of questions they like to ask, the tricks for guessing, how fast to go. Your score goes up. But you’re not really better at reading, writing or solving new problems. You’re just better at the SAT.
AI has started doing the same thing.
First, what is a harness?
An AI model on its own can only do one thing: read some text and write some text back. To get real work out of it, someone has to write a small program around it. That program decides what to tell the model, which tools it is allowed to use, what it should remember, and when it is done. That program is the harness.
The key idea is that the harness is the code around the model that manages the interaction loop:
notes = []
while True:
# what we tell the model
prompt = INSTRUCTIONS + task + notes
# the model reads and writes back
reply = model(prompt)
# the model asks to use a tool, e.g. "open this file", "run the tests"
if reply.wants_tool:
result = run_tool(reply.tool)
# remember what happened
notes.append(result)
else:
# finished
return reply.answer
So the model is not the harness. The model is the part that thinks. The harness calls the model over and over, gives it tools, remembers what happened, and sets what it is allowed to do.
The model in the middle stays the same. But change the harness, with better instructions, a new tool, or smarter notes, and the same model can do a much better job.
Then people had an idea: let the AI rewrite its own harness.
The AI makes a small change to its own harness, like tweaking its instructions or adding a tool. Then it takes a test. If the score goes up, it keeps the change. If not, it throws it away. Then it tries again, and again. The AI model itself is no more capable than before. The AI is just rewriting the program around it.
This works, at least on paper: the AI’s score on the benchmark, a fixed set of test tasks, goes up. The catch is that it keeps practicing on the same tasks. Do that long enough and it starts learning the tasks instead of the skill. Its harness fills up with little tricks that only help on those exact tasks: special cases, hints, shortcuts. Give it new tasks it has never seen, and the gains shrink or disappear. It’s the SAT problem all over again. The AI gets very good at the test, not at the job. Maybe that’s fine if all you want is a specialist agent for one small, narrow job, where the test tasks and the real tasks are basically the same.
A new paper from Google Cloud AI Research and collaborators, called RRSI (Regularized Recursive Self-Improvement), tries to stop the AI from just learning the test. One obvious approach would be to limit what the AI is allowed to put in its harness. RRSI does something different: it leaves the harness alone and puts rules on the improvement loop itself, the cycle of changing the harness, testing it and keeping or throwing away the change.
In plain terms, RRSI is built so that tricks have a hard time surviving. Picture a tutor going through a student’s study notes and crossing out every line that is just a memorized answer to a past exam question, keeping only the notes that explain how to actually solve problems. RRSI has a checker that does this to the harness: any change that only makes sense for specific test tasks gets thrown out. The AI may only make a few changes at a time, so a useless change can’t sneak in alongside a good one. A higher score has to be clearly bigger than what luck alone could produce, because a lucky run is another way to “improve” without getting better. Changes that make the AI more expensive to run have to pay for themselves, and parts that stop helping get removed, so the harness stays lean instead of piling up special cases.
The real proof is what happens on tasks the AI never practiced on. The researchers let RRSI improve the harness on one set of benchmarks, then tested it on other benchmarks it had never seen. The harness did better on those unseen tasks too, not just on the ones it practiced on. That’s the difference between learning the skill and learning the test. It’s like a student whose SAT prep also makes them better in their actual classes. On top of that, the improved harness was cheaper to run, handling about 30% less text going in and out of the model than the same loop without RRSI’s rules. A harness stuffed with special-case tricks tends to grow, not shrink, so that’s another good sign.
Conclusions?
Self-improving harnesses will probably become more common, and that isn’t a bad thing. Specializing is fine. A student who studies only medicine isn’t cheating, and good notes and a solid method help any student. In the same way, a harness with good procedures, safety checks and reference material can make an AI better at new problems, not just old ones.
The trouble starts when the notes turn into an answer key. I’ve seen the same thing at work with machine learning models: they look great on the data they were built with, then struggle once they’re out in the real world.
So picture two students who get the same score on an open-book exam. One brought a single page of notes. The other brought a binder full of answers to past questions. Change the questions, and you’ll quickly find out who understands the subject. That’s why RRSI’s cost rule matters. If a harness needs more and more machinery to hit the same score, that’s a hint the machinery, not the AI, is doing a lot of the work. The goal isn’t the smallest harness. It’s the smallest one that still leaves the AI doing the thinking.
It’s scary to think where we will be in 10 years.