Skip to content
Apixo
Blog
news· 2 min read· via The Decoder

Google Researchers Address AI Agent Memorization in Self-Improvement

New Google research reveals how AI agents can avoid memorizing test tasks during self-improvement using a regularized approach called RRSI.

Google Researchers Address AI Agent Memorization in Self-Improvement

Recent advancements in autonomous artificial intelligence agents stem largely from improvements made to their surrounding frameworks, known as harnesses, rather than updates to the underlying models themselves. Traditionally, developers patched these harnesses by hand after analyzing failed runs. Newer automated approaches use language models to repeatedly rewrite the harness based on test feedback, creating a form of recursive self-improvement.

However, a recent research paper from Google highlights a major drawback to this automated optimization. Because agents continually work on a restricted set of training tasks, they tend to memorize the benchmarks. While their training scores improve, performance on unseen tasks remains stagnant or declines. The research notes that systems can inadvertently memorize patterns specific to single benchmarks, favor chance-based candidates, or accumulate unnecessary complexity.

The RRSI Solution

To counter this tendency toward memorization, the researchers developed Regularized Recursive Self-Improvement of Agent Harnesses (RRSI). This method regulates both ends of the optimization loop while keeping the harness fully editable. RRSI imposes a shrinking edit budget that limits the number of bundled changes a candidate can make, forcing smaller, traceable adjustments over time. Additionally, the system tracks prior attempts to avoid repeated failures and explores untouched components when progress stalls.

A strict critic model evaluates every proposal, discarding any changes that hardcode solutions or rely on benchmark-specific tricks. Furthermore, rules dictate that increased compute costs must yield measurable performance gains, and unhelpful components are systematically removed.

During testing across eight benchmarks covering coding, office tasks, and engineering design with a frozen Claude Opus 4.8 model, RRSI demonstrated notable improvements. The method achieved up to 14.1 points on training tasks and up to 4.7 points on five unseen benchmarks, notably improving by 4.7 points on JobBench. Furthermore, RRSI used about 30 percent fewer tokens at runtime compared to the unregularized version, avoiding the performance drops typically seen with over-optimized harnesses.

Interestingly, harnesses optimized using one model also benefited others. A coding harness optimized with Gemini 3.5 Flash improved the accuracy of the less capable Gemini 3.1 Flash Lite from 11.2 to 14.6 points. Developers can try top AI models cheaply through one API at https://apixoai.online.

What it means for developers

For developers building and deploying autonomous AI agents, these findings emphasize that optimizing agent performance requires careful guardrails to ensure generalizability. Relying on automated self-improvement without regulation can easily lead to models that excel at specific benchmarks while failing in real-world scenarios. By implementing constraints like shrinking edit budgets and independent critics, developers can build robust agents that maintain performance across diverse, unseen tasks while optimizing token efficiency.


Source: Google researchers find a way to keep self-improving AI agents from memorizing their tests — The Decoder. Written by the Apixo team from that report.

#ai-news#ai-agents#google-research#machine-learning#llm#developers
Try it with your own tools

One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.

Get your API key

Keep reading