RRSI lets an LLM agent evolve its own harness (prompts, tools, memory, control flow) using a regularized search that curbs overfitting to its training tasks.
Yes, RRSI is open source under the Apache-2.0 license.
RRSI is free to use.
Yes, RRSI can be self-hosted (the source is available under the Apache-2.0 license).
RRSI (Regularized Recursive Self-Improvement) is a Google Research framework for evolving an LLM agent's harness — the prompts, control flow, tools, memory, skills, and sub-agents wrapped around a frozen model — while a regularized search curbs the overfitting that plain harness-evolution produces. Open source: yes (Apache-2.0); self-hostable; free OSS. It has ~1.1k stars, ships runnable code rather than just paper text, and is backed by a paper (arXiv 2609.24972).
An agent's capability is largely set by its harness, and you can improve the harness automatically by searching over edits against a fixed task set — but that search tends to memorize the training tasks, so in-distribution gains shrink or vanish out of distribution. RRSI keeps the edit space fully open (prompts, control flow, configuration, context management, tools, skills, memory, sub-agents can all change) and instead regularizes how the search moves: an annealed budget caps edits per candidate, the proposer sees the full edit history so a falsified hypothesis is not redrawn, a critic screens each candidate for suite-specific logic before evaluation, a noise-adjusted floor blocks gains inside evaluation variance, and components that stop helping are pruned. Every candidate harness is drafted and evaluated in its own git worktree, and the edit history records the component, hypothesis, measured score/cost, and verdict per edit, so the trajectory is auditable. Open source: yes (Apache-2.0); self-hostable; free OSS.
The cheapest capability upgrade for an agent is often a better harness, not a bigger model — and the hard part is improving it without quietly tuning to your eval set. RRSI is a credible, permissively-licensed take on exactly that: automated harness improvement with explicit anti-overfitting machinery and an auditable trail of what changed and why. For a GTM-infrastructure team building agents on frozen models, it is a method worth studying and testing against your own task suites — the worktree-per-candidate design and per-edit evidence log are reusable ideas even if you never adopt the whole loop. The honest read: this is research code, not a product — early, with reference domains you adapt — and the headline claim that regularization preserves out-of-distribution gains is the paper's finding, so validate it on your tasks before building on it.