US Techs Register

America's Tech, Logged

Breaking News
Fund Rounds

AI agents refactor massive codebase in three weeks

By Imogen Fairfax September 30, 2026
Handwritten document with agent name and ID code, showcasing personal handwritten text.
Handwritten document with agent name and ID code, showcasing personal handwritten text. Photo: cottonbro studio/Pexels

Coding agents refactored a 300,000-line C codebase in just three weeks, according to a new case study from CodeScene. The work transformed Street Fighter III: 3rd Strike, an open-source decompilation, producing 2,903 commits across 726 files. The agents modified 252,055 lines of code, boosting the codebase’s Code Health score from 5.6 to 10.0 at a token cost of roughly $4,000.

Adam Tornhill, CodeScene’s founder and author of Your Code as a Crime Scene, called the result the first instance of “superhuman AI performance at scale” he had witnessed in his three decades of working on large systems. The project relied on two key mechanisms: a quality signal from the CodeHealth MCP Server, which gave agents a deterministic score to optimize, and a replay-trace harness that compared rollback state hashes frame by frame to ensure correctness after every change.

Agents Built a Custom Refactoring Playbook

The agents did not follow a fixed catalog of transformations. Instead, they developed a custom refactoring playbook over time, ending with 22 recipes and 82 supporting notes. While familiar patterns like Extract Function and Guard Clauses appeared, the agents also created codebase-specific recipes. One, called Shared Index Range, captured repeated loops differing only in start and end ranges. Another, Action Parameter, handled duplicated control structures that mainly differed in which function they invoked. Failed attempts were recorded too, including transformations that worsened Code Health.

The choice of model had a clear impact. The team used Claude Opus for most of the work, finding that Claude Code with Opus outperformed Codex with Sol at capturing and documenting emerging patterns. Smaller models often caused files to plateau at a local optimum they could not surpass, the report noted.

Divergent Reactions

Reaction from practitioners has been sharply divided, with disagreement focused on what the result proves rather than whether it occurred. Paolo Perrone argued that replaying traces against a fighting game sets a much higher bar than the green test suites typical of most refactor claims. Mats Iremark, CTO at Omda Response, called the CodeScene MCP combined with current agents “almost like cheating.” Others raised pointed questions about the project’s scope and methodology.

Konrad Otrębski, a tech lead and consultant, asked whether the work was merged and how it was delivered. Daniel Webb, the CTO at NeoSee and one of the engineers who did the work, confirmed it was merged to main on a fork through 54 pull requests. Otrębski then proposed a higher-stakes test: applying the same courtesy refactoring to a famous open-source project like Grafana and merging it to master.

Some critics took issue with the framing. Tracy Bannon, a software architect, objected to describing the outcome as “perfect.” Denis Baltor questioned the discovered recipes, arguing that concerns about duplication should focus on knowledge and intent, not just identical lines of code. Asko Nõmm raised a methodological point: since Claude Code and Codex are tuned to their own models, it’s unclear how much of the result measures the model versus the harness.

Read Also: Cloudflare Releases Python Workers to Public

The authors themselves acknowledged unanswered questions. Marc Bouvier asked if non-functional behavior, like framerate and memory use, had improved. Webb said a performance specialist was being brought in. On the harness’s ability to catch subtle frame timing regressions, he was less certain, noting some failures were fixed without being observed. He also questioned how large a diff can be without line-by-line review, offering no firm assertion, only questions.

Risks and Safeguards

Ultimately, the debate hinges on the conditions that made the work possible. The replay-trace harness succeeded because a decompiled game offers deterministic, frame-by-frame replay. Most legacy systems lack such an oracle, which is why refactoring them is inherently risky. Tornhill emphasizes that automated tests and equivalence checks are essential safeguards, placing the burden on safeguards that unhealthy codebases often lack.

That research is the next step. The uplift produced two functionally equivalent versions of the same system, one at Code Health 5.6 and one at 10.0, for a study with Lund University in which students will implement features in both using frontier models and compare cost and quality.

Future Research

The study with Lund University aims to measure the practical impact of the codebase transformation. Students will use both versions of the Street Fighter III code to implement new features. The goal is to see if working with the refactored code, now at a Code Health score of 10.0, is faster and produces better results than working with the original version, which scored 5.6.

CodeScene has already projected some potential benefits from this kind of improvement. Based on its earlier research, the company forecasts a roughly 70% reduction in AI-induced defects when building on healthier code.

Unanswered Questions and Methodological Concerns

Several methodological questions remain unanswered. The cost of such a project would likely vary depending on which combination is used.

Commit Rates and Unanswered Questions

Webb suggested there might be no definitive answer, as the harness ran as a pre-commit hook and some failures were fixed without being observed. On diff sizes, he posed a direct question: if one is no longer reviewing every line, how large can a diff be?

Leave a Reply

Your email address will not be published. Required fields are marked *

© 2026 US Techs Register. All rights reserved.