BT

Facilitating the Spread of Knowledge and Innovation in Professional Software Development

Write for InfoQ

Topics

Choose your language

InfoQ Homepage News Agents Refactor 300K Lines in Three Weeks, and Practitioners Ask What It Proves

Agents Refactor 300K Lines in Three Weeks, and Practitioners Ask What It Proves

Listen to this article -  0:00

CodeScene has published a case study in which coding agents refactored a 300,000-line C codebase over three weeks, at a token cost of roughly $4,000. The work produced 2,903 commits across 726 files, modified 252,055 lines, and moved the codebase from a Code Health score of 5.6 to 10.0. The codebase is Street Fighter III: 3rd Strike, taken from an open-source decompilation.

Adam Tornhill, CodeScene's founder and the author of Your Code as a Crime Scene, wrote that this was the first time he had seen what he called superhuman AI performance at scale, after three decades working on large systems.

Two mechanisms carried the work. The first was a quality signal: the CodeHealth MCP Server, which gave agents a deterministic score to optimize and to judge whether a transformation had helped. The second was correctness: a replay-trace harness that compared the rollback state hash frame by frame, so behavior could be checked after every change.

The more novel result is what the agents built along the way. Rather than applying a fixed catalogue, they accumulated a refactoring playbook, ending with 22 recipes and 82 supporting notes. Familiar transformations appear, including Extract Function and Guard Clauses, but so do recipes specific to this codebase. Shared Index Range captures repeated loops differing only in start and end ranges. Action Parameter handles duplicated control structures differing mainly in which function they invoke. Uniform Step Table converts heterogeneous calls into table-driven dispatch. Failed attempts were recorded too, including transformations that made Code Health worse.

Model choice mattered. The team settled on Claude Opus for the bulk of the work, reporting that Claude Code with Opus was significantly better than Codex with Sol at capturing and documenting the emerging patterns. Files often plateaued when smaller models ran the task, appearing to reach a local optimum they could not move past.

AI-commits-example_Streetfighter_codebase_casestudy

Non-merge commits per day by model across the three-week project (Source: CodeScene)

Reaction from practitioners on LinkedIn has been sharply divided, and the split runs along what the result proves rather than whether it happened.

Paolo Perrone put the case for taking it seriously:

most refactor claims i've read rest on a green test suite, which only tells you the tests survived. replaying traces against a fighting game sets a much higher bar.

Mats Iremark, CTO at Omda Response, reported doing something comparable, writing that the CodeScene MCP combined with current agents is almost like cheating.

The skeptics concentrated on scope. Konrad Otrębski, a tech lead and consultant, asked whether the work was merged, whether it arrived as one enormous merge or many, and whether this was an experiment on open-source code rather than production code earning money. Daniel Webb, CTO at NeoSee and one of the two engineers who did the work, replied that it was merged to main on a fork, through 54 pull requests.

Otrębski's follow-up set a higher bar:

I think the real true test of AI capability here would be to offer this courtesy of refactoring to some famous open source project, say Grafana. The definition of done is ofc merging it to master.

Tracy Bannon, a software architect and researcher, objected to the framing rather than the method, noting that describing the outcome as perfect is pretty bold.

Denis Baltor questioned the discovered recipes themselves, arguing that DRY concerns duplication of knowledge and intent rather than identical lines of code, and that recipes defined by loops differing only in ranges, or control structures differing only in the function invoked, may be collapsing the two.

Asko Nõmm raised the methodological question. Since Claude Code and Codex are tuned to their own models, he argued it is unclear how much of the result measures the model and how much measures the harness, and that cost would vary the same way. He added that architecture remains unmeasured, so code can look healthy while fundamental problems surface later.

Several unanswered questions came from the authors themselves. Marc Bouvier asked whether non-functional behavior had improved, given that framerate, memory use and input latency matter in a game; Webb said a performance specialist was being brought in. Asked whether the harness caught subtle frame timing regressions, he said there might be no definitive answer, because it ran as a pre-commit hook and some failures were fixed without being observed. On diff sizes, he was equally direct: if you are not familiar with the code and no longer review every line, how big can a diff be? He made no assertion, only questions.

If these reactions are representative, the disagreement is not about the numbers but about the conditions that produced them. The replay-trace harness worked because a decompiled game offers deterministic frame-by-frame replay. Most legacy systems have no equivalent oracle, which is the same reason refactoring them is risky in the first place. Tornhill writes that automated tests and equivalence checks are absolutely essential safeguards, which places the burden on exactly the thing unhealthy codebases tend to lack.

The team's own selection process illustrates it. Webb said they considered a Gov.UK marine licensing codebase he had worked on, but it was too healthy to be useful for the research that follows, and chose the game partly because they play it and are therefore its users.

That research is the next step. The uplift produced two functionally equivalent versions of the same system, one at Code Health 5.6 and one at 10.0, for a study with Lund University in which students will implement features in both using frontier models and compare cost and quality.

Two headline figures belong to that future work rather than this case study. CodeScene projects a roughly 70% reduction in AI-induced defects and roughly 45% less token waste from the uplift, both extrapolated from its earlier research rather than measured here. The $4,000 and the three weeks were measured.

About the Author

Rate this Article

Adoption
Style

BT