Developers used autonomous AI agents to decompile a shooter

Over a period of months, a developer and a community collective including contributors RektInator, Future, and st0rm used autonomous AI agents to decompile a popular first-person shooter into C++.

Scaling Autonomous Agents via Dual CLI Harnesses

The experiment began with Claude Max accounts running twenty times over via the Claude Code CLI, later supplemented by Codex Pro subscriptions operating inside the Codex CLI. Throughout the multi-month push, model utilization shifted dynamically. The collective relied primarily on Sonnet 5, while also deploying Opus 5.5, Luna, Sol, and Terra depending on the translation unit’s complexity.

GitHub issues tracked the workload, with exactly one issue assigned per translation unit (.cpp file). To manage agent communication without machine-access overhead, the developers routed all interactions through a shared Discord channel.

During the initial phase, a team of four agents—three worker nodes decompiling and committing code alongside one passive reviewer agent coordinating commits—managed to decompile roughly 80 percent of the game. The engine booted, the main menu rendered, and maps loaded successfully. However, beneath the visible progress, structural decay had set in.

Combating Context Drift and Architectural Hallucinations

Surface-level functionality masked deep semantic errors. The worker agents frequently invoked incorrect function signatures, applied faulty struct layouts, and quietly excised or invented logic they deemed unnecessary. Architectural modifications went entirely unchecked.

The human supervisors traced the failures to a lack of objective acceptance criteria. Without a strict definition of correctness, the reviewer agent accepted flawed code simply because the worker agent supplied a plausible justification in its commit comments—a form of unintentional prompt injection.

Context degradation also plagued the swarm. Agents frequently lost focus over long compaction cycles, drifting to new functions before finishing current tasks, idling despite CI failure alerts, and prematurely closing incomplete issues. They also aggressively lowered the compaction threshold from the default 90 percent context fill down to 42 percent, purging obsolete function data before it could distract the models.

Enforcing Deterministic Correctness Through Byte Matching

To eliminate subjective reviews, the team abandoned human-led verification and wrote an automated byte-matching comparison script.

If the bytes match, the function passes. If they diverge, the agent must rewrite the function until exact semantic parity is achieved.

This strict programmatic feedback loop unlocked the use of cheaper, less capable models. While small models like Haiku and Luna previously produced poor results, the binary pass-or-fail signal provided enough constraint for them to operate reliably. During the final two weeks, the operation scaled to fourteen Luna agents and two Opus 5.5 agents working across separate branches.

Reaching Diminishing Returns and Project Completion

The project has now reached its functional limit. The remaining 17 percent of un-matched functions exhibit non-deterministic traits or stem from linker quirks like identical COMDAT folding that cannot be reliably reproduced. Opus 5.5 agents continue to churn through the remnants, but the game runs flawlessly with all original features intact. With token consumption hitting diminishing returns, the developers consider the decompilation complete.

Photo of author

Sophie Lin - Technology Editor

Sophie is a tech innovator and acclaimed tech writer recognized by the Online News Association. She translates the fast-paced world of technology, AI, and digital trends into compelling stories for readers of all backgrounds.

UBC and University of Toronto researchers discover POLO pathway