Porting a Large Codebase With Coding Agents
Published October 11, 2026 · by the AQ team
Porting a large codebase to another language with AI coding agents works when, and only when, you have an oracle: the original implementation and its test suite, turned into a conformance harness the agents iterate against. The playbook behind 2026's successful ports (Bun's Zig to Rust rewrite, the community Rust port of the TypeScript compiler, Ladybird's LibJS port from C++ to Rust) is four moves: build the harness first, give each parallel agent its own git worktree with gated merges, write the porting rules into one document every session reads, and account for cost in API-priced tokens. Each move below comes with the numbers those projects published.
Why ports work better than greenfield agent projects
A port has something no greenfield project has: a complete, executable definition of "correct." The original program is a reference you can run on any input, which enables differential testing (running both implementations on the same input and diffing the outputs; any difference is a bug in the port by definition). The agent loop becomes "write code, run the harness, fix the named failure": tokens go to measurable gaps, not design debates.
The 2026 ports all leaned on exactly this. Ladybird's Andreas Kling ported LibJS's parser pipeline to Rust (about 25,000 lines in roughly two weeks, written up in February 2026) under a standing requirement of byte-for-byte identical output, starting where the test262 conformance suite is densest. Bun reused its own test suite, written in TypeScript and so indifferent to the runtime's language; the team reports zero tests skipped or deleted. The TypeScript port to Rust (ts-rust, released October 7, 2026) ported the Go compiler's own tests and reports all 181,711 passing.
Step 1: Build the conformance harness first
Stand up the machinery that will judge the port before any porting:
- Pin the source. Port a fixed revision, not a moving branch: ts-rust pins a specific September 29, 2026 TypeScript commit, so upstream changes become a scheduled rebase, not silent drift.
- Make the oracle one command. Agents will run the old-vs-new diff thousands of times, and mechanical checks (does it compile, is the diff empty, what share of the suite passes) cost no tokens; save the model's context for failures.
- Carry the original's tests across. Reuse the suite as is when it is implementation-independent like Bun's; otherwise porting the tests is the first agent task, because every later task is graded by it (executable specs discipline).
- Track one number. "Percent of conformance suite passing" is the progress bar: Bun watched failing test files drop from 972 to 23 in two days after the first full CI run.
The harness is also the token budget: naming the exact failing case keeps each session a short, bounded loop instead of rework (verifying agent work at scale).
Step 2: One worktree per lane, and gate the merges
Parallel agents in one checkout destroy each other. Bun's July 8, 2026 postmortem is blunt: one Claude instance ran git stash, another ran git stash pop, a third ran a hard reset, and work vanished. The fix: four lanes, each in its own git worktree running 16 agent instances (64 at peak), plus a standing rule that agents never run git stash, git reset, or any git command beyond committing specific files (one worktree per agent was out: 64 working copies would have exhausted the disk).
A git worktree (a second working directory sharing one repository, so each lane gets its own files and branch without a second clone) is the unit of isolation that makes this safe; see git worktrees for AI coding agents and why parallel sessions collide. Split lanes along module boundaries, and merge a lane only when the conformance number holds. Bun also gated every merge on review: two reviewer instances in separate context windows, shown only the diff and told to assume it was wrong, caught a use-after-free and a panic-prone unwrap before they landed (the same structure works with humans).
Step 3: Write the porting rules down once
Every session on a weeks-long port must make the same idiom decisions, so the decisions live in a file. Bun spent about three hours discussing how Zig patterns map to Rust (banned concurrency primitives, lifetime handling, keeping Bun's own event loop) and had the agent serialize the discussion into a PORTING.md every later session read. Day to day the same job is done by AGENTS.md and CLAUDE.md files; a port just raises the stakes, since an undocumented convention becomes a thousand inconsistent files. Two of Bun's rules earn a permanent place there: define what "make it compile" means (its agents initially read it as permission to stub out the functions that had errors), and ban the justifying comment (a workaround needing a paragraph of explanation is wrong code to fix, not prose to admire).
Step 4: Budget long runs and real money
These projects are measured in agent-weeks: Bun's rewrite ran May 3 to 14, 2026 with over 6,500 commits, and ts-rust and Ladybird's port each took about two weeks. Sessions that long run on machines that stay up, not laptops: see overnight coding agent runs and running multiple agents in parallel.
The published numbers are clarifying. The ts-rust README reports over $400,000 of API-priced tokens across months of GPT-5.6 Sol and GPT-6 Astra runs that stalled near 84 percent compatibility; the Opus 5.5 restart that shipped cost about $24,047 over two weeks, which the author measured at nine to ten times a $200 subscription plan's weekly limits. Bun's port came to about $165,000 at API pricing, with 72 billion cached input token reads against 5.9 billion uncached: prompt caching doing most of the work. The lessons: measure convergence early and be willing to restart, since a failed approach can cost more than the successful one; subscriptions and API billing are different economics (usage limits, token costs for teams); and model choice dominates the bill, so A/B test models on your own repo first.
The 2026 ports at a glance
| Project | Direction | Scale | Time | Verification |
|---|---|---|---|---|
| Bun (postmortem July 8, 2026) | Zig to Rust | 535,496 lines of Zig to about 780,000 of Rust | May 3 to 14, 2026 | Own TypeScript test suite; zero tests skipped or deleted |
| ts-rust (released October 7, 2026) | Go to Rust (TypeScript 7 compiler) | Compiler, checker, language server | Two weeks, after a $400k+ stalled first attempt | All 181,711 ported Go tests passing; output diffed against tsc on real repos |
| Ladybird LibJS (written up February 23, 2026) | C++ to Rust | About 25,000 lines of Rust (parser pipeline) | About two weeks | Byte-for-byte output match against the C++ pipeline; test262 |
What still goes wrong
The harness proves behavior, not maintainability. The ts-rust author says he has never read a line of the resulting code, and the Hacker News discussion split on exactly that point: someone eventually extends the code. Plan the post-port phase (ownership, an architecture map, a conventions cleanup) as part of the project.
Review concentration is the quieter risk. A July 2026 LeadDev analysis of 25,264 agent-generated pull requests across 2,361 popular GitHub repositories found that in 79 percent of agentic PRs the same developer both reviewed and modified the agent's contribution, and only about one in eight workflows involved multiple humans (LeadDev, July 2026). On a port, where thousands of commits land in days, one person absorbing all of it is a bottleneck; give each lane a human owner. And known regressions ship (Bun lists 19 traced to the port, since fixed): the harness catches what the old tests covered, production telemetry catches the rest, so stage the rollout.
Where AQ fits
AQ is the multiplayer coding harness where engineering teams run AI coding agents like Claude Code and Codex together: shared live terminals, a code editor, and app previews, in your own cloud. A port is the workload AQ's shape was built for. Agents run as real CLIs (Claude Code, Codex, Cursor Agent, Kimi, Grok, or plain shells) in persistent tmux sessions on your team's VM, so a two-week lane keeps running when laptops close and resumes from any device, and each workspace gets its own isolated git worktree on its own branch (ai/{id}-{slug}, dependencies installed automatically, one-click rebase onto main): the lane isolation above, without the shell scripting, in parallel.
The team layer maps to the review problem: teammates open the same workspace and watch the same live session, driving someone else's terminal only after its owner approves a control request, so a stuck lane can be handed to whoever knows that module. Agents commit, push, and open PRs under per-user GitHub auth, tracked per workspace, and each person logs into the CLIs with their own Claude or OpenAI account (AQ never marks up usage on your own subscriptions), so the cost accounting above stays as legible as on a bare VM. The Free plan is a personal sandbox (AQ creates a private machine in an isolated network, nothing to install, no time limit), enough to pilot one lane; the Team plan is $50 per user per month early access (standard $200, billed monthly), covering VMs you connect from your own cloud or a dedicated always-on AQ-managed VM, rate locked for your first 12 months.
Frequently asked questions
Can an LLM really port an entire large codebase to another language?
Yes, with qualifications that matter. As of October 2026 the public examples include Bun (535,496 lines of Zig to about 780,000 lines of Rust, May 2026), a community Rust port of the TypeScript 7 compiler with all 181,711 ported tests passing (October 2026), and Ladybird's LibJS parser pipeline (February 2026). All three had a strong conformance suite and a runnable reference implementation. Without that oracle, repository-scale translation still fails routinely: the first ts-rust attempt spent over $400,000 in API-priced tokens and stalled near 84 percent compatibility.
How do you verify that an AI-ported codebase is correct?
Differential testing against the original: run both implementations on the same inputs and treat any output difference as a bug in the port. Concretely, pin the source revision, carry the original's test suite across (or reuse it directly if it is implementation-independent, like Bun's TypeScript test suite), and track percent-passing as the project's single progress number. Ladybird went further and required byte-for-byte identical ASTs and bytecode from the old and new pipelines.
How much does it cost to port a codebase with coding agents?
The two public 2026 data points: Bun's port consumed about 5.9 billion uncached input tokens and 690 million output tokens (plus 72 billion cached reads), roughly $165,000 at API pricing, for a half-million-line codebase in 11 days. The TypeScript to Rust port cost about $24,000 of API-priced Claude Opus 5.5 usage over two weeks, after a $400,000+ failed first attempt on other models. Budget for the possibility of a restart, and remember subscription plans meter differently than API billing: the ts-rust author measured his two weeks at nine to ten times a $200 plan's weekly limits.
Should parallel agents share one checkout or use separate git worktrees?
Separate worktrees, one per lane of work, every time. Bun's team documented what happens in a shared checkout: one agent ran git stash, another ran git stash pop, a third ran a hard reset, and work was destroyed. Their working setup was four worktrees with 16 agent instances each, plus a standing rule that agents never run git stash, git reset, or any git command beyond committing specific files. Split lanes along module boundaries and merge only when the conformance suite holds.
Is AI-ported code maintainable if nobody has read it?
That is the honest open question. A conformance suite proves the port behaves like the original today; it says nothing about how hard the code is to extend tomorrow, and the ts-rust author states he has never read a line of the result. If the codebase has a future, treat post-port ownership as part of the project: assign human owners per module, have agents generate an architecture map, and schedule a conventions cleanup pass. Bun's port also shipped 19 known regressions (since fixed), so stage the rollout behind real telemetry.