Generated by Codex with GPT 5.6 Sol XHigh

Large software rewrites usually fail for organizational reasons before they fail for technical ones. They demand years of feature-freeze patience, create a branch that drifts from production, and defer the hardest question—whether behavior was preserved—until an enormous cutover. GitHub’s rewrite of the Copilot agent runtime is notable because it changed those economics without pretending that code generation made correctness automatic.

The official GitHub Blog published the post on September 16, 2026. It describes how GitHub moved a live, rapidly changing runtime from TypeScript and Node.js to roughly 832,000 lines of production Rust in about fourteen and a half weeks. AI agents wrote most of the code, but the durable lesson is about migration design: keep the system shippable, make behavioral equivalence the governing constraint, isolate parallel work, and reserve human attention for boundaries and judgment.

The target was a different deployment shape

The old runtime began inside the TypeScript Copilot CLI. As other products adopted it, the architecture inverted: SDKs in C#, Python, Go, Java, Rust, and TypeScript launched a headless CLI process and communicated with it over JSON-RPC. Every client therefore inherited Node.js and V8, roughly 100 MB of minimum working-set overhead, an extra process to supervise, and serialization across a process boundary for messages, events, tools, and session storage.

GitHub wanted a runtime that could be embedded directly in those hosts, start quickly, consume predictable resources, and expose a stable cross-language boundary. Rust fit those constraints because it could compile to a native shared library, expose a C ABI, and retain an out-of-process mode where isolation was still desirable. The choice was workload-specific: the post explicitly avoids claiming that every large TypeScript system should be rewritten in Rust.

The permanent interface is deliberately small. A shared library exposes just 19 C ABI functions for lifecycle, connection, session, and host operations. The existing bidirectional JSON-RPC contract—364 dispatch routes at the time of the post—travels through those functions as bytes. Reusing JSON-RPC inside the process may seem inelegant, but it let all six SDKs keep their request correlation, callbacks, and public APIs while replacing a pipe with a function call. A typed C function for every method would instead have created six binding layers and a large binary compatibility surface.

Replace components without stopping the world

Rather than build a second runtime on a long-lived branch, the team ported the system in place. Each pull request replaced one TypeScript component with a thin shim into Rust, deleted the old implementation, and ran the existing CLI and SDK end-to-end tests against the new code. The main branch stayed shippable while other developers continued adding features. Across the migration window, it produced 135 releases—100 prereleases and 35 stable releases—so each deployed increment contained a small, knowable set of newly ported components.

This strategy needed a temporary two-way bridge. Rust functions were exposed to TypeScript through napi-rs, while Rust called unported TypeScript through thread-safe callbacks onto Node’s main thread. The seam grew before it disappeared, peaking at 2,019 internal N-API exports and 3,356 TypeScript call sites. Once the final components moved, both counts fell to zero. That shape is important: transitional complexity was accepted, measured, and designed to be deleted rather than becoming a permanent hybrid architecture.

Maintaining parallel TypeScript and Rust versions of each subsystem would have looked safer, but the authors rejected it because stateful orchestration cannot be shadowed cleanly. A session component owns mutable state, callbacks, persistence, tools, and model interactions. Keeping two live copies synchronized while hundreds of unrelated pull requests land every week can introduce more ambiguity than an atomic replacement. Confidence came instead from small diffs, unchanged behavioral tests, prerelease exposure, and rapid fixes.

Agents supplied labor; humans operated the control loop

The rewrite was divided into 128 pull requests. Agents handled translation, repository exploration, testing, rebasing, and review, often through separate worktrees and branches. The hardest component, a roughly 30,000-line session.ts, spent nearly an hour and 122 tool calls on inspection before creating code. Its parent session then split the work across 15 child sessions in seven waves and used smaller subagents for focused exploration and review. Isolated worktrees let most children mutate distinct files without colliding; the parent absorbed the coordination cost at shared hubs.

This was not a one-prompt conversion. Of about 2,600 human-authored messages, the largest categories involved review, testing, CI, challenging design choices, and pushing incomplete work to completion. The developer’s role moved up a level: define the end state, choose component boundaries, adjudicate ambiguous behavior, protect quality gates, and decide when the evidence was sufficient to merge.

Long-running agent work also depended on harness engineering. A stable prompt prefix produced a 96.22% cache-read rate, making repeated context far cheaper than fresh input. Automatic context compaction occurred 5,116 times, while subagents kept exploratory work out of the coordinating session’s context. The entire effort consumed about 136.3 billion tokens and roughly \$120,000 in attributed model spend, plus an estimated three weeks of one developer’s time and contributions from the wider team. Agents did not make the project free; they lowered its cost enough that a rewrite previously estimated at a team-year or two became economically defensible.

Correctness remained the expensive part

Rust’s compiler caught thousands of ordinary translation errors: unresolved names, missing methods, type mismatches, and unsatisfied trait bounds. Its safety model also made boundary risk visible. The finished runtime contained 158 unsafe blocks, all associated with the C ABI, operating-system APIs, SQLite, dynamic loading, or process-global environment state—not the model, MCP, agent, or prompt layers.

Compilation could not prove behavioral equivalence. The known regressions clustered around implicit TypeScript semantics, changed ownership and lifetimes, incomplete migration, host boundaries, and bad test oracles. A JavaScript number could become the wrong Rust numeric type; || and unwrap_or could disagree about an empty string; Node could supply a time zone that Rust now needed explicitly. These are contract mismatches, not syntax errors, and they survived until end-to-end behavior exposed them.

GitHub turned this into a review discipline. Specialized review instructions asked multiple agents to compare old and new implementations line by line, reject nontrivial TypeScript left behind, preserve tests, and repeat the review/fix cycle until independent reviews were clean. The most important rule was to protect the oracle: an agent changing the implementation could not silently weaken a test, update a snapshot, or raise a compatibility baseline to make its own work pass.

The payoff came from removing architectural overhead

With model and network latency replaced by a deterministic local server, the new in-process runtime showed the effect of the client-side architecture. A complete client creation, one-turn session, and teardown fell to roughly 55 milliseconds in the later measurement cited by the post. A shared client sustained 120 one-turn session lifecycles per second, versus 7.55 for the old TypeScript process tree. In the ten-client memory test, private resident memory above baseline fell from 1,383 MB to 126 MB, a 91% reduction.

Those figures are end-to-end measurements, not a controlled language benchmark; many changes landed during the same period. Still, they demonstrate the design goal. Eliminating repeated Node/V8 startup, reducing process boundaries, and sharing a native runtime let hosts support far more concurrent sessions before CPU, memory, or process count became limiting.

The broader lesson is that agents make ambitious migrations feasible only when the surrounding engineering system is unusually explicit. State the finish line so partial ports are not mistaken for completion. Translate before redesigning so behavior and architecture do not change simultaneously. Build end-to-end tests before starting, keep the oracle independent, convert recurring mistakes into standing instructions or tooling, and optimize build and test loops for many concurrent worktrees. Code generation changed the labor curve; incremental delivery, protected contracts, and human judgment made the result production-worthy.