Generated by Codex with GPT 5.6 Sol XHigh

The Pragmatic Engineer surfaced the Asana Engineering Team’s August 7 post, “We migrated off Enzyme in 2 weeks. It should have taken five years.” The eye-catching result is real: Asana removed its remaining use of an obsolete frontend testing library in about a week and a half of engineering time. The more useful story is why this particular job was unusually well suited to coding agents.

A five-year tail collapses

Asana began moving its frontend tests from Enzyme to React Testing Library in 2022. Enzyme had lost community support, did not fit newer React releases, and encouraged tests that knew too much about component internals. React Testing Library instead pushes engineers to test behavior visible to users.

The company had already devoted several multi-engineer-year efforts to the migration, while product teams converted the tests they owned. That work reduced the backlog but did not finish it. At the existing pace, Asana estimated that the remaining long tail would take roughly five more years.

Engineers then set a deliberately extreme target: eliminate Enzyme in a week. They ran as many as four Codex agents at once, assigned different directories to each, and let them work during the day and overnight. Humans reviewed progress and opened pull requests in the morning and evening. The migration finished in roughly one and a half engineering weeks spread over two calendar weeks, and the team also removed remnants of an even older test framework.

The reported direct cost was about \$11,000 for model use and \$1,000 for infrastructure. Asana contrasts that \$12,000 total with a rough \$5–7 million estimate for three senior engineers working over the projected five-year period. That comparison illustrates the order of magnitude, but it is not a clean experiment: years of earlier human work had already established the target framework, converted many examples, built helpers, and encoded the conventions the agents copied. The post also discloses that it is part of an Asana–OpenAI collaboration, so its economics should be read as a useful company case study rather than a neutral benchmark.

The prompt was not the system

Asana’s instruction to the agents was short. It identified the old and new testing styles, pointed each agent at a directory, asked it to follow local norms, required tests, and suggested starting with easy files. More elaborate attempts—detailed convention documents, tracked subtasks, running notes, and deeper agent hierarchies—generally performed worse.

That does not mean context was unimportant. It means the important context already lived in the repository. The codebase contained working React Testing Library examples, established helpers, and consistent patterns. The task also had an unusually crisp finish line: no Enzyme imports remained. Type checking, linting, tests, and continuous integration could quickly reject bad conversions. Directories supplied natural parallel work units.

In other words, the five-sentence prompt sat on top of a much larger engineering system. The agents were effective because they could infer the desired transformation from good examples and receive fast, mechanical feedback. The case is less a lesson in prompt minimalism than in making a repository legible and testable.

The failures reinforce that point. Some recent internal documentation still recommended Enzyme, so agents treated stale guidance as current policy. A flaky lint step could take more than ten minutes, and differences between local checks and CI required the most human intervention. Once agents can produce changes quickly, contradictory documentation and slow feedback become production bottlenecks rather than minor developer annoyances.

Where the result generalizes—and where it does not

Framework migrations often have the right shape for agentic work: thousands of similar edits, a known destination, abundant examples, automated verification, and limited need to invent product behavior. Other maintenance backlogs—dependency upgrades, API replacements, test modernization, lint cleanup, and repetitive compatibility work—may share those properties.

The case does not show that an agent can compress any five-year project into two weeks. Ambiguous redesigns, migrations with weak tests, changes whose correctness depends on tacit business knowledge, and transformations with hard-to-reverse data effects provide much less reliable feedback. Parallelism also helps only when work units are sufficiently independent and reviewers can absorb the resulting changes.

The practical question for an engineering organization is therefore not simply whether to “use AI.” It is whether a neglected project can be expressed as a bounded transformation, divided safely, checked automatically, and reviewed at the rate agents generate work. If it cannot, the highest-leverage investment may be better tests, cleaner conventions, faster tooling, and current documentation.

The quiet shift in engineering economics

The most consequential result is not the volume of code produced. Asana completed maintenance work that had remained rationally deprioritized because its cost competed with new product development. Coding agents changed that tradeoff enough to make an old liability worth eliminating.

That suggests a better measure of AI’s engineering value than lines of code or pull-request counts: valuable work pulled above the planning line. The teams that benefit most may not be those with the cleverest prompts. They may be those that turn engineering judgment into reusable examples, make correctness observable, and reserve human attention for exceptions, integration, and accountability. Asana’s migration is striking because the agents moved quickly, but it succeeded because years of human engineering had made the correct path easy to see.