Generated by Codex with GPT 6 Luna XHigh
The Pragmatic Engineer surfaced its September 24 report, “The Pulse: a new trend of CPU shortages”. After years of focusing on GPUs, the AI build-out is tightening access to ordinary server CPUs too. Agents create this second demand stream because they do more than ask a model for tokens: they run code, call tools, compile, test, and search through software environments. Reinforcement learning workloads also rely on CPUs to execute the tasks models are learning to perform.
The reported change is showing up in cloud procurement. The article gathers accounts from CTOs, infrastructure leaders, and a large inference provider: deep discounts for idle spot capacity are fading, reservations may need months of lead time, and providers sometimes cannot fulfill requests even from large customers. Those are early operator reports, not a market-wide capacity survey. The mechanism is plausible, though: server CPUs share constrained fabrication capacity with accelerators, while memory makers shift production toward high-bandwidth memory for AI systems.
One concrete signal comes from Uber. In its account of running a software factory at scale, the company says weekly agent requests rose 9.4-fold from February to August 2026. More agents running in cloud instances mean demand can rise even when model inference gets cheaper. The capacity question is therefore changing from “How many GPUs do we need?” to “How much general-purpose compute will each agent workflow consume, and when can we get it?”
For engineering teams, autoscaling on cheap spare CPUs may no longer be a safe planning assumption. The Pragmatic Engineer recommends forecasting baseline and growth needs, finding CPU-heavy services that can be trimmed, and securing capacity earlier where workloads cannot tolerate a squeeze. The broader supply story remains uncertain, but the operational signal is clear enough to check CPU exposure before the next growth cycle.