Generated by Codex with GPT 5.6 Sol XHigh
Techmeme surfaced SemiAnalysis’s August 25 deep dive “OpenAI Jalapeño: Better Than Nvidia Blackwell”. The headline result is striking: OpenAI’s first custom inference chip reportedly serves large language models with both higher throughput per watt and lower latency than the Nvidia systems tested. The more consequential story, however, is how a frontier AI lab combined custom silicon, rack-scale networking, serving software, and AI-written low-level code quickly enough to make a first-generation chip competitive.
Jalapeño is an inference accelerator. It is designed to run trained models and generate answers, not to perform the enormous training runs that create frontier models. That distinction matters because inference is the recurring cost behind every ChatGPT response, API call, and multi-step agent task. If OpenAI can produce more useful tokens from the same power envelope, it can serve more demand without waiting for an equivalent expansion in grid capacity.
What the tests actually show
OpenAI tested Jalapeño on InferenceX, SemiAnalysis’s public benchmark, using three open-weight models: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. According to OpenAI’s published results, Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. On selected highly interactive operating points, the advantage rose to between 2.1 and 4.1 times.
Those are system-serving measurements, not a simple count of arithmetic operations. Large-model inference has two different phases. Processing the prompt, called prefill, leans heavily on computation. Generating the answer token by token, called decode, is often constrained by how quickly model weights and cached context can move through memory. Communication between chips adds another bottleneck. A system can look excellent at one phase and stall at another.
Jalapeño attacks the whole path. Its 700-watt package is paired with high-bandwidth memory and a network designed as part of the accelerator rather than an attachment. A rack contains 128 chips, and the scale-up network can connect as many as 2,048. OpenAI says the architecture keeps model state local when possible and can shift resources between prefill and decode as the workload changes. SemiAnalysis reports 15.4 terabytes per second of memory bandwidth per chip, an especially important figure for low-latency generation.
The result is a better position on the throughput-latency curve. Operators normally increase total throughput by batching more requests together, but that makes individual users wait longer. Jalapeño reportedly sustains more output per unit of power without accepting the same latency penalty. That is particularly valuable for agents, where pauses accumulate across many sequential model calls.
The benchmark has real limits
“Better than Nvidia Blackwell” is not the same as universally better than Nvidia. SemiAnalysis says its staff observed the InferenceX runs in OpenAI’s lab, but OpenAI supplied the numbers. The publication did not execute its complete benchmark suite and has not seen AgentX results, which stress long-context, multi-turn workloads and cache management more like production agents do. The public tests use three open-weight models with nominal single-turn prompts, not OpenAI’s proprietary frontier models.
The comparison also captures young hardware and fast-changing software stacks at particular operating points. OpenAI normalized results using published package power ratings; actual facility-level cost also depends on host CPUs, networking, cooling, utilization, yields, and the price of the complete rack. OpenAI has not published an independent production cost per token. SemiAnalysis estimates Jalapeño is roughly level with Nvidia’s newer Rubin platform on current total-cost performance, but it expects Jalapeño to improve once OpenAI adds techniques such as speculative decoding.
The claims are therefore best read as credible early engineering results, not a settled market verdict. They show that the chip works, that an outside benchmark team observed it, and that a first design can compete at relevant workloads. They do not yet show fleet reliability, manufacturing economics, broad model compatibility, or performance under sustained production traffic.
AI is compressing the software moat
The unusual part of the development process is that OpenAI used AI on both sides of the chip. Earlier models helped with design exploration and verification; newer ones helped write and tune the kernels that map model operations onto the hardware. OpenAI says the core design moved from its initial phase to tapeout in nine months. SemiAnalysis measures about 16 months from early hiring to tapeout, a different starting point but still an unusually short cycle for a new ASIC team.
The chip deliberately exposes low-level control. Its kernels can resemble assembly code and run to thousands of lines, which traditionally demands scarce specialist labor. OpenAI’s Gluon programming language supplies explicit abstractions for data layout and hardware resources, while Codex searches for efficient implementations. The team says Codex with GPT-Astra brought the three benchmark models—none of which had been in the original production plan—to high performance in two months. On selected GPT-OSS attention and mixture-of-experts blocks, AI-generated implementations ran 1.5 to 1.8 times faster than existing human-written versions. That improvement applies to particular blocks, not entire models.
This development loop challenges one part of Nvidia’s advantage. CUDA’s value is not only hardware performance; it is the mature software, libraries, tools, and developer knowledge accumulated around Nvidia GPUs. A clean-sheet chip normally struggles because that ecosystem must be rebuilt. A frontier lab that can use its own models to generate, test, and retune kernels may compress the software bring-up period enough to make custom hardware practical.
That does not erase Nvidia’s broader moat. Nvidia sells general-purpose systems to many customers, supports far more workloads, and has an established manufacturing and deployment machine. OpenAI can optimize around its own traffic, accept a narrower market, and consume the hardware internally. Jalapeño shows that the largest AI labs may be able to reduce dependence on merchant accelerators at the margin; it does not show that they can replace them.
The strategic payoff is power, not just price
The chip’s most important economic unit may be tokens per megawatt rather than tokens per dollar. Money can buy more servers, but grid interconnections, substations, cooling systems, and backup generation take years to expand. When a data center reaches its power ceiling, a more efficient accelerator creates capacity that additional capital alone cannot immediately provide.
OpenAI also captures value that would otherwise become a chip supplier’s margin and gains control over the schedule for the models it expects to serve. It plans to begin deploying Jalapeño internally by the end of 2026, while SemiAnalysis expects production to ramp gradually through 2027, with most volume late in the year. OpenAI says second- and third-generation designs are already under way, but it also expects to keep deploying Nvidia and other partners’ accelerators.
The durable takeaway is not that one benchmark has crowned a new chip champion. It is that custom inference silicon has become a plausible full-stack strategy for frontier labs. Jalapeño combines a specialized business need, power-constrained infrastructure, early access to future model workloads, and AI-assisted programming. If the production system preserves the laboratory gains, the competitive boundary in AI infrastructure will move upward: from who can design the fastest chip to who can co-design the model, compiler, network, rack, and data center as one continuously optimized machine.