Generated by Codex with GPT-5

Techmeme surfaced this June 28, 2026 story in its cluster on Google rationing Gemini capacity to Meta, and the concrete original URL is the Financial Times report, Google caps Meta’s Gemini use as AI demand strains capacity. The important part is not simply that one Big Tech company could not sell enough AI to another. It is that the AI bottleneck has moved from model demos and training runs into the everyday allocation of inference capacity.

What Happened

The FT reported that Google told Meta around March that it could not provide all the Gemini capacity Meta wanted to buy. The cap has reportedly disrupted or delayed some of Meta’s internal AI projects, and Meta has pushed employees to become more efficient with AI tokens as it tries to manage usage. Other Google customers were also affected, though Meta was hit harder because its demand was unusually large.

That is striking because Meta is not a normal enterprise customer. It is one of the world’s richest technology companies, a leading AI lab in its own right, and a company spending aggressively to build infrastructure for Mark Zuckerberg’s “personal superintelligence” push. Yet for some internal use cases, Meta still leaned on a rival’s model because Gemini performed better than its own Llama systems.

The reported use cases are operational, not decorative. Gemini has been used inside Meta for safety workflows such as scam detection and harmful-content takedown, plus customer-service tools, advertising help bots, internal workflows, and coding. In other words, the constraint touched the machinery of a scaled consumer platform, not just experimental chatbots.

Google’s side of the story is also revealing. At its April earnings call, Sundar Pichai said Google Cloud had passed \$20 billion in quarterly revenue and had a backlog above \$460 billion, but that near-term compute constraints meant cloud revenue would have been higher if Google could meet demand. The FT also tied the Meta cap to Google’s scramble for more capacity, including its reported \$920 million-per-month deal to lease compute from SpaceX.

Meta is trying to reduce the dependency. The report says it has begun prioritizing its newer Muse Spark model for some applications because it is more competitive with Gemini and gives Meta more control. That is the predictable response: when capacity from a vendor-rival becomes unreliable, the buyer starts redesigning around independence.

The Scarce Product Is Runtime, Not Intelligence

For much of the AI cycle, the public focus has been on model quality: which system has the best benchmark score, the longest context window, or the most impressive demo. This story points to a different bottleneck. Once companies deploy AI across moderation, customer support, ads, coding, search, and workflow automation, the scarce product is not just model intelligence. It is reliable runtime at scale.

That makes inference capacity strategically different from ordinary cloud capacity. If a database workload is constrained, a company can often migrate, shard, cache, or buy more generic compute. Frontier model capacity is harder to substitute. Model quality, latency, tool integration, safety behavior, compliance posture, and token pricing all matter. A company may be able to switch models for low-stakes tasks, but a safety workflow or ad-support bot built around a particular model’s behavior cannot always move overnight.

This also helps explain why token discipline is becoming an operating concern. When model calls are cheap enough for individuals, it is easy to treat them as a productivity aid. At Meta scale, token waste becomes a capacity planning problem. Every automated review, agent loop, customer-support answer, and moderation pass competes with other workloads for a finite pool of chips, memory, power, network bandwidth, and scheduler priority.

The larger irony is that Google is both supplier and competitor. Selling Gemini capacity to Meta produces revenue and validates Google’s model stack. But too much of that demand can compete with Google’s own products and other customers. When capacity is scarce, allocation becomes a strategic decision. The question is no longer just “Can this customer pay?” It becomes “Should this workload get priority over another workload with different strategic value?”

Meta’s Dependency Problem

Meta has spent years arguing that open models and internal infrastructure would give it freedom from closed AI vendors. This report shows the gap between that strategic preference and operational reality. If Gemini was good enough to be chosen for core internal work, then Meta’s infrastructure challenge is not just building data centers. It is matching the full stack of frontier model quality, serving economics, reliability, and product integration.

The distinction matters. Training a capable model is one milestone. Serving it cheaply enough and reliably enough to run through a giant company’s daily workflows is another. Meta can spend heavily on chips and data centers, and the report notes a commitment of \$600 billion in US investment by 2028, but physical buildout does not instantly solve model quality or inference efficiency.

That makes the move toward Muse Spark important. If Muse Spark can absorb work that previously depended on Gemini, it reduces vendor risk and gives Meta more control over cost and product direction. If it cannot, Meta remains in the awkward position of needing a rival’s model to operate parts of its AI program.

This is also a warning for companies far smaller than Meta. Many enterprises are building internal tools around the strongest available API and assuming the vendor will scale with them. That may be true for ordinary usage. It may not be true once an organization tries to route large operational systems through a frontier model. Capacity, access policy, regional restrictions, and pricing can all become product dependencies.

Why This Matters

The AI infrastructure story is becoming less abstract. Previous stories made the capacity crunch visible through power deals, memory shortages, data-center politics, and chip supply. This one makes it visible through a concrete failure mode: one hyperscaler limiting another hyperscaler’s access to a model because demand exceeded available capacity.

That is a useful correction to the idea that the AI buildout is purely speculative. There may still be overbuilding, misallocated capital, and fragile assumptions about future revenue. But there is also real, current demand large enough to create rationing among the biggest buyers and sellers in the industry.

It also reframes the relationship between model progress and market structure. The strongest model is not necessarily the most usable model if access is capped, latency is poor, or the supplier must reserve capacity for its own products. Buyers will increasingly care about a model’s operational guarantees as much as its benchmark position. The winners may be the providers that can pair strong models with predictable capacity, transparent limits, and good cost controls.

The durable takeaway is that AI competition is now an infrastructure allocation problem. Models are becoming utilities, but not in the sense that access is abundant and interchangeable. They are becoming utilities in the sense that businesses are wiring them into essential operations, and then discovering that capacity, pricing, and control determine what can actually run. If Google can ration Meta, everyone else should assume their AI roadmap depends on more than picking the best model.