Generated by Codex with GPT 5.6 Sol XHigh
Techmeme surfaced SemiAnalysis’s September 14 deep dive, “A Brain Too Big to Carry — On-Device vs Datacenter Inference,” about a decision that may define the first useful fleets of general-purpose robots: how much intelligence should travel on the machine, and how much should run on shared datacenter hardware.
The answer is unlikely to be all-or-nothing. Fast motion, balance, collision avoidance, and emergency behavior must remain local because they cannot wait for a wireless round trip. High-level planning can sometimes move to a datacenter, where larger models and pooled accelerators offer more capability at lower cost per unit of compute. The unresolved problem is the link between those layers. Ordinary WiFi is designed for people downloading data while mostly stationary, not metal machines continuously uploading video and depending on predictable response times.
The robot sets a hard real-time boundary
Robotics reverses the usual relationship between models and hardware. A cloud language model can grow until serving it becomes an infrastructure problem. A robot manufacturer must buy the compute for every machine, fit it inside a limited power and cooling envelope, and still meet physical deadlines. If a chatbot pauses, the conversation waits. If a robot pauses, the object, person, or robot may have moved before its action arrives.
That difference creates a natural hierarchy. Servo and safety loops run at hundreds of hertz; a 100 Hz controller has only 10 milliseconds to produce its next output. A typical wireless round trip of 10 to 50 milliseconds can consume that entire budget before inference begins. These loops therefore belong on the robot. A planning model running at 5 Hz has roughly 200 milliseconds per decision and can tolerate a network hop, so long as the delay is consistent.
Consistency is the catch. Fixed latency can be predicted or buffered. Jitter—the occasional late packet or one-second spike—makes the next update unknowable. Buffering for the worst case makes every action slow, while ignoring the tail can leave a robot frozen mid-task. The article’s most useful distinction is therefore not edge versus cloud, but which decisions can survive an irregular network and which must remain deterministic when the link fails.
Model scale pulls in the opposite direction. Today’s generalist robot models are mostly measured in billions of parameters, but more open-ended physical reasoning may require models too large for an onboard accelerator. SemiAnalysis says Jetson Thor provides about one-tenth the compute and one-thirtieth the memory bandwidth of a GB200. NVIDIA’s 14-billion-parameter DreamZero reportedly needs two off-robot GB200s for real-time inference, while the later 3-billion-parameter RoboTTT shows how model design can swing the boundary back toward the machine. The placement decision will keep moving as both models and embedded silicon improve.
Pooling compute can win—if the fleet stays busy
Shared datacenter inference has an economic advantage that a chip bolted to one robot cannot match: utilization. Home robots may work only a few hours a day, leaving expensive local silicon idle. A server can batch requests from many machines and remain busy even when individual robots stop. SemiAnalysis estimates that shared GPU inference begins using less leading-edge silicon at roughly seven robots per accelerator and less DRAM at roughly five, because one server’s resources are pooled across the fleet.
Its cost model points the same way, but the assumptions matter. The authors reconstructed the compute and memory profile of NVIDIA’s RoboTTT because public code and weights were unavailable; the benchmark measures serving cost, not whether the robot completes tasks correctly. In a modeled 96-robot deployment, one B300 can sustain 12 robots within a 500-millisecond chunk deadline. After applying a 90% utilization assumption to the shared server and 40% to onboard Jetson Thors, the B300 case reaches about 46% of the Thor case’s cost per unit of dense compute. At five or more robots per B300, offloading becomes cheaper by that measure.
Those results do not prove that cloud-hosted robots are cheaper products. The calculation excludes mechanical components common to both cases, depends on assumed useful lives and utilization, and does not price every network retrofit, outage, privacy obligation, or safety consequence. It also compares compute economics rather than task-level output. Still, it exposes the central tradeoff: local compute buys independence, while pooled compute buys access to larger models and spreads scarce silicon across more robots.
Supply incentives reinforce the datacenter side. Jetson Thor now uses the same leading-edge TSMC process family that high-margin datacenter accelerators compete for, while its 128 GB of LPDDR5X draws from a memory market increasingly shaped by HBM demand. SemiAnalysis argues that current robot volumes are too small to strain wafer capacity directly. The strategic question is whether future fleets should reserve advanced silicon and memory inside every robot or share them in racks.
Deployments are already splitting along product boundaries
The companies examined in the article do not converge on one architecture because they sell different kinds of reliability. Boston Dynamics splits Atlas into an onboard “System 1” for vision-to-action control and an off-robot “System 2” for planning and supervision. The planner runs through the company’s Orbit platform on Google infrastructure, interpreting work orders and translating them into simpler instructions the local controller can execute. That gives Atlas access to a much larger reasoning model without asking the network to maintain balance or generate every motor command.
Agility Robotics keeps Digit’s learned control stack onboard. Factory radio conditions are hostile: moving metal, machinery, multipath interference, changing inventory, and access-point handoffs all create unpredictable gaps. More importantly, safety cannot depend on connectivity. A biped that loses an offboard safety model would have to stop blindly, so collision detection, controlled deceleration, and human-proximity behavior stay local even when slower fleet orchestration and updates use the cloud.
Home deployments make the same point less dramatically. Sunday Robotics initially expected to use cloud inference, then moved its model onboard after real homes exposed dead zones, neighbor interference, poor mesh setups, and unreliable alternatives such as 5G and satellite connections. Weave Robotics also expects local execution, although training-data uploads and teleoperation still depend on connectivity. “On-device” removes the network from the critical control path; it does not make the product network-independent.
The network must become part of the robot
SemiAnalysis calls the remaining obstacle the “network wall.” Robot traffic inverts the assumptions of consumer networking: cameras continuously send large uplink streams, clients move, the radio environment changes, and a rare dropped packet matters more than average throughput. Existing access points generally offer best-effort contention when robots need scheduled, predictable service.
The proposed remedy is closer to co-designing a site than buying a faster router. Robot-aware access points would reserve uplink slots, steer beams based on a machine’s location and path, coordinate fast handoffs, use multiple WiFi and 5G links, and share a precise clock across the fleet. The robots would compress perception data, use radios designed for uplink capacity, and capture observations on the same cadence. Datacenter schedulers could then batch synchronized frames without waiting unpredictably for stragglers.
That design may work in factories and other bounded sites where the radio environment can be surveyed and changed. It is less plausible on public roads and harder in homes where every layout and network is different. It also expands the product from a robot into a robot-plus-network-plus-compute system, with corresponding installation, security, privacy, and operational burdens.
The lasting insight is that a general-purpose robot’s “brain” is not one model in one place. It is a layered system whose placement follows deadlines, failure modes, model size, fleet utilization, and control over the environment. As physical AI scales, the winning architecture may be the one that makes the boundary boring: reflexes and safety remain local, large-model reasoning is pooled only where the network can be engineered, and every loss of connectivity has a predictable outcome.