Generated by Codex with GPT 5.6 Sol XHigh
Large language model serving has an awkward failure mode: a software process can die while its GPUs, node, and model weights are still perfectly usable, yet the replacement process must behave as if it were starting from nothing. Reloading weights into high-bandwidth memory, compiling kernels, rebuilding communicators, and capturing CUDA graphs can take minutes. During that interval, the remaining replicas inherit the traffic and often miss latency or throughput targets.
The official NVIDIA Technical Blog published this post on August 25, 2026. It describes Shadow Engine Recovery, a preview feature in NVIDIA Dynamo that moves most restart work out of the failure path. A fully initialized standby process waits on the same GPUs as the active engine, shares the active model’s physical weight memory, and takes over when a kernel-managed lock is released. In NVIDIA’s controlled test, this reduced the return of a second serving worker from 283 seconds to 7.3 seconds.
The useful idea is broader than a faster restart button. The design separates durable machine state from disposable process state, precomputes what cannot be transferred, and uses operating-system primitives to make failover depend less on the health of the failed process.
Preserve the expensive state outside the engine
A conventional inference engine owns its GPU allocations through its CUDA context. When the process exits, the driver destroys that context and releases the allocations, including model weights that may have taken minutes to load. A new process cannot simply attach to the old engine’s memory. Other startup products are equally process-specific: NCCL and torch.distributed communicators belong to the process that created them, while captured CUDA graphs assume the virtual addresses used during capture.
Shadow Engine Recovery addresses these two categories differently. It moves the reusable physical weight memory under a separate service, then creates the non-transferable process state in advance inside a standby engine.
The GPU Memory Service, or GMS, is a per-GPU sidecar that owns physical GPU pages on behalf of inference engines. It uses the CUDA Virtual Memory Management API to decouple physical allocations from the virtual addresses in any one CUDA context. GMS hands engines importable handles; each engine maps the same pages at addresses in its own context. Because the physical allocations are reference-counted, the weight bytes remain resident when one engine disappears as long as another mapping survives.
This is sharing rather than copying. The active and standby engines see the same physical tensor data, so the second process adds no separate model-weight footprint in HBM. Once mapping is complete, GMS is not on the inference data path. A GPU kernel follows an ordinary pointer to the same weight bytes it would have read from an engine-owned allocation. NVIDIA integrated this mechanism into vLLM, SGLang, and TensorRT-LLM through a custom torch.cuda.CUDAPluggableAllocator, leaving weights visible to the engine as normal tensors.
That narrow interface is an important implementation choice. Persisting state helps only if adopting the persistence layer does not require an inference framework to rewrite how every kernel accesses memory. GMS changes allocation and ownership while preserving the engine’s normal execution model.
A standby precomputes what cannot survive a crash
The shadow is a complete engine process, not a container image waiting to start. Before parking, it imports the shared weight mappings, establishes NCCL and NIXL communicators, creates its CUDA context, compiles and warms the engine, and captures CUDA graphs. Those artifacts cannot be inherited safely from a failed peer, so the design pays their cost before a failure occurs.
The standby does not keep every serving allocation materialized. In particular, the KV cache is the largest reclaimable block and today belongs only to the active engine. The shadow reserves the cache’s virtual-address range without backing it with physical pages, then materializes the cache only when promoted. Its steady-state cost is therefore limited to the CUDA context, captured graphs, communication buffers, communicators, and mappings to the one shared copy of the weights.
Each Kubernetes worker pod combines two engine containers, the GMS sidecar, and a shared file lock. At steady state, the active engine holds the lock, owns a materialized KV cache, and is registered with the frontend router. The shadow is fully initialized but dormant, has no KV cache, and blocks on the lock.
Failover follows a short four-phase state machine. If the active engine crashes, or if a liveness probe kills a hung process, the kernel closes its file descriptors and releases its POSIX flock. The shadow acquires the lock, remaps the GMS-backed weights, materializes its KV cache, and registers with the router. Kubernetes restarts the failed container in the background; after initialization, that new process parks as the next shadow. The two engines have swapped roles without putting a full cold start on the serving path.
Using flock is more than a convenience. Mutual exclusion prevents both engines from serving at once, while kernel-mediated release avoids asking a corrupted or deadlocked process to announce its own failure. A process that is alive but stuck eventually reaches the same path when the liveness probe sends SIGKILL. This turns promotion into a small leader election backed by lifecycle semantics the operating system already guarantees.
Fault injection measures the customer-visible difference
NVIDIA compared the design with a cold restart in a two-worker GLM-5.2 deployment. Each worker occupied one NVIDIA B200 node with tensor parallelism of eight, a 200,000-token maximum context, NVFP4 model weights, and an FP8 KV cache. The synthetic requests contained 32,000 input tokens and requested 1,000 output tokens, arriving at 0.7 requests per second. After the system reached steady state, the test sent SIGKILL to one worker and observed the fleet for 600 seconds.
The cold-started worker took 283 seconds to serve again. The shadow configuration took 7.3 seconds: 1.7 seconds to detect the fault and 5.6 seconds to promote the standby. That distinction mattered because a two-worker fleet loses half its capacity after one failure. In the baseline, median time to first token after the fault reached 23,815 milliseconds, versus 1,311 milliseconds with the shadow. Median decode rate was 12 tokens per second per user in the baseline and 46 with the shadow.
The request counts make the tail impact easier to see. During the post-fault window, 201 of 399 cold-restart requests waited more than five seconds for their first token; the shadow arm recorded one such request out of 398. The baseline also produced 226 requests below 20 tokens per second per user, compared with none in the shadow configuration.
These results isolate process recovery rather than every availability problem. The shadow lives on the same GPUs and node, so it cannot help when the hardware, node, or a multi-node deployment fails. Those cases still need ordinary rescheduling and replica-level redundancy. The experiment is also a two-node synthetic workload, not proof that every production traffic mix will see identical numbers. Its strength is the controlled comparison: identical engine configurations and load, with the standby as the only intentional difference.
Recovery time is a state-placement decision
The preview has practical boundaries. It currently targets vLLM, requires Kubernetes 1.34 or newer with Dynamic Resource Allocation and NVIDIA’s GPU DRA driver, and promotes a shadow with an empty KV cache. That cold cache causes a small time-to-first-token bump after cutover. NVIDIA is working to place KV-cache pages under GMS as well, so a future standby could inherit both cache memory and the prefix-cache index instead of rebuilding them as requests arrive.
Even with those limits, the architecture offers a general reliability pattern. A restart is slow when durable, expensive state is accidentally bound to a disposable process. The remedy is not always another replica consuming a full copy of every resource. It can be a deliberately layered design: externalize state whose lifetime should span processes, pre-initialize state that cannot move, keep the standby’s reclaimable footprint sparse, and let an independent authority detect ownership loss.
For LLM inference, this reframes capacity recovery as memory and lifecycle engineering. Model weights, virtual addresses, communicators, graphs, KV caches, locks, containers, and routers have different lifetimes and failure boundaries. Once those differences are made explicit, a software crash no longer has to trigger a hardware-sized restart. The broader takeaway is that fast failover comes less from accelerating every initialization step than from deciding which steps should never be on the recovery path at all.