Generated by Codex with GPT 5.6 Sol XHigh

Quantizing a language model is often treated as a final conversion step: train in high precision, compress the finished checkpoint, measure the damage, and accept the best compromise. NVIDIA’s latest work makes a more useful engineering move. It turns quantization error into a condition the model can train against, allowing the deployment format to influence the model before release rather than merely constraining it afterward.

The official NVIDIA Technical Blog explains how the team built the NVFP4 version of Nemotron 3.5 Lightning with quantization-aware distillation, or QAD. The resulting checkpoint is about 22 GB rather than 66 GB, and NVIDIA reports up to four times the throughput of the BF16 model. The important part of the post, however, is not the headline compression ratio. It is the recipe for deciding where precision can be removed, how the resulting error can be recovered, and when the extra training is actually worthwhile.

A full-precision model teaches its quantized counterpart

Ordinary post-training quantization, or PTQ, maps a trained model’s weights and sometimes its activations onto lower-precision values. It is inexpensive because it does not substantially retrain the model, but it freezes in the approximation error introduced by the low-precision grid. Conservative PTQ recipes minimize that error by leaving sensitive operations at higher precision. The tradeoff is that they also leave some memory and throughput gains unrealized.

QAD adds a second stage. The original BF16 checkpoint becomes a frozen teacher, while a PTQ-produced low-precision copy becomes the student. Each training batch passes through both models. The student experiences simulated inference-time quantization during its forward pass, and a Kullback-Leibler divergence loss pushes its output logits toward the teacher’s. The student therefore learns how to behave within the numerical constraints it will face in production.

This differs from simply fine-tuning the quantized model on next-token labels. The teacher supplies a dense target distribution over possible tokens, preserving more of the original model’s behavior than a single hard label conveys. It also keeps the optimization goal closely tied to the exact model being compressed: the student is not trying to discover a new policy so much as reproduce its own full-precision ancestor under noisier arithmetic.

That changes how aggressively the first quantization pass can be configured. A PTQ-only workflow commonly aims to retain more than 99% of the baseline median benchmark score. NVIDIA instead targeted roughly 95% to 99% before QAD. A small but visible regression was evidence that the recipe had actually captured worthwhile size and latency savings; the training stage would then try to close the deliberately opened quality gap.

Precision is allocated by operation, not by slogan

Nemotron 3.5 Lightning combines sparsely activated mixture-of-experts components with Mamba and attention layers, so “make it 4-bit” is not one uniform decision. NVIDIA evaluated five PTQ recipes that varied the calibration method, the precision of Mamba projections, and the KV-cache format. All used W4A16 NVFP4 weights for the mixture-of-experts, shared, and output-head layers, while attention projections stayed in BF16. The most aggressive variant also moved the K and V cache tensors to NVFP4 while leaving the attention matrix multiplications in BF16 and Q unquantized.

The team found that quantizing Mamba linear layers to W4A16 instead of the safer FP8 setting offered the best useful performance tradeoff once QAD was available to repair the loss. Calibration context length mattered too. It tested sequences from 8K to 128K tokens and selected 32K for the chosen four_over_six PTQ recipe. These choices illustrate why low-precision deployment is a model-architecture problem, not merely a file-format conversion: different operations amplify numerical error differently, and the calibration workload has to expose the behavior that matters.

The quantization-scale strategy follows from the calibration method. A checkpoint calibrated with simple maximum values can use dynamic-scale QAD, recomputing its scales as the weights change. Checkpoints calibrated by an MSE search, including the selected four_over_six recipe, use frozen-scale QAD. Repeating that search on every training step would be too expensive, so the PTQ-discovered scales remain fixed while the underlying weights adapt around them.

This is a subtle but broadly applicable design principle. Training should not blindly make every deployment parameter learnable. If a parameter was selected by a costly offline optimization, it may be better treated as part of the environment that the model must accommodate. QAD then spends compute on the weights that can adapt cheaply, not on continuously re-solving calibration.

Long-context quality must be trained at long context

NVIDIA’s ablations found that short QAD sequences left performance on long-context evaluations unrecovered. Initial experiments used 256K-token sequences to reduce iteration cost, but preserving the model’s long-context behavior required roughly the 522K-token length used in post-training; the final run was configured at about 524K.

This matters because quantization noise is not necessarily uniform across workload shapes. A recipe that looks sound on short prompts can fail once recurrent state, attention caches, routing behavior, and small numerical errors accumulate over a much longer execution. The calibration and distillation workloads therefore need to resemble the serving envelope, not merely fit conveniently into an experiment.

The post also makes the workflow reproducible rather than stopping at a result table. NVIDIA Model Optimizer provides the PTQ and QAD pipeline; Megatron-LM runs the distillation loop; and Megatron-Bridge handles conversion between Hugging Face checkpoints and Megatron Core while exposing tensor, expert, context, and data parallelism. The published launchers cover student creation, frozen-teacher distillation, and export to a deployable Hugging Face checkpoint. NVIDIA recommends its open post-training datasets as a starting point, although the exact final internal data mix is not released.

The evaluations show both the value and the boundary

The cleanest QAD result came from an aggressively quantized intermediate supervised-fine-tuning checkpoint. PTQ alone recovered 96.33% of the BF16 model’s median score. Two hundred QAD iterations raised recovery to 99.72% at the same 21.19 GB footprint and improved 10 of 11 reported benchmarks. Large recoveries appeared on AIME 2025, SciCode, and Humanity’s Last Exam, suggesting that reasoning and coding were among the capabilities most exposed by the aggressive quantization.

A later reinforcement-learning checkpoint showed the same general pattern but less uniformly: median recovery rose from 95.84% with PTQ to 98.53% with QAD, while only five of nine individual benchmarks improved. Some measures, including normalized GDPval Elo, regressed. Holding the PTQ and QAD students at the same size is important because it isolates retraining as the source of the change rather than giving the QAD model more memory.

The final production-oriented checkpoint supplies the most valuable caveat. Its PTQ recipe was intentionally conservative and already achieved 99.24% median recovery. QAD’s aggregate median was slightly lower at 98.97%, not higher. It did improve several agentic and coding evaluations—Terminal-Bench v2.1 rose from 22.05 to 25.84, and smaller gains appeared on SWE-Bench Multilingual, BrowseComp, SWE-Bench Verified, PinchBench, and Humanity’s Last Exam—but other reasoning and long-context scores fell.

QAD is therefore not a general promise that another training stage makes every metric better. Its economic value grows when aggressive PTQ creates a meaningful, recoverable gap. When conservative PTQ is already close to the teacher, a single aggregate score can hide task-specific movement in both directions. Deployment teams need an evaluation suite weighted toward their real workload and should compare the simpler PTQ baseline before paying for distillation.

Quantization becomes part of model development

The broader takeaway is that inference optimization works best as co-design. Hardware formats, model architecture, calibration data, context length, distributed training, and evaluations form one system. Treating any of them as an isolated afterthought leaves either performance or quality on the table.

QAD makes that system easier to reason about. First choose an aggressive but plausible low-precision configuration. Measure where it breaks. Train the model while it experiences the same simulated numerical constraints. Then evaluate the restored checkpoint across the workloads that matter, retaining the simpler PTQ result when QAD does not earn its cost.

That pattern extends beyond NVFP4 or Nemotron. A production constraint can become a training condition rather than a fixed penalty imposed at export time. The useful engineering question is no longer only how much accuracy compression loses, but which parts of the model can learn to operate reliably inside the deployment budget.