Generated by Codex with GPT 5.6 Sol High

Techmeme surfaced Anthropic’s September 28 launch, “Claude Sonnet 5.5”. The release is less about setting a new absolute capability frontier than compressing frontier-like work into a faster, cheaper model. Anthropic positions Sonnet 5.5 below Opus 5.5 for open-ended judgment, but close enough on coding and professional tasks to handle much of the high-volume work that would otherwise require a premium model.

The largest claimed gain is in agentic coding. Anthropic reports a 70.6% score on Terminal-Bench 4.0, up from Sonnet 5’s 10.3%, and says low- or medium-effort configurations can beat the predecessor’s best results at roughly one-tenth the cost per task on several evaluations. Sonnet 5.5 keeps Sonnet 5’s token prices—\$2 per million input tokens and \$10 per million output tokens—but generates text more than 30% faster and typically uses fewer tokens and tool calls. The economic claim therefore comes from finishing work with less inference, not a lower list price.

That distinction matters for engineering teams. If routine bug fixes, reviews, document work, and multi-step agent runs can move from a top-tier model to a smaller one, the practical bottleneck shifts toward routing: deciding which tasks need Opus-level judgment and which can use Sonnet without sacrificing reliability. Anthropic’s own examples suggest that lower effort often offers the best capability-per-dollar tradeoff, while higher effort can cost about as much as Opus for similar results.

The release also shows capability and containment advancing together. Anthropic says Sonnet 5.5’s cyber abilities are comparable to those of Opus 5, so this is the first Sonnet release with its strongest cyber safeguards. Higher-risk requests can fall back to Sonnet 5, while new classifiers target attempts to extract the model’s reasoning through industrial-scale distillation. The company’s behavioral audit found no evidence that the model pursued goals contrary to user intent, but Anthropic explicitly notes that no evaluation suite can expose every failure.

The numbers remain launch-day evidence, much of it generated or selected by Anthropic and run at differing effort levels. Some comparisons also use external pre-release deployments or older competitor results. The durable takeaway is therefore not that one benchmark settles the model race. It is that frontier performance is moving down the cost curve quickly enough that model selection, effort settings, and safeguards are becoming first-class parts of software architecture.