Generated by Codex with GPT 5.6 Sol XHigh
Model pricing tables make an awkward engineering problem look deceptively simple. A low price per token can still produce an expensive application if the model fails often, takes more turns to finish, repeats a growing context on every turn, or creates work that humans must repair. The useful unit of comparison is therefore not the token. It is a successful outcome at the quality level the application actually requires.
The official AWS Artificial Intelligence Blog published the post on September 11, 2026. It presents an open-source benchmarking harness for comparing OpenAI models on Amazon Bedrock with cost-oriented OpenAI API baselines, then uses three different workload shapes to show why the model with the lowest nominal rate is not necessarily the cheapest model to operate.
Normalize cost by completed work
The harness sends every model through the OpenAI Responses API and holds the evaluation code constant while switching the backend and model identifier. It compares GPT-5.6 Luna, Terra, and Sol on Amazon Bedrock with GPT-5.4 Mini and Nano on the OpenAI API. This common interface matters because it removes one large source of experimental drift: each model receives the same task representation, measurement logic, and result-recording path.
The comparison is intentionally practical rather than a claim about intrinsic model capability. The Bedrock configurations ran with reasoning disabled, while the API baselines used their defaults; provider infrastructure and model-specific settings also differ. Every run writes timestamped JSON, the grading prompts are frozen and hashed, and the charts are generated from those result files. That makes the evidence inspectable while preserving an important warning: these are measured deployment configurations, not laboratory-isolated model weights.
The first metric is cost per correct answer. Instead of dividing spend by calls or tokens, the harness divides total spend across both successful and unsuccessful attempts by the number of correct results. It applies this calculation to AIME, GPQA Diamond, and MMLU-Pro. The distinction is decisive because a cheap failed attempt contributes cost without producing value, while a more capable model can amortize a higher per-token rate across more successes.
The results separate two questions that pricing comparisons often collapse. Sol delivered the strongest accuracy in the tested reasoning benchmarks, including 75% on AIME versus 37% for Mini. Luna, after the price assumptions recorded for the July 2026 reductions, delivered the lowest observed cost per correct answer across those samples. One model therefore occupied the quality frontier while another occupied the cost-efficiency frontier. The right choice depends on whether failure is cheap or accuracy is a hard gate.
Agent turns compound the bill
Multi-turn agents add a second multiplier. The benchmark’s web-research agent uses client-managed history with store: false, so each turn resends the system prompt, earlier messages, and accumulated tool results. Context size grows roughly linearly with the number of turns, which means cumulative billed input can grow roughly quadratically. An extra turn costs more than one additional model response: it also retransmits everything the agent has already seen.
The harness tests this mechanism on a 50-question sample from DeepSearchQA using live search and page-fetch tools. A deterministic pre-pass and a frozen GPT-5.5 autorater score the answers at an F1 threshold. Mini averaged 7.6 turns and consumed about 2.3 times Terra’s input-token volume per question, largely because of repeated search loops. Luna combined fewer turns with a much lower observed cost per passing answer; Terra’s higher token price was partly offset by stronger results and shorter trajectories.
This reveals a systems property that is invisible on a rate card: tool policy and stopping behavior are cost controls. A model that chooses the right search earlier, avoids redundant calls, or finishes in five turns instead of eight reduces latency and also prevents repeated context from accumulating. Model selection, prompt design, tool routing, memory strategy, and context compression therefore belong in the same cost model.
Grade the artifact, not just the answer
Many production tasks do not have a short ground-truth string. They produce compliance briefs, care plans, financial analyses, and other artifacts whose quality depends on structure, completeness, caveats, and professional judgment. To represent that workload, the harness evaluates a 48-task slice of GDPval against human-authored rubrics.
The tested GPT-5.6 configurations earned higher rubric scores than the two baselines with reasoning disabled. Luna passed 27 of 48 deliverables, compared with 20 for Mini, while also recording the lowest observed cost per passing deliverable under the experiment’s price assumptions. Sol passed more tasks still, at a premium. This framing makes review and rework visible: if a weaker result must be regenerated or repaired by a person, the apparent savings in inference spend can disappear.
Latency measurements reinforce the need to benchmark the deployed path rather than infer performance from model size or price. In the authors’ July 2026 single-region runs, median time to first token on Bedrock was lower for matched Luna and Terra configurations, and Luna’s throughput was higher for longer outputs. The article carefully treats those numbers as a point-in-time snapshot. Shared services vary with load, observed maxima are not p99 estimates, and Sol’s deeper reasoning profile behaves differently.
The durable result is the method
The post does not establish a universal winner. Its samples are modest, prices can change, reasoning settings were not identical, and an 8,192-token cap truncated several professional deliverables. Those caveats are not defects to hide; they define where the results stop transferring. The accompanying scripts are valuable because a team can replace the public benchmarks with 50 to 100 representative tasks, set its own acceptance threshold, and rerun the same cost-per-success analysis when models, prices, prompts, or workload patterns change.
The broader engineering lesson is to optimize the complete path from request to accepted result. Token rates are only one input. Accuracy determines how often work succeeds, trajectory length determines how much context and latency accumulate, and rubric quality determines how much review or rework remains. A sound model-selection process measures all three, chooses against the application’s real failure costs, and keeps the benchmark close enough to production that changes in agent behavior become visible before they become changes in the bill.