Anthropic has released Claude Sonnet 5.5, a mid-tier model that outperforms the company's own flagship Opus 5.5 on coding benchmarks — and does it at half the price per token. The model also tops Terminal-Bench 4.0, surpassing Anthropic's flagship in that test as well. The release inverts the usual product ladder, where the most capable model sits at the top of the price sheet.
Benchmarks put Sonnet ahead of Opus on code
Coding benchmarks are the clearest win for Sonnet 5.5. The model beats Opus 5.5 on those tests, according to the facts released with the launch. Terminal-Bench 4.0 also goes to Sonnet, which means the cheaper model handles the terminal-style tasks that tend to trip up coding assistants. For developers choosing between the two, the benchmark results point in one direction.
The pricing gap is what makes the benchmark result matter. Sonnet 5.5 costs half as much per token as Opus 5.5. Teams that were paying flagship rates for coding work now have a cheaper option that scores higher on the tests Anthropic is using to compare the two.
Token burn is the catch an independent tester found
There's a wrinkle. An independent tester found that Claude Sonnet 5.5 burns more tokens than any model they have measured. Token burn is the amount of model output consumed to complete a task. A model that's cheaper per token can still cost more per job if it uses enough extra tokens to finish. That's the tension in the launch: lower sticker price, higher usage per task.
Anthropic hasn't addressed the tester's finding in the release materials. The two data points sit side by side without reconciliation — half the per-token price, more tokens consumed per job than anything the tester has measured.
What developers will have to work out
The practical question is whether Sonnet 5.5's higher token burn eats the savings from its lower per-token rate. That depends on the workload. Coding tasks with clear stopping points may not run long enough for the burn to erase the price advantage. Open-ended or agentic tasks, where the model keeps generating until it decides it's done, are where token consumption tends to pile up.
Terminal-Bench 4.0 is one of those agentic-style evaluations, and Sonnet 5.5 leads it. But leading the benchmark and leading on cost-per-completed-task aren't the same measurement, and only one of those has an independent number attached so far.
Anthropic hasn't published a cost-per-task comparison between Sonnet 5.5 and Opus 5.5. Until it does, or until enough developers run their own workloads, the token burn finding stays an open variable in what otherwise looks like a straightforward upgrade.
The pricing structure itself is the story
Anthropic's lineup has, until now, followed the standard pattern: more capability, more cost. Sonnet 5.5 breaks that on coding benchmarks and Terminal-Bench 4.0, where the cheaper model wins. Opus 5.5 still carries the flagship name and the flagship price, but the benchmark sheet doesn't back the hierarchy on the tests Anthropic chose to highlight.
That leaves buyers with a decision the company hasn't made for them. The per-token math favors Sonnet. The token-burn finding cuts against it. Which one dominates depends on what the model is being asked to do.
Anthropic has released the model and the benchmark results. It hasn't released a token-efficiency comparison against Opus. The independent tester's measurement stands as the only outside data point on consumption so far. Developers evaluating the switch will likely want their own numbers before committing a production workload — particularly if that workload runs long.




