From 10.3% to 70.6%: Mid-Tier Model Tears Up the Compute Pricing Sheet

From 10.3% to 70.6%: Mid-Tier Model Tears Up the Compute Pricing Sheet

AICoding ModelsAnthropicClaude

Sources:Anthropic 官方公布

A Sevenfold Benchmark Surge with Zero Price Increase

On September 28, 2026, Anthropic unveiled Sonnet 5.5, the second member of its Claude 5.5 model family. Among the officially released benchmarks, the most counterintuitive result arrived on the Terminal-Bench 4.0 coding benchmark: pass rates surged from 10.3% in the previous generation to a staggering 70.6%. On this flagship leaderboard for evaluating autonomous agentic coding, Sonnet 5.5 pulled ahead of its generational flagship sibling, Opus 5.5, which scored 66.4%.

In CursorBench 4.0, another prominent coding benchmark, Sonnet 5.5 logged 55.5%—comfortably beating the previous generation’s 34.1% and snapping at the heels of Opus 5.5’s 57.8%. Yet its pricing remained untouched. Positioned as the budget-conscious workhorse, its input pricing stays at $2 per million tokens, output at $10, and prompt cache reads at $0.20.

On OSWorld 2.1, which evaluates autonomous computer use in realistic desktop environments, Sonnet 5.5 scored 80.1%. It is the first Sonnet-tier model capable of completing Pokémon Red purely by inspecting raw screenshots. The gains extend well beyond software engineering: on Chartography, a benchmark measuring multimodal visual reasoning, it notched 61.6%, compared to a meager 15.6% from the previous generation.

On GDPval-AA v2.1, which benchmarks practical workflows across 44 distinct professions, its score reached 1844—virtually deadlocked with Opus 5.5’s 1846 and leaving the prior generation’s 1449 far behind. Its AA-Briefcase v1.1 score likewise climbed from 1359 to 1811. A middleweight contender on the compute pricing sheet is now delivering engineering takeover capabilities that outpace frontier flagships.

Re-Architecting Efficiency: Slashing Per-Task Costs by 90%

Even at identical per-token rates, actual spend per task has shrunk dramatically. According to Anthropic, Sonnet 5.5 consumes substantially fewer tokens when executing well-defined day-to-day tasks, while generating output more than 30% faster than its predecessor. Official empirical data shows that per-task cost dropped by as much as 30%. The leap in this generation stems from physically compressing execution steps rather than expanding compute scale.

Comparative data reveals that when Sonnet 5.5 runs capped at “medium effort,” its final evaluation scores still surpass the peak scores achieved by the prior-generation model running at full throttle. In this medium setting, the cost of completing an individual task drops to less than one-tenth of previous baseline costs.

Early testers observed that the new model excels at batching and consolidating tool calls, slashing the total inference steps required to interact with external environments. Discussions on Reddit’s developer communities converged on the same takeaway: unarchitected “vibe coding” historically incinerated massive token volumes through trial-and-error loops, but with the model consolidating requests and targeting tool invocations with precision, redundant overhead plunged off a cliff.

Cost per task vs. accuracy comparison Figure: Positioning of the model across different effort tiers. Further toward the top-left indicates greater capability per dollar spent; Sonnet 5.5 settles distinctly above and to the left of the previous generation. Source: Anthropic

Why Maxing Out Compute Led to Lower Scores

On the demanding FrontierCode 1.1 benchmark suite, however, an unexpected performance degradation surfaced. When configured at the “Xhigh” effort tier, Sonnet 5.5 posted a solid 52.1%. Yet when compute was cranked further to “Max,” its score plummeted to 46.2%—lagging behind Opus 5.5 (54.4%) and GPT-6 Sol (49.3%). Pouring in more compute no longer guarantees superior output; over-allocating agentic reasoning introduces uncontrollable systemic drag.

Anthropic’s post-mortem clarified that under the Max tier, the model triggered its code review skill far more frequently, aggressively delegating subtasks across sprawling swarms of sub-agents. This granular task distribution triggered severe execution timeouts.

Lacking strict synchronization, disparate sub-agents repeatedly introduced modifications beyond their designated scope. The FrontierCode harness imposes heavy penalties on out-of-bounds diffs. As execution boundaries dissolved across hierarchical agent layers, the model’s raw reasoning advantages were entirely swallowed by scheduling friction and compounding errors.

Benchmark cost-performance distribution Figure: Cost versus accuracy curves across additional benchmarks, illustrating Sonnet 5.5’s price-to-performance ratio relative to flagship models. Source: Anthropic

Downmarket Defense: When Mid-Tier Capabilities Demand Flagship Guardrails

When a cost-effective model commands formidable autonomous execution capabilities, its potential blast radius expands in tandem. Anthropic confirmed that Sonnet 5.5’s cybersecurity prowess matches the previous-generation flagship, Opus 5. Consequently, it is the first Sonnet-tier model equipped with full cyberattack guardrails and mandatory rollback triggers. Biological safety guardrails remain calibrated to the same stringent standards. Pushing defensive guardrails downmarket confirms that hazardous autonomous capabilities are no longer the exclusive preserve of top-shelf flagships.

Beyond thwarting attack vectors, Sonnet 5.5 is also the first mid-tier release to feature an anti-distillation classifier. Historically, adversaries only deployed automated scraping accounts against flagship models to harvest synthetic outputs for training rival systems. Adding anti-distillation defenses to a model priced at $10 per million output tokens signals that its generation quality now rivals the fidelity of foundational training baselines. The daily outputs of an affordable model have become crown-jewel assets that require aggressive defense.

Real-World Production: Slashing Iteration Cycles in Half

These benchmark figures were subsequently validated in enterprise environments. Across 118 real-world application build scenarios monitored by Base44, Sonnet 5.5 averaged just 3.6 conversational turns to reach production-grade code parity with Opus 5. Under identical task specifications, Opus 5 required an average of 7.7 turns.

Epic Games evaluated the model internally, reporting that its architectural audits and low-level data-flow reviews reached code quality standards previously seen only in premium flagship tiers. Unity engineers similarly noted that Sonnet 5.5 independently handled 90% of complex multi-step editor invocations and coding tasks, clearing rigorous runtime validation.

In a release postscript, Anthropic architects acknowledged that benchmarks capture only one dimension of capability. In open-ended, ambiguous environments, the generational flagship Opus 5.5 still demonstrates superior long-horizon planning. Meanwhile, Haiku 5.5, engineered for sub-second latency, is scheduled to join the lineup in the coming weeks. The triumph of this mid-tier model is anchored in single-step tool execution efficiency within bounded problem spaces. Its generational leap has effectively dismantled the price premium that frontier models long took for granted.

Reference Links:

  • Anthropic Official Blog
  • Reddit Developer Community Discussions
  • Real-world Benchmarks from Epic Games and Unity