Claude Sonnet 5.5 promises lower AI bills without a price cut

Anthropic says Claude Sonnet 5.5 is faster and more token-efficient. I tested it against Sonnet 5 on the same marketing task to see what changed.

Anthropic has released Claude Sonnet 5.5, saying that the model can complete work faster while using fewer tokens than its predecessor. In its own testing, the company says that can reduce the cost of a task by up to 30%.

The reduction is not in terms of cost per token as the prices for tokens remain unchanged: $2 per million input tokens, $10 per million output tokens, $0.20 per million cache-read tokens and $2.50 per million cache-write tokens.

Improvements on Sonnet 5.5

Anthropic says token efficiency is one of Sonnet’s 5.5 advantages when it comes to completing the same work compared to Sonnet 5. In coding workflows, the company says early testers also saw fewer steps and tool calls.

On Anthropic’s reported benchmarks, Sonnet 5.5 scored 70.6% on Terminal-Bench 4.0, compared with 10.3% for Sonnet 5. Anthropic describes Terminal-Bench as an evaluation of agentic coding performance

It demonstrates better judgment across long-horizon work, catches data errors by re-checking source materials, and avoids over-relying on web search tools compared to Sonnet 5.

There are strong improvements in image understanding and visual reasoning. One example Anthropic lists is playing Pokémon Red operating solely from screenshots.

It is also the first Sonnet model integrated with Anthropic’s elevated cybersecurity safeguards and automated fallbacks previously reserved for flagship Opus models.

I tested Sonnet 5.5 against Sonnet 5

To see whether Anthropic’s speed and efficiency claims were visible in a simple marketing workflow, I gave Sonnet 5 and Sonnet 5.5 the same 150-word promotional-email brief for a fictional coastal hotel.

Both completed the task in two seconds, and neither produced a materially different result: both incorporated the core information supplied in the brief. Claude’s free interface did not expose token usage, so the test could not independently verify Anthropic’s claim that Sonnet 5.5 uses fewer tokens or costs less per task. Neither model reached the free-tier usage limit during the test.

Anthropic provides its own example of the model’s knowledge-work performance. The company says it gave Sonnet 5.5 a public company’s quarterly earnings materials and call transcripts, together with a slide template, and asked it to create a 10-slide operating review. From the response, two experts judged the first draft ready to send as is.

On external-company testing, Anthropic cites that Slack says it performed better than Sonnet 5 on almost all of its offline Slackbot evaluations while using about 14% fewer output tokens. Zendesk says its support tickets were processed 20% faster, while Box reports that Sonnet 5.5 was 2.4 times faster and used 12% fewer total tokens in its testing. These are company-reported results, rather than independent SEW tests.

Key takeaway

Because our consumer-interface test couldn’t verify token usage, the real question isn’t just whether Sonnet 5.5 is faster. It is whether it uses fewer tokens while producing an output that still meets the same quality bar for the task you actually run.

For marketers considering the upgrade, the useful measurement is therefore not simply the model’s listed token price. Compare the cost and token consumption of the actual tasks you run, alongside output quality and completion time, if your interface or API exposes those measurements.

 

 

Want SEW higher in your Google results?Add as a preferred source

More in SEO News

View more
SEO News

Google AI Mode now showing more blog carousels

Google AI Mode is showing more blog carousels than before, giving websites that publish articles regularly another way to earn visibility inside AI Mode answers.

Start the conversation by posting the first comment

Join the conversation

Posting publicly · your email is never shown