Anthropic introduced Claude Opus 5.5 on September 22, 2026, with input and output prices of $4 and $20 per million tokens. The company says those rates are 20% below Opus 5, while typical workloads in its testing cost 40% less to run. Those are different claims.
The release announcement also lists cache reads at $0.20 per million tokens and reports faster generation. Anthropic’s Transparency Hub summarizes model evaluations and safeguards. Neither source substitutes for evaluating a particular application’s quality and complete cost.
Why task savings can differ from token savings
A model can consume fewer tokens, reuse more cached input, or complete a task with fewer attempts. Those changes can reduce total spend beyond the nominal price cut. Conversely, a demanding configuration or a longer output can offset cheaper rates.
For budgeting, separate input, output, cached input, retries, and any external tool usage. Compare matched tasks and settings. A vendor’s average savings should not be applied automatically to every request type or treated as a contractual prediction for a customer’s monthly bill.
Quality needs to survive the cheaper execution
Anthropic reports improvements in coding, professional work, and behavioral evaluations. The release also acknowledges limits. A claim of safer behavior does not authorize an agent to make irreversible changes without application-level controls.
Measure the result a reviewer accepts. For engineering work, include tests, behavior preservation, and the clarity of the change. For documents, check factual accuracy, source use, and required formatting. Record tasks that need substantial human correction rather than counting the initial answer as a success.
Migration needs configuration checks
The announcement identifies claude-opus-5-5 for developers and states that thinking mode can no longer be switched off. Existing integrations that depend on a particular thinking configuration should inspect the migration guidance before changing the model.
Preserve the old configuration and a reproducible task set during evaluation. Check streaming, tool behavior, and error handling alongside response quality. An application that parses outputs or coordinates several models needs to verify those boundaries, not only that a single prompt produces a plausible answer.
When Opus 5.5 deserves a trial
Teams running complex tasks where review and retries are expensive have a concrete reason to evaluate the release. The useful comparison includes latency, accepted-result cost, and the amount of expert attention required.
Nerova’s assessment is that the pricing change is meaningful, particularly when cache reads make up a large share of a workflow. The exact value remains workload-specific. Keep the 20% rate cut, claimed 40% task savings, and the application’s measured outcome as separate numbers.