Anthropic announced Claude Opus 5.5 on 22 September 2026. This article is covering that release on 30 September because it is a major model update that had not yet been covered here; it did not launch today.
The model is positioned for long-running agentic coding and knowledge work. The practical API story combines lower prices with a 1-million-token context window, up to 128K output tokens and several breaking changes. A production team cannot safely migrate by changing only the model string.
Anthropic says Opus 5.5 costs 40% less than Opus 5 on typical workloads and generates output more than 30% faster. Those are company-reported measurements, not independent guarantees. The useful question is whether the model lowers cost per successful task on your own workload after thinking tokens, retries, tool calls and human review are included.
Price changes the routing decision
Standard Claude API pricing is $4 per million input tokens and $20 per million output tokens, down from $5 and $25 for Opus 5. Cache reads cost $0.20 per million tokens instead of $0.50, while cache writes are $5 instead of $6.25. Fast mode is available at $8 input and $40 output per million tokens.
This pricing makes Opus 5.5 more plausible for long agent sessions, but it does not automatically make it the default for every request. A short extraction job may still belong on a smaller model. A repository migration, multi-step investigation or high-value knowledge workflow may justify Opus. Route by task difficulty and failure cost, then measure successful outcomes.
Build a cost record per candidate run: uncached input, cache writes, cache reads, thinking and visible output, tool requests, wall-clock time, retries and final task status. Divide total spend by accepted tasks rather than by API calls. A model that uses fewer calls can be cheaper even with a higher token price; a model that produces long failed trajectories can be more expensive than its headline suggests.
The specification is materially different
The official model page lists a 1M-token context window, 128K maximum output, adaptive thinking that is always on and a default effort of medium. The Claude API model ID is `claude-opus-5-5`. The context-window beta header used by some older integrations is unnecessary.
The 1M context window is capacity, not a reason to send every available document. Retrieval quality, ordering, prompt-cache boundaries and privacy still matter. Large prompts increase latency and create more places for conflicting instructions. Test the smallest context that preserves task quality, then add evidence only when it changes outcomes.
Maximum output also includes thinking. A request that previously budgeted only for visible text can stop early after migration because `max_tokens` covers thinking plus the response. Anthropic recommends controlling depth through effort and starting with a large output allowance for xhigh or max effort before tuning down.
Thinking can no longer be disabled
Opus 5.5 accepts no thinking field or adaptive thinking. Requests that send disabled thinking or a manual thinking-token budget return HTTP 400. The only supported depth control is `output_config.effort` with low, medium, high, xhigh or max.
This changes response parsing. A response can begin with a thinking block, so code that assumes `content[0].text` is fragile. Select blocks by type. In tool loops, return the assistant message—including signed thinking blocks—complete and unmodified when submitting tool results. Dropping, editing or reordering those blocks can invalidate the conversation.
Thinking text is omitted by default. If a product previously streamed narration between tool calls as progress, users may now see a long quiet period. Use the documented display settings for summarized or progress updates when that experience is appropriate, and make clear that progress text is not a durable execution log.
Forced tool selection now fails
Tool choice `any` or a forced named tool is rejected. Anthropic directs integrations to automatic tool selection, combined with strict tool schemas or structured outputs. That is not a cosmetic request change: an application that relied on forced tool choice as an authorization boundary needs a new design.
Keep authorization outside the model. The agent may decide that a tool is relevant, but the application must still validate the tool name, arguments, user permission, idempotency key and risk policy. Use a narrow strict schema, set `additionalProperties: false` for objects where required, and reject mutations that lack an explicit product approval.
Sampling overrides such as non-default temperature, top-p and top-k are also rejected. Assistant-message prefills are not supported. Replace prefills with structured output or clear system instructions, and test behavior rather than assuming a request accepted by Opus 5 remains valid.
Computer-use agents need a toolset migration
On the Claude API and Google Cloud, the older `computer_20251124` tool is rejected. Integrations must declare `computer_toolset_20260801`, remove the former beta header and adapt the loop to member tool-use blocks, potentially several in one turn. Each tool result must echo the documented toolset name.
Anthropic notes that Amazon Bedrock keeps different compatibility behavior, so do not apply one payload blindly across providers. Maintain provider-specific conformance tests around model identifiers, tool versions and response shapes. Computer use should run in a constrained environment with domain, action and data boundaries; a stronger model is not a substitute for sandboxing and approvals.
Routing and fallback need conversation rules
Thinking blocks are tied to the model and conversation. Opus 5.5 can read thinking blocks from several earlier Claude families, while most fallback models cannot consume Opus 5.5 thinking blocks. A router that changes models mid-conversation may therefore lose that hidden state.
Treat the visible conversation, tool results and application state as the portable record. Keep conversations append-only where the API requires it. If a fallback must take over, summarize the verified task state into a fresh request instead of assuming internal reasoning transfers. Test refusals explicitly: Opus 5.5 may return `stop_reason: refusal` with a category, and not every refusal is eligible for the same server-side fallback.
Benchmark claims need local evaluation
Anthropic reports gains on agentic coding, computer use and knowledge work and publishes detailed benchmark settings and caveats. It also says benchmark margins are becoming less reliable guides to real-world differences. That warning should shape the migration.
Create a release set from real tasks: repository changes that must pass tests, research reports whose facts can be checked, tool workflows with exact side effects, and long-context cases that reflect production. Evaluate correctness, critical failure rate, completion time, token use, cost per accepted task and reviewer effort. Run every candidate at more than one effort level; medium is the new default, while Opus 5 defaulted to high.
Separate vendor claims from your evidence. Record the model ID, effort, prompt and tool versions, dataset version, cache state and retry policy. A one-time score without those inputs cannot support a routing decision.
A migration plan that limits risk
First, inventory every request builder and response parser. Search for the old model ID, explicit thinking settings, manual budgets, forced tools, sampling parameters, prefills, beta headers, positional text parsing and the older computer-use type.
Second, update the contract in a development environment. Read blocks by type, preserve thinking blocks in tool loops, set effort explicitly, validate strict schemas and handle refusal and context-limit stop reasons. Revisit `max_tokens` so thinking has room without creating an uncontrolled output budget.
Third, replay a frozen evaluation set on the old and new bundles. Measure total task cost and latency rather than token prices alone. Inspect regressions by workflow slice and verify that user-visible progress, cancellation, timeouts and audit trails still behave correctly.
Fourth, run shadow traffic with mutations disabled. Then canary a small eligible cohort behind a reversible route. Keep the complete previous bundle—model, prompts, tools and parser—deployable. Roll back the bundle, not only the model name.
The practical decision
Claude Opus 5.5 offers a meaningful price and capacity change for demanding agents, but it also tightens the API contract. Teams with long coding, research or computer-use workflows have a strong reason to evaluate it. Teams with simple tasks should compare it against smaller models rather than assume the flagship is the economical choice.
The migration is ready when request validation passes, response blocks are parsed safely, tool authorization remains external, cost per accepted task improves, failure paths work and rollback restores a previously evaluated bundle. Until then, the lower price is an invitation to test—not permission to swap production traffic blindly.
Official references
These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.
Prepared by: Noor Yasser
Working through a similar engineering challenge?
I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.




