Opus or Sonnet 5.5? Choose by the task.

A practical way to pick your Claude model, set its thinking effort, and understand what the benchmarks actually tell you.

In this guide

The decision in 20 seconds

Clear task, clear finish line? Start with Sonnet 5.5.
Open-ended work where better judgment changes the outcome? Start with Opus 5.5.

Then check the result. The right model is the one that meets your quality bar with the least time, cost, and rework.

Came here from video 3? This is your MODEL guide. Start with the checklist, then see the evidence behind the cost comparison. Examples and prompts are Michelle’s practical applications of the guidance.

Start with the task

Bounded means you can describe the job and recognize when it is done well enough. Unbounded means there is room for a much better answer: sharper reasoning, better tradeoffs, or a more useful recommendation.

Sonnet 5.5

For a clear finish line

  • You have the information it needs.
  • The output and constraints are specific.
  • You can check whether it succeeded.

Try: correcting typos, extracting dates, translating a routine email, or summarizing agreed actions.

Opus 5.5

For work that needs judgment

  • The question needs framing or research.
  • Several answers could work.
  • Choosing well matters more than a quick draft.

Try: a market study, a client proposal with competing priorities, or a strategy memo.

A detailed brief can still require difficult reasoning. Classify the work, not just the format. “Summarize this meeting in five bullets” is different from “Explain why this project is failing, using conflicting interviews.”

Official reference: How Anthropic positions Sonnet and Opus 5.5

Supplied webinar slide comparing a flattening value curve for bounded tasks with a rising curve for unbounded tasks.View full size
Slide supplied with video 3. Project notes date the Anthropic webinar to 1 October 2026; the exact capture time is unverified. Curves are illustrative, not measured shares of people’s work. Related official model guide

Make your default match your work

If you mostly make decisions, research, and develop ideas, Opus is a sensible starting point. If you mostly transform supplied material into a known format, start with Sonnet. A stronger model may still help on a bounded task if the first result fails your checks.

Official reference: Choose a starting model and evaluate it

Same document, different model

The verb in your request often tells you more than the file type. These are starting recommendations, not promises that one model always wins.

Choose a model for the task
Your requestA useful starting point
Meeting notes: list decisions and named ownersSonnet · Preserve the record and flag missing details.
Meeting notes: work out why the team cannot agreeOpus · Weigh conflicting explanations and propose a way forward.
Client proposal: fix typos in an approved draftSonnet · Make only the specified edits.
Client proposal: decide which offer best fits the briefOpus · Compare options, objections, and tradeoffs.
Market research: extract prices from supplied pagesSonnet · Follow a fixed schema and cite each page.
Market research: recommend which segment to enterOpus · Assess evidence, uncertainty, and alternatives.

Mixed task? Split it. Have Sonnet gather or format well-defined pieces; use Opus to assess the evidence and make the recommendation. Review the source material yourself before relying on the conclusion.

Official reference: Balance capability, speed, and cost

Choose the model, then the effort

Effort controls how much work the model puts into a response. More effort can help, but it also consumes more time and tokens. It does not guarantee a better answer.

  1. In Claude, open the model menu next to Send.

    Choose Sonnet 5.5 or Opus 5.5 from the options available to your account. Check the displayed effort setting too; model and effort are separate controls.

  2. Start with the default.

    For these models, Claude apps and Claude Code start at Medium. Sonnet 5.5 on the API defaults to High; Opus 5.5 on the API defaults to Medium. Available controls can vary by surface and account.

  3. Change one thing when you see a problem.

    Missing source? Supply it. Unclear brief? Clarify it. Weak reasoning on a complex task? Compare Opus before pushing Sonnet to Extra high or Max.

Official reference: Change models and effort in Claude

Official reference: API effort levels and defaults

Supplied effort slide showing Low, Medium, High, Xhigh and Max, with different defaults for Opus 5.5 and the Sonnet 5.5 API.View full size
Slide supplied with video 3; webinar dated 1 October 2026 in project notes, exact capture time unverified. Its “no limit” wording describes Max’s effort budget: model output limits and account limits still apply. Official effort documentation

Anthropic’s Sonnet building guide specifically suggests considering Opus when very high effort erodes Sonnet’s speed and cost advantage.

Official reference: Building with Claude Sonnet 5.5

The “twice as much” example

A cheaper token price does not always produce a cheaper finished task. In the Artificial Analysis comparison below, Sonnet uses a higher effort setting to reach a similar overall score.

The two benchmark configurations compared in video 3
Model and effortArtificial Analysis results
Opus 5.5 · Medium · Default fallbackIntelligence Index: 51 · Cost per benchmark task: $1.34
Sonnet 5.5 · Xhigh · Default fallbackIntelligence Index: 52 · Cost per benchmark task: $2.75

About 2.05× the cost, at similar scores

$2.75 ÷ $1.34 ≈ 2.05. That supports this specific comparison in the video. It does not mean Sonnet always costs twice as much, or that the two models give equally good answers to every task.

Values checked 3 October 2026: Opus Medium · Sonnet Xhigh. Both configurations are labelled “Default Fallback” by Artificial Analysis.

Artificial Analysis scatterplot of intelligence versus cost per benchmark task for Opus and Sonnet 5.5 effort variants.View full size
Unaltered chart supplied for video 3, recorded in project notes on 2 October 2026; exact capture time unverified. The two values in the table were separately checked on the model pages on 3 October. Tap to read the original at full size. Artificial Analysis · interactive chart

Read the chart in three steps

  1. Up means a higher benchmark score. It is an aggregate index, not a percentage of tasks answered correctly.
  2. Left means a lower benchmark cost. The dollar axis is logarithmic; equal spacing does not mean equal dollar increases.
  3. Check the effort label. Medium versus Xhigh compares two configurations, not just two model names.

The cost is a weighted average across the benchmark’s tasks, including input, cached tokens, reasoning, and answers. It is not a quoted price for your email or proposal. The index mixes knowledge work, coding, reasoning, and other evaluations; your own tasks can give a different result.

Independent methodology: How the Intelligence Index is measured.

Official reference: Compare models by completed work, not token price

API cost or Claude subscription?

These are different ways to pay for Claude. Use the benchmark dollars to understand the tradeoff, not to predict how many messages remain in your plan.

API pricing versus subscription usage
How you use ClaudeWhat the numbers mean
API · Sonnet 5.5Standard input: $2 per million tokens. Output: $10 per million.
API · Opus 5.5Standard input: $4 per million tokens. Output: $20 per million.
Claude subscriptionYour plan has usage limits. Model, effort, conversation length, and features affect consumption; the benchmark ratio is not a quota conversion.

Standard API rates checked 3 October 2026. Caching, batch processing, tools, and provider-specific options can change the bill. For repeat work, compare the cost of a result you can actually use, including retries and your own editing time.

Official reference: Official Sonnet and Opus token prices

Official reference: Understand and manage Claude usage

Try it on your next task

Choose one real task and write down what “good” looks like before you run it. Here are two prompts to adapt.

A bounded task → try Sonnet

Prompt · A clear finish line
Turn the meeting notes below into an action list.

Use a table with Action, Owner, and Due date. Include only explicitly agreed actions. Preserve names and dates. Write "Not specified" where an owner or date is missing. Do not turn suggestions into commitments.

After the table, list any ambiguities I should check. Do not add recommendations.

Meeting notes:
[paste your notes]

Check: every action exists in the notes, no owner is invented, and suggestions remain separate.

An open-ended task → try Opus

Prompt · Reason through a decision
Help me decide which of these three client proposal options best fits the brief.

Client goal: [goal]
Options: [A, B, C]
Constraints: [budget, timing, non-negotiables]
Evidence: [attach the brief and relevant notes]

First identify the decision criteria and any missing information that could change the recommendation. Compare the options against those criteria, explain the strongest case against your preferred option, and recommend a next step.

Separate evidence from assumptions. Reference the supplied sources for factual claims. Do not invent client priorities or numbers. If a missing fact is decisive, ask me before making a firm recommendation.

Check: the recommendation respects the brief, considers a credible alternative, and says what would change it. A fluent answer with unsupported claims fails the test.

Make a small comparison

Run the same prompt with the same files in two fresh chats. Record model, effort, time, factual errors, missed requirements, and edits needed. For API work, record total cost too. Repeat on a few representative tasks before changing a recurring workflow.

Official reference: Test the choice against your requirements

For agents: plan, delegate, review

If you build agents, a useful pattern is Opus for the plan and final judgment, Sonnet for clearly specified work. The orchestrator is the agent that assigns tasks and brings the results together.

  1. Opus defines the job.

    For a market study, decide which competitors to compare, which sources count, and what evidence each worker must return.

  2. Sonnet handles independent pieces.

    Each worker extracts the same facts about one competitor: pricing, target customer, source URLs, dates, and missing information. These tasks can run in parallel.

  3. Opus reviews and synthesizes.

    Check the evidence, resolve contradictions, and produce the recommendation. Send unclear or difficult pieces back for further work.

Prompt · Define a worker task
Complete only this subtask: [specific task].
Inputs and allowed sources: [files or URLs].
Return: [exact format and required fields].
Success checks: [conditions that must be met].
Do not: [scope boundaries].
If information is missing or conflicting, flag it with its source instead of guessing. Return the result, supporting references, and unresolved questions for review.

The builder must configure which model each agent uses; writing “use Sonnet workers” in a normal chat does not create that setup. Parallel agents add coordination costs. Use them when the pieces are independent and substantial enough to justify it, and measure the whole run.

Official reference: Official orchestrator strategy and its tradeoffs

Go straight to the sources

Checked 3 October 2026. Model options, prices, and benchmark results change. The linked documentation is the place to recheck.

Anthropic’s model selection guideOfficial documentation · Where to start and how to evaluate your choice.Sonnet 5.5: release and benchmarksOfficial documentation · Positioning, pricing, results, and evaluation footnotes.Find the model and effort controlsOfficial documentation · Help for using Claude’s model menu.Understand thinking effortOfficial documentation · API settings, defaults, and limitations.Optimize cost and intelligenceOfficial documentation · Completed-task costs and multi-model architectures.Compare Opus and Sonnet on Artificial AnalysisIndependent benchmark · Your selected models and effort levels.

Read the benchmark methodology before treating an index score as a prediction for your own work.

Your next small step

Pick the next task on your list. Define the finish line, choose a starting model, and check the answer against the original material. Keep the model that earns its place in your workflow.