Skip to main content
Elios logo
Elios
Elios InsightsAcademyAbout Us
Talk with Elios

BLOG POST

How I choose Claude models to manage team costs

Matthew Groff, VP of Technology ยท 9 min read

Claude logo above Haiku, Sonnet, Opus, and Fable, surrounded by question marks with dollar bills burning below.

Opus 5 Medium is my recommended default for substantive work in Claude Cowork and Claude Code. If you're responsible for your team's Claude budget, I would start there, keep Sonnet 5 Medium for simple tasks, and teach people when to change models or reasoning effort.

Clients keep asking me about this. I've heard the same concern from two different IT teams: people keep using Opus, and Opus costs too much. I understand the concern. But the model name alone doesn't tell you what it costs to get the work right.

I've personally found that Sonnet can take more explanation, corrections, and rewrites to produce something usable. You may eventually get the result you want, but you've consumed more tokens and employee time getting there. Those extra attempts can erase the saving from choosing a model with a lower price per token.

These are my recommendations as of September 6, 2026, based on personal experience and published benchmarks. Your results may vary, and models, prices, included usage, and settings change quickly. I haven't measured the cost of my own retries; the benchmark figures below come from the named evaluators.

Choose the model and the reasoning effort together

Reasoning effort controls how much work the model puts into solving a task. Medium, High, and Max are settings within a model, and increasing effort can increase token consumption and time. It can also change how extensively an agent explores a problem and uses its tools. More effort doesn't guarantee a better result. Anthropic explains the trade-offs in its effort documentation.

That distinction matters when someone calls Opus expensive. Opus Medium and Opus High have the same token prices, but they can consume very different amounts of tokens to complete the same work. Sonnet at a high reasoning setting can cost more per attempt than Opus at Medium.

My general recommendation is:

Matthew Groff's Claude model and reasoning effort recommendations
Model and effortClaude CoworkKnowledge workClaude CodeSoftware work
Sonnet 5Medium effort
Simple email edits, proofreading, and short rewrites you can readily review.Narrow, mechanical edits with clear checks.
My defaultOpus 5Medium effort
Default for substantive work: spreadsheet analysis, documents, presentations, and research synthesis.Default for substantial coding: features, bug fixes, tests, refactoring, and repository investigation.
Opus 5High effort
Difficult analysis, conflicting sources, or substantive errors at Medium.Complex debugging, changes across systems, or substantive failures at Medium.
Fable 5.1Medium effort
My final escalation for difficult work that Opus High isn't resolving.My final escalation for persistent implementation or architectural problems.
Sonnet 5Medium effort
Claude CoworkKnowledge work
Simple email edits, proofreading, and short rewrites you can readily review.
Claude CodeSoftware work
Narrow, mechanical edits with clear checks.
My defaultOpus 5Medium effort
Claude CoworkKnowledge work
Default for substantive work: spreadsheet analysis, documents, presentations, and research synthesis.
Claude CodeSoftware work
Default for substantial coding: features, bug fixes, tests, refactoring, and repository investigation.
Opus 5High effort
Claude CoworkKnowledge work
Difficult analysis, conflicting sources, or substantive errors at Medium.
Claude CodeSoftware work
Complex debugging, changes across systems, or substantive failures at Medium.
Fable 5.1Medium effort
Claude CoworkKnowledge work
My final escalation for difficult work that Opus High isn't resolving.
Claude CodeSoftware work
My final escalation for persistent implementation or architectural problems.

I don't recommend moving everyone to Max. I also don't include Haiku, older Opus models, Opus Low, or Fable above Medium in this standard guidance. This is the starting policy I'd give a team, not an exhaustive account of every setting that might suit a specialist workload.

Why I favor Opus Medium for knowledge work

Artificial Analysis's AA-Briefcase benchmark evaluates professional work involving spreadsheets, presentations, documents, and other files. Its overall Elo score combines rubric performance, analytical quality, and presentation. Higher is better; Elo isn't a percentage of tasks completed correctly.

AA-Briefcase knowledge-work results
Model and effortOverall EloCost per task attemptApproximate USD
Sonnet 5Medium effort
1,056$1.73
My defaultOpus 5Medium effort
1,444$5.25
Sonnet 5Max effort
1,360$14.43
Opus 5High effort
1,561$10.41
Sonnet 5Medium effort
Overall Elo
1,056
Cost per task attemptApproximate USD
$1.73
My defaultOpus 5Medium effort
Overall Elo
1,444
Cost per task attemptApproximate USD
$5.25
Sonnet 5Max effort
Overall Elo
1,360
Cost per task attemptApproximate USD
$14.43
Opus 5High effort
Overall Elo
1,561
Cost per task attemptApproximate USD
$10.41

Sources: Artificial Analysis results for Sonnet 5 Medium, Opus 5 Medium, Sonnet 5 Max, and Opus 5 High. Costs reflect the benchmark's token usage and pricing assumptions, including caching.

Sonnet Medium was much cheaper per attempt, with a lower score. My Opus recommendation accepts that higher initial cost for better substantive results. Whether it saves money after corrections depends on your work.

Opus Medium scored higher overall than Sonnet Max at approximately 64% lower cost per attempt. Sonnet Max scored better on presentation than Opus Medium, so the comparison doesn't favor Opus on every dimension. Opus High exceeded Sonnet Max's presentation score while still costing less.

That supports my recommendation for substantive knowledge work. It doesn't establish a 64% saving on your company's Claude bill. AA-Briefcase uses an independent agent environment, not Claude Cowork or its Office integrations, and it doesn't measure your employees' correction time.

The coding results also make a case for Medium

Datacurve's DeepSWE benchmark tests substantial repository changes with programmatic checks. These results use the same mini-swe-agent setup across models, which makes them useful for comparing model and effort choices within that environment.

Datacurve DeepSWE results using mini-swe-agent
Model and effortDeepSWE pass@1Cost per attemptApproximate USD; * Sonnet estimates
Sonnet 5Medium effort
39.8%$2.72*
Sonnet 5High effort
48.2%$4.95*
Sonnet 5Max effort
53.8%$17.60*
My defaultOpus 5Medium effort
68.9%$3.29
Opus 5High effort
72.8%$6.08
Sonnet 5Medium effort
DeepSWE pass@1
39.8%
Cost per attemptApproximate USD; * Sonnet estimates
$2.72*
Sonnet 5High effort
DeepSWE pass@1
48.2%
Cost per attemptApproximate USD; * Sonnet estimates
$4.95*
Sonnet 5Max effort
DeepSWE pass@1
53.8%
Cost per attemptApproximate USD; * Sonnet estimates
$17.60*
My defaultOpus 5Medium effort
DeepSWE pass@1
68.9%
Cost per attemptApproximate USD; * Sonnet estimates
$3.29
Opus 5High effort
DeepSWE pass@1
72.8%
Cost per attemptApproximate USD; * Sonnet estimates
$6.08

Source: Datacurve's downloadable results. Pass@1 estimates success from a single attempt, using the benchmark's repeated runs.

The Sonnet costs are my repricing estimates, not newly measured bills. Datacurve reports approximately $4.08, $7.43, and $26.40. Those figures appear consistent with $3/$15 input/output pricing. I multiplied the unrounded costs by two-thirds to estimate costs at Sonnet 5's current $2/$10 prices, assuming the same proportional reduction applies throughout the recorded usage.

Under that assumption, Opus Medium achieved a substantially higher success rate than Sonnet Medium for about 21% more money per attempt. It also outperformed Sonnet High and Max at a lower estimated cost.

Artificial Analysis separately tested Opus inside Claude Code:

Artificial Analysis results for Opus 5 inside Claude Code
Model and effortDeepSWE pass rateCoding Agent IndexOverallAverage cost per taskFull coding suite, USD
My defaultOpus 5Medium effort
63%64$3.17
Opus 5High effort
61%66$3.92
Opus 5Max effort
63%67$8.94
My defaultOpus 5Medium effort
DeepSWE pass rate
63%
Coding Agent IndexOverall
64
Average cost per taskFull coding suite, USD
$3.17
Opus 5High effort
DeepSWE pass rate
61%
Coding Agent IndexOverall
66
Average cost per taskFull coding suite, USD
$3.92
Opus 5Max effort
DeepSWE pass rate
63%
Coding Agent IndexOverall
67
Average cost per taskFull coding suite, USD
$8.94

Source: Artificial Analysis's Claude Code results and Coding Agent Index methodology.

Medium matched Max's rounded DeepSWE pass rate at approximately 65% lower average cost across the full suite. That cost comparison covers the combined suite, not DeepSWE alone. High improved the overall coding index while scoring lower on DeepSWE in this particular evaluation.

This is why I start at Medium and increase effort selectively. These are separate evaluations with different agent setups; their scores shouldn't be combined into one ranking. All attempt costs include unsuccessful work, rather than guaranteeing an accepted result.

Escalate before repeated rewrites consume the saving

If the model is missing information, give it the information. If you haven't decided what you want, settle that before asking it to produce another version.

But when the model has what it needs and still struggles to understand the task or repeats substantive mistakes, I would escalate before spending an hour correcting it. A small edit is normal. Repeatedly regenerating an entire report or implementation is a reason to reconsider the model and effort setting.

In my experience, Fable 5.1 is more capable and makes fewer mistakes on the difficult work I escalate to it. That's why Fable 5.1 Medium is my final escalation after Opus 5 High. If Fable resolves a task in one attempt that would otherwise take repeated Opus attempts, the higher token price can be worthwhile. The cited comparisons don't establish Fable Medium as a universal winner over Opus High.

At Anthropic's current base API prices, Sonnet 5 costs $2 per million input tokens and $10 per million output tokens. Opus 5 costs $5/$25, and Fable 5.1 costs $10/$50. Output costs five times as much per token as uncached base input for each model. Repeated writing and rewriting adds up, although long input histories and reasoning also contribute to the total.

Included usage and metered spend need different checks

For Team and legacy seat-based Enterprise plans, usage credits let people continue after included limits, with those credits billed at standard API rates. Current usage-based Enterprise plans charge separately for all usage; their seat fee covers access.

Fable also needs attention. As of this post's date, it uses paid credits from the start on standard Team and legacy standard Enterprise seats. Premium seats on those plans include Fable within a limited share of weekly usage. Check Fable's plan-specific terms before making it part of your team's escalation guidance.

A benchmark's API-dollar saving doesn't translate directly into the same percentage of included usage saved. Check your actual plan and usage records. For metered work, compare spending alongside accepted results. For included usage, look at whether people finish their work before reaching limits. In both cases, include the time people spend reviewing and correcting the output.

Another way to save: agree on the plan before generating the work

Model selection is only part of my advice. For complex work, especially software development, I use Grill Me to get clear on what we're doing before asking the agent to do it.

Matt Pocock created the original Grill Me skill. My version expands the places the agent should check before asking me a question. It asks one question at a time, gives a recommended answer, and makes me resolve decisions that would otherwise become assumptions in the work.

Here's the version I made for Claude Code/Cowork

/grill-me
The Elios adaptation of Matt Pocock's pattern for broader AI-native workflows.
View the full skill
---
name: grill-me
description: >
  Interview the user relentlessly about a plan or design until reaching shared
  understanding, resolving each branch of the decision tree. Use when the user
  wants to stress-test a plan, get grilled on their design, or mentions "grill me".
---

Interview me relentlessly about every aspect of this plan until we reach a shared understanding. Walk down each branch of the design tree, resolving dependencies between decisions one-by-one. For each question, provide your recommended answer.

Use `AskUserQuestion` to ask me questions. Ask one question at a time so each answer can inform the next.

Before asking me a question, check whether you can answer it yourself from the conversation history, files or documents on this machine, through your connectors, or the `WebSearch` tool. Only ask me questions that require human judgment, context I hold, or a design decision.

Use it after describing the task and providing the relevant files or context. Agree on the scope, constraints, and what an acceptable result looks like before moving into execution. The prompt names Claude's tools; in another environment, use its available question and search tools. For more on saving this as a reusable skill, see how to write a good AI agent skill.

Planning consumes tokens too. My reason for doing it is to avoid spending far more time and usage producing work against assumptions I never agreed to.

Plan with Opus or Fable, then try Sonnet for execution

Once a strong model and I have agreed on a clear plan, Sonnet can be a reasonable choice for carrying out well-defined work with clear checks. I would consider this when the hard part is deciding what to do and the implementation is comparatively routine. Opus Medium remains my general default when the work still requires substantial judgment throughout.

Claude Code supports this pattern through opusplan, which uses Opus in plan mode and Sonnet for execution. Confirm the actual model versions your account uses, because aliases and organization settings affect the selection.

I recommend this as an option to test on suitable work. Keep the plan available, check the result against it, and move back to a stronger model if execution starts requiring repeated correction.

Caching is another way to reduce costs. It's outside the scope of this post, but Anthropic's explanation of how Claude Code uses prompt caching is worth reading.

Give your team a default, then check the results

My advice to clients is to start substantive work on Opus 5 Medium, use High when the task warrants it, and reserve Fable 5.1 Medium for difficult work that still isn't getting resolved. Keep Sonnet 5 Medium for simple tasks and consider it for execution after a clear plan is agreed.

Apply that guidance to recurring work your team actually does. Record the model and effort, whether the result was accepted, how much correction it needed, and the usage or spend involved. Use those observations to adjust the default as the models and your team's needs change.

The question I want a leader to answer is what it took to get useful work completed. That's the basis on which I'd manage a Claude budget.

Subscribe to our newsletter

New posts delivered to your inbox. Playbooks, signals, and hard lessons in AI deployment.

By clicking โ€œSubscribe Nowโ€ you agree to our Terms of Use and Privacy Statement. You can unsubscribe anytime.

FEATURED BLOG POSTS

Latest Insights from Elios

How I choose Claude models to manage team costs
How I choose Claude models to manage team costs
Matthew Groff, VP of Technology ยท 9 min read
The One-Pizza Team
The One-Pizza Team
J. Campos ยท 10 min read
What Is an MCP Server?
What Is an MCP Server?
Matthew Groff, VP of Technology ยท 12 min read
See All Blog Posts

Wherever you are with AI, weโ€™ll help you deploy it.

Weโ€™ll deploy a qualified specialist or AI Pod in less than seven days, and leave the capability behind.

Talk with Elios
Elios logoElios

Make your team AI-native. We deploy AI into your business and the people who run it.

llms.txt

Solutions

  • Forward Deployed Engineers
  • Forward Deployed Specialists
  • Accelerators

Engage

  • Talk with Elios
  • How We Engage
  • Executive AI Roadmap
  • AI-native Journey
  • How We Screen AI-native Specialists

Candidates

  • Elios Academy
  • Explore Jobs
  • Join the Network

Company

  • About Us
  • Elios Insights
  • Contact

Resources

  • Blog
  • Agent Tools

1โ€œConvertible to full-timeโ€ describes an option that may be available on certain engagements. Any conversion is subject to a separate written agreement, eligibility, and applicable terms; Elios does not guarantee conversion.

2 Source: RAND Corporation, 2024, The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed.

ยฉ 2026 Elios, Inc.PrivacyTerms