Start at the cheapest model that meets the contract, then prove you need more
The matcher below records RossLabs routing heuristics rather than universal rankings. Context length, tools, modality, privacy, availability, latency, and budget can each override the route it suggests, so the deciding evidence is a task-specific evaluation on your own workload.
Start with the lowest-cost model that meets the task contract. Use task-specific evaluations to decide whether a higher tier improves the result enough to justify its cost. The matcher below records RossLabs routing heuristics, not universal rankings. Context length, tools, modality, privacy, availability, latency, and budget can override its starting route.
Find a model for your task
Pick the task and how open-ended it is. The result is a starting route with a required check, not a benchmark verdict.
build from a clear spec
A lower-cost coding model is a reasonable starting point when the specification and tests are explicit.
Why models differ
Start with measured capability and labeled RossLabs observations, then apply hard constraints such as tools, context, privacy, latency, availability, and cost.
The hardest evaluated task it can pass. Use your acceptance tests to find the lowest-cost model that clears the required bar.
How it works. What it pays attention to. What it reaches for. This is what counts once the work gets hard.
The five models
Open a card for the full picture. Behavior statements are RossLabs observations, not vendor guarantees or benchmark results.
Fable RossLabs observation: favors grounded, inspectable answers
Opus 5 RossLabs observation: maps structure and failure paths
Sonnet Builds exactly what the plan says
Haiku Fast, cheap, and literal
Codex (GPT-5.6) Designed for long coding-agent runs with active steering
Where each family stands as of August 11, 2026 List price per 1M tokens, standard context · all three Codex tiers share 1.05M context and 128K max output
- Fable
claude-fable-5$10 in / $50 out Anthropic's most capable widely released model. - Opus
claude-opus-5$5 in / $25 out Thinking is on by default. 1M context, 128K max output. - Sonnet
claude-sonnet-5$2 in / $10 out Near-Opus quality on coding and agentic work. Introductory rate through 31 Aug 2026; $3 / $15 from 1 Sep. - Haiku
claude-haiku-4-5$1 in / $5 out 200K context. The cheap tier. - Codex — Sol
gpt-5.6-sol$5 in / $30 out Top tier. The one to reach for on the hardest runs. - Codex — Terra
gpt-5.6-terra$2 in / $12 out Balanced agentic coding tier. Long-context requests use higher rates. - Codex — Luna
gpt-5.6-luna$0.20 in / $1.20 out Cost-sensitive agentic coding tier. Long-context requests use higher rates.
Per-token rates are not directly comparable across the Claude rows: Haiku 4.5 uses the older tokenizer, while the 5-series produces roughly 30% more tokens for the same text. The cards above describe observations that may change with a version bump. This table also changes — check the source before you budget against it: Anthropic pricing · OpenAI pricing.
Should you use one model for everything?
No. Even if the best model were also the cheapest. Three reasons.
- 1 Style does not rank cleanly
Model behavior varies by task, prompt, tools, and version. No single style should be treated as a universal winner without comparative evaluation.
- 2 Most work does not need a top model
When a task has a clear contract and strong checks, a lower-cost model may meet the requirement. Measure the pass rate before paying for a higher tier.
- 3 An outside opinion is worth keeping
A model from a different provider can supply an independent implementation and training lineage. Disagreement is a reason to inspect the evidence, not proof that either model is correct.
A way to think about the next model
When a new model lands, you do not have to ask if it is the best. Ask four things.
What is the hardest task it gets right? Everything below that line is a job for the cheapest model that also clears it.
What does it pay attention to, and where does that style slip? The strength and the weakness are the same habit.
Most of your work is everyday building. Price it on the common job, since that is where the volume is.
A different source buys you an independent check the others cannot give.