Claude Code Model & Effort Pairing: Real-World Test Across 5 Tasks, 4 Models, 5 Effort Levels

The author benchmarks four Claude Code models (Haiku 4.5, Sonnet 5.5, Opus 5.5, Fable 5.1) across five effort tiers on five typical coding tasks, revealing that Sonnet, Opus, and Fable succeed at low effort while higher tiers only add verification and cost; Haiku misses boundaries that Sonnet low catches; Fable at low effort modifies out-of-scope code; switching effort preserves cache but switching models forces full recomputation.

Tech Ocean
Tech Ocean
Tech Ocean
Claude Code Model & Effort Pairing: Real-World Test Across 5 Tasks, 4 Models, 5 Effort Levels

Two Settings: Model and Effort

Claude Code uses /model to select the model and /effort to set the effort level (low, medium, high, xhigh, max). The model determines capability ceiling and per-token price; the effort level controls how much work the model does per turn — how many files it reads, how many verifications it runs, how far it pushes before asking. Default effort levels: Opus 5.5 and Sonnet 5.5 default to medium; Fable 5.1 defaults to high. The same effort name means different compute budgets across models.

Test Methodology

Five small Node.js repositories, each with an ISSUE.md describing a typical task:

Rename a function (add parameter, update callers and tests).

Fix a CSV import that fails on quoted fields and Windows line endings.

Fix a month-end renewal date bug (rules in a separate document, service runs in multiple time zones).

Fix overselling in a flash-sale endpoint.

Optimize a slow deduplication function.

Hidden tests (not visible to the model) were run after each attempt. Scoring principle: cases in the same code block that conventionally belong together count (e.g., month-end preservation, Windows line endings in CSV); rules in other files not mentioned in the issue do not count. Each run used a clean repo copy, non-interactive mode, no user CLAUDE.md, skills, plugins, MCP, or memory. Models tested: Haiku 4.5, Sonnet 5.5, Opus 5.5, Fable 5.1. Fable costs 2.5× Opus 5.5; costs shown are Claude Code's official estimates for relative comparison.

Key Findings

1. Sonnet 5.5, Opus 5.5, Fable 5.1 pass at low effort

All three models scored full marks on all five tasks at low effort. Raising effort did not improve scores; it only added verification steps, time, and cost. Across the five tasks, max effort cost 4–10× low effort. Opus 5.5 at max even lost one point on the CSV task because it deliberately excluded Windows line endings from scope.

2. Fable 5.1 low effort already does extensive verification

On the renewal task, Fable at low effort ran tests in three time zones (Beijing, US West, UTC) — behavior the docs note is typical for Fable. Yet its low-effort average cost was ~3× Opus 5.5 low.

3. Haiku 4.5 misses boundaries that Sonnet 5.5 low catches

Haiku failed two tasks:

CSV import: it split on newlines first, breaking fields that contain embedded newlines (which the exporter quotes). The other three models parsed by quote rules at low effort.

Renewal date: it never opened the billing rules document, so it missed the "month-end stays month-end" rule and used local-time methods that passed in Beijing but failed in US West.

Haiku has no effort knob; the only fix is switching models. Sonnet 5.5 low effort passes both at comparable estimated cost.

4. Fable at low/medium effort modifies out-of-scope code

On the renewal task, a fourth rule ("expired members renew from today") lived in another file. Sonnet and Opus at all efforts only mentioned it; Fable low edited it both runs, medium sometimes, high/max never. On the flash-sale task, Fable medium/high added validation to the existing decrement interface (disallowing zero/negative) unprompted. At xhigh/max it stopped. The author categorizes out-of-scope edits into three types: adjacent logic in same block (all models handle), unspecified behavior like unknown currency (Sonnet decides only at xhigh/max, Opus/Fable from low), and edits to unrelated files (mostly Fable at low–high). Official docs say higher effort makes the model take more initiative; here the pattern is inverted for unrelated-file edits.

5. Switching effort preserves cache; switching models invalidates it

In a single session, changing effort (e.g., Opus 5.5 low → high) reused the cached conversation (17k tokens). Changing model (Opus 5.5 → Sonnet 5.5) forced a full recompute (~18k tokens). Docs confirm: Opus 5.5, Sonnet 5.5, Fable 5.1 keep cache on effort change when using API key or subscription; other models or cloud platforms (Bedrock, Google Cloud) lose cache on effort change. Every model has its own cache; entering/exiting plan mode ( opusplan) counts as a model switch. Longer sessions make model switches more expensive.

Practical Configuration Tips

/effort

without args opens a slider; /effort high sets and saves as default for that model. Press s on the slider for a one-off change. /model lets you pick model and adjust effort with arrow keys. claude --model opus --effort high applies only to that launch.

Fable must be selected explicitly ( /model fable or claude --model fable); some plans charge extra quota. max is session-only; not persistable in settings.

Legacy top-level effortLevel in user settings does not affect Opus 5.5+.

Sub-agents inherit the parent session's model/effort unless their own file specifies model and effort. ultrathink in the prompt deepens thinking for that turn without changing effort; think, think hard, etc. are ignored. /model opusplan uses Opus for planning, then switches to Sonnet for execution.

Recommended Defaults per Model

Sonnet 5.5 (default medium): low/medium/high similar cost and score; xhigh doubles cost, max ~10× low with no score gain. Daily use: default medium.

Opus 5.5 (default medium): low already perfect; medium/high +30–40% cost; max ~8× low and lost a point on CSV. Daily use: default medium (low also works).

Fable 5.1 (default high): official guidance says start at high, push to xhigh/max for hardest agentic tasks, drop to medium/low if quality holds. Here low scored full marks with three-zone verification; high +50% cost, xhigh ~3×, max ~4×. Fable low cost sits between Opus xhigh and max. For daily work, start at low; clarify scope in prompt to avoid extra edits.

Haiku 4.5 : no effort levels. Fine for mechanical refactors; misses hidden boundaries — switch model to fix.

Decision Framework: When Results Are Wrong

The author maps failure layers (from Lydia Hallie's blog):

Prompt ambiguity (e.g., whether to edit out-of-scope rules).

Insufficient effort (not reading enough files, not verifying enough).

Model capability limit (doesn't know how).

Model read/verified but still failed (not observed in this test).

Practical takeaway: pick model at session start (Sonnet 5.5 default for daily work), adjust effort per task; use Fable low for daily work, reserve higher Fable tiers for large, ambiguous tasks; put scope constraints in the prompt rather than relying on effort level to control out-of-scope edits.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

model comparisonBenchmarkingAI coding assistantClaude CodeHaiku 4.5Sonnet 5.5Fable 5.1Opus 5.5cache behavioreffort levels
Tech Ocean
Written by

Tech Ocean

Focused on AI programming, sharing ready-to-use development efficiency solutions.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.