Skip to content
SOLVENCY

Independent cost data · 6 measured models · verified 2026-08-21

Token price is not task cost.

Pricing pages compare dollars per million tokens. Solvency ranks AI coding models by cost per solved task — what it costs to get a task finished — with a source and a verification date on every number.

Per input token

11x

DeepSeek V4 Flash vs Claude Opus 5, the top-scoring model

Per output token

19x

cheaper

Per solved task

100x

$0.12 vs $12.01

Spread, cheapest to dearest

38x → 146x

token price → cost per solved task

Source: Artificial Analysis (artificialanalysis.ai) · verified 2026-08-21 · measured on agentic coding tasks; prices from provider pricing pages

Cost per solved task

measured

USD per task that passes · lower is better

Source: Artificial Analysis (artificialanalysis.ai) · verified 2026-08-21

Full leaderboard ↓

Leaderboard

Cost per solved task, ranked

Cost per attempt divided by pass rate. Lower is better. Measured and modelled rows are ranked separately and never averaged together.

How this is computed

Measured · cost per solved task

The benchmark ran the model and observed the cost. No Solvency assumption is inside these figures.

measured · 6 harness + model pairs
#ModelHarnessIndex$ / task$ / solved task
1DeepSeek V4 FlashCodex50$0.06$0.12
2Gemini 3.7 FlashOpencode60$1.27$2.12
3Grok 4.5Grok Build64$2.44$3.81
4GPT-5.6 SolCodex65$6.42$9.88
5Claude Opus 5Claude Code68$8.17$12.01
6Claude Fable 5Claude Code67$11.70$17.46

Source: Artificial Analysis (artificialanalysis.ai) · verified 2026-08-21 · Coding Agent Index v1.4, 326 tasks, 3 attempts each. Rows are harness + model pairs, not bare models. $ / solved task = $ / task ÷ (index ÷ 100).

Modelled · pass rate published, cost estimated

These sources publish no cost. Cost is Solvency's loop model at today's verified prices — an assumption, labelled as one. Not comparable with the measured table.

modelled · heavy tier
#ModelSourcePass rate fromPass$ / solved task
1GPT-5.4SEAL2026-08-2159%$4.19
2Gemini 3.1 Pro (preview)SEAL2026-08-2146%$4.30
3Claude Opus 4.6SEAL2026-08-2152%$8.67
4Claude Opus 4.5SEAL2026-08-2146%$9.81
5GPT-5Aider2025-08-23 · stale88%$1.66
6Gemini 2.5 ProAider2025-06-06 · stale83%$1.76
7o3Aider2025-06-25 · stale81%$1.99
8Claude Sonnet 4Aider2025-05-24 · stale61%$4.40
9GPT-4.1Aider2025-04-14 · stale52%$5.15
10Claude Opus 4Aider2025-05-25 · stale72%$18.75
11o3-proAider2025-06-28 · stale85%$19.08

Sources: Scale SEAL (SWE-bench Pro) and Aider polyglot · verified 2026-08-21 · Aider pass rates predate 2026 and are marked stale. Loop model and frontier-efficiency assumptions are listed in the methodology.

Not ranked — no published pass rate, reported as missing rather than estimated: Claude Sonnet 5, Claude Haiku 4.5, GPT-5.6 Terra, GPT-5.6 Luna, GPT-5.3 Codex, DeepSeek V4 Pro, Grok 4.6, Mistral Medium 3.5.

Highlights

Three ways to rank the same six models

Same measured rows, three axes. The order on the left is the one that predicts your bill; the order on the right is the one pricing pages publish.

Cost per solved task

USD per task that passes · lower is better

Coding Agent Index

Composite score, 0–100 · higher is better

Output price

USD per million output tokens · what pricing pages show

Source: Artificial Analysis (artificialanalysis.ai) · verified 2026-08-21 · index and per-task cost; output prices from each provider's pricing page

Calculator

Your task mix, your volume

Pick a task tier and a monthly volume. Measured rows cannot be moved by the assumptions; modelled rows can, and say so.

Sign in to run your own scenario

Free account. Set your task tier, volume, cache rate and takeover cost; share the result by link.

DeepSeek V4 Flash costs $0.12 per solved task against Claude Opus 5 at $12.01 — 100.1x more for 18 more points of pass rate. Over 200 tasks that difference is $2.4k a month.

Measured

The benchmark ran the model and observed this cost. No Solvency assumption is inside these figures.

#ModelPass $ / solved task$ / month
1 DeepSeek V4 Flash measured 50% $0.12 $24.00
2 Gemini 3.7 Flash measured 60% $2.12 $423
3 Grok 4.5 measured 64% $3.81 $763
4 GPT-5.6 Sol measured 65% $9.88 $2.0k
5 Claude Opus 5 measured 68% $12.01 $2.4k
6 Claude Fable 5 measured 67% $17.46 $3.5k

Modelled

Pass rate published, cost estimated by Solvency's loop model — an assumption.

#ModelPass $ / solved task$ / month
1 GPT-5.4 modelled 59% $1.86 $372
2 Gemini 3.1 Pro (preview) modelled 46% $1.91 $382
3 Claude Opus 4.6 modelled 52% $3.85 $771
4 Claude Opus 4.5 modelled 46% $4.36 $872

Modelled from stale pass rates

Pass rates published before 2026. Cost recomputed at current prices; the pass rate is old.

#ModelPass $ / solved task$ / month
1 GPT-5 stale 88% $0.74 $148
2 Gemini 2.5 Pro stale 83% $0.78 $156
3 o3 stale 81% $0.89 $177
4 GPT-4.1 stale 52% $1.37 $275
5 Claude Sonnet 4 stale 61% $1.96 $392
6 Claude Opus 4 stale 72% $8.33 $1.7k
7 o3-pro stale 85% $8.48 $1.7k
Not shown — no published pass rate, reported as missing rather than estimated: Claude Sonnet 5, Claude Haiku 4.5, GPT-5.6 Terra, GPT-5.6 Luna, GPT-5.3 Codex, DeepSeek V4 Pro, Grok 4.6, Mistral Medium 3.5.

Measured rows carry a cost the benchmark observed, so the tier, cache and efficiency controls cannot move them. Modelled rows are priced by an assumed loop model.

Finding 1

Token price is not task cost

Log scale, normalised to the cheapest model on each axis. Measured rows only, so no Solvency assumption is inside these numbers.

Read in the note
Token price is not task costToken price is not task cost

Cost per solved task = measured cost per task / pass rate. Source: Artificial Analysis (artificialanalysis.ai) — Coding Agent Index v1.4. Prices verified 2026-08-21.

Finding 2

What one index point costs

Coding Agent Index vs cost per solved task. Dashed line: the frontier — nothing cheaper scores higher.

Read in the note
What one index point costsWhat one index point costs

Rows are harness+model pairs, not bare models. Source: Artificial Analysis (artificialanalysis.ai) — Coding Agent Index v1.4. Verified 2026-08-21.

Finding 3

Newest entry, by source

Bar runs from each source’s newest published entry to today. Longer is worse.

Read in the note
Newest entry, by sourceNewest entry, by source

Aider verified against its raw leaderboard YAML. SEAL publishes no update date, so its staleness is unknown, not zero.

Sources

Where every number comes from

In preference order: fewest Solvency assumptions first, then freshness. Benchmark data is cited and linked, never redistributed.

SourceTasksCovers 2026 modelsPublishes costBasisNewest entryVerified
Artificial Analysis Coding Agent Index v1.4326yesyes, measuredmeasured by source2026-08-212026-08-21
Scale SEAL leaderboard - SWE-bench Pro (public)1,865yesnot capturedmodelled by Solvencyunknown2026-08-21
Aider polyglot benchmark225nohistorical onlyhistorical at run date2025-10-032026-08-21

· Prices are verified against each provider's own pricing page; recalled prices are never used. Missing is printed as missing.

Share the finding

The tweet text is generated from the data, so it matches the tables above.

Share on X