Solvency research
Research notes
Each note is re-derived from the datasets by the test suite before it is published, so a table in a note can never drift from the data behind it.
Headline finding
DeepSeek V4 Pro is 7.6x cheaper than Claude Fable 5 per input token and 13x per output token, but ▼ 83x cheaper per solved task ($0.21 vs $17.45).
Note 04
published 2026-08-26
Composing the Stack
One workload, three ways to staff it. Give every role to a frontier model and the template month bills $249.00; let the frontier model conduct while a value model types and it bills $36.16 — a 6.9x spread from composition alone.
Read the note →
Share Note 04
Note 01
published 2026-08-26
Cost Per Solved Task
Per-token pricing does not predict what a coding task actually costs. Measured across thirteen current configurations, the gap widens from 42x to 145x.
Read the note →
Share Note 01
Note 02
published 2026-08-26
Same Model, Ten Harnesses
Three benchmarks, ten coding harnesses, one result three times over — hold the model constant, change only the scaffold, and the bill moves 3.8x, 2.8x, and in Solvency's own measurements 9.1x.
Read the note →
Share Note 02
Note 03
published 2026-08-25
What Is a Task
Solvency's calculator asks "how many tasks?" Measured across 60 shipped GitHub repos in six use cases, a solo mobile app ships in dozens of tasks (median 48); a data/ML pipeline needs hundreds (median 479).
Read the note →
Share Note 03