BEP Research Research preview

Your AI is wrong about 1 in 16 times.
Which model is wrong least — and what is that costing you?

We test 13 AI models on real work — banking, legal, medical, support and code — and measure the two things nobody publishes: how often each one is actually right, and what the wrong answers cost you.

Price is what you pay for tokens. Cost is what you pay for mistakes.

benchmark 0.2.0 · pricing 2026-07-09 · billed cost · baseline gpt-oss 120B · 32 tasks · computed 2026-07-31 04:01
17 AI models tested1,627 graded testsJuly 2026 last updated
What are you trying to automate?
How much of it, and what does a mistake cost?
Tasks per month
A wrong answer that reaches a customer costs us
Most teams underestimate this. $200 is a reasonable starting point for a customer-facing mistake that needs finding, apologising for and redoing.
Accuracy we require

Workload calculator

The benchmark measures what an attempt costs and how often it is accepted. Everything else that decides your bill is yours: volume, review time, what a wrong answer costs when it reaches a customer. Supply those and the ranking can change — which is the point.

Retries assume a failed task fails again, because that is what we measured — asking a second time usually returns the same wrong answer at twice the price. Why that is, and when it is not true →

Your workload assumption
Human review assumption
Risk and retries assumption

Success rates and cost per attempt are measured on the benchmark tasks shown below. Volume, review time, error cost, detection rate and retry behaviour are assumptions you supplied. Totals combine both and are estimates, not quotes.

The evidence

Everything below is how that recommendation was reached: what each model cost, how often a grader accepted its output, and where the numbers are too thin to act on. The Octane score explains the ranking; it is not the product.

$0.62 per 10,000 completed tasks · gpt-oss 120B · 81.1× cheaper than GLM-5
baseline
observed success of that option
97%
312 of 10,000 still need redoing
weakest vertical
80%
Legal — weakest measured vertical
widest gap in one vertical
132×
how much the choice is worth
kinds of work tested
5
1,627 graded tests

Which model won each kind of work

These cards reflect performance on the specific benchmark tasks listed under “What we test” — they are starting points for a model evaluation, not blanket recommendations for an entire industry. The badge is the observed success rate on those tasks, with the run counts beneath it; not a claim about production reliability. Each card leads with the highest-scoring model above the 80% threshold, then names a lower-cost option that also cleared it. Where no model reached the threshold, the card says so instead of recommending one.

Banking & Finance
100%observed success
18 / 18 benchmark runs
GPT-5.6 Luna
$0.35 per 10,000 successful tasks
Lower-cost option: DeepSeek V3.2 — 17/18 (94%), $0.33, 1.1× cheaper
Customer Service
100%observed success
15 / 15 benchmark runs
gpt-oss 120B
$0.32 per 10,000 successful tasks
Lowest measured cost above the threshold
Legal
100%observed success
15 / 15 benchmark runs
GPT-5.6 Luna Pro
$4.38 per 10,000 successful tasks
Lower-cost option: GPT-5.6 Luna — 14/15 (93%), $0.73, 6.0× cheaper
Medical
100%observed success
27 / 27 benchmark runs
gpt-oss 120B
$0.31 per 10,000 successful tasks
Lowest measured cost above the threshold
Software Engineering
100%observed success
21 / 21 benchmark runs
gpt-oss 120B
$1.14 per 10,000 successful tasks
Lowest measured cost above the threshold

Sticker price is not the price

A pricing page sells you a rate per million tokens. That rate is one of three terms, and it is the only one published. A model can advertise the lowest rate, write two or three times as many tokens to answer the same question, fail a share of the time — and finish as the most expensive option here.

Cost per Successful Task = Total Billed API Cost Across All Attempts ÷ Successful Outputs

what the price page implies →$ per 10,000 completed tasksgpt-oss 120B$0.10$0.54$0.56$0.62×5.3 tokens · ×1.0 failures · ×1.1 work mix → ×6.0 the stickerGPT-5.6 Luna$0.33$0.73$0.73$0.76×2.2 tokens · ×1.0 failures · ×1.0 work mix → ×2.3 the stickerDeepSeek V3.2$0.46$0.73$0.88$1.43×1.6 tokens · ×1.2 failures · ×1.6 work mix → ×3.1 the stickerGemini 2.5 Flash$1.23$2.53$2.83$2.89×2.1 tokens · ×1.1 failures · ×1.0 work mix → ×2.4 the stickerLlama 4 Maverick$0.52$1.37$2.39$3.30×2.7 tokens · ×1.7 failures · ×1.4 work mix → ×6.4 the stickerGPT-5.4 mini$2.46$2.67$3.09$3.36×1.1 tokens · ×1.2 failures · ×1.1 work mix → ×1.4 the stickerGPT-5.6 Luna Pro$0.33$4.86$4.86$4.95×14.8 tokens · ×1.0 failures · ×1.0 work mix → ×15.1 the stickerGPT-5.6 Terra$3.28$6.34$6.54$6.72×1.9 tokens · ×1.0 failures · ×1.0 work mix → ×2.0 the stickerClaude Haiku 4.5$2.93$4.89$6.25$8.19×1.7 tokens · ×1.3 failures · ×1.3 work mix → ×2.8 the stickerGLM-4.7 Flash$0.21$4.58$5.94$10.05×21.7 tokens · ×1.3 failures · ×1.7 work mix → ×47.7 the stickerGLM-5.2$2.82$9.88$10.54$11.78×3.5 tokens · ×1.1 failures · ×1.1 work mix → ×4.2 the stickerQwen3.5 Flash$0.17$8.93$9.74$13.01×53.2 tokens · ×1.1 failures · ×1.3 work mix → ×77.4 the stickerClaude Sonnet 4.6$8.79$16.86$17.59$19.32×1.9 tokens · ×1.0 failures · ×1.1 work mix → ×2.2 the stickerGPT-5.6 Sol$16.39$30.01$31.03$32.44×1.8 tokens · ×1.0 failures · ×1.0 work mix → ×2.0 the stickerClaude Opus 4.6$14.66$31.27$32.63$33.32×2.1 tokens · ×1.0 failures · ×1.0 work mix → ×2.3 the stickerKimi K2.5$1.67$34.91$37.23$45.56×20.9 tokens · ×1.1 failures · ×1.2 work mix → ×27.3 the stickerGLM-5$2.02$29.62$36.00$50.23×14.7 tokens · ×1.2 failures · ×1.4 work mix → ×24.9 the sticker$0$13$26$40$53

Each bar adds one thing the price page does not tell you. The faintest is the advertised rate on this exact workload, priced as if the model were as concise as the leanest one measured. Then the tokens it really wrote. Then the attempts you paid for and could not use. Then your actual work mix, since the index weights every vertical equally instead of letting a model's weak vertical be diluted by its strong one. Only the faintest bar is visible when you pick a model.

Overall

The leading model in each panel sets 100; every other grade is its share of that leader's successful work per dollar. Log axis — a real panel spans orders of magnitude. Whiskers are 95% clustered bootstrap intervals, resampled by vertical, then task, then repetition. Raw baseline-relative scores are in the CSV and JSON exports.

gpt-oss 120BClaude Haiku 4.5Claude Sonnet 4.6Claude Opus 4.6Qwen3.5 FlashLlama 4 MaverickDeepSeek V3.2Kimi K2.5GLM-5GLM-5.2GLM-4.7 FlashGPT-5.4 miniGPT-5.6 LunaGPT-5.6 Luna ProGPT-5.6 TerraGPT-5.6 SolGemini 2.5 Flash
0.101.0010.01001,000100 = scope leadergpt-oss 120B100GPT-5.6 Luna81.5DeepSeek V3.243.4Gemini 2.5 Flash21.5Llama 4 Maverick18.855/96 — below thresholdGPT-5.4 mini18.4GPT-5.6 Luna Pro12.5GPT-5.6 Terra9.23Claude Haiku 4.57.5775/96 — below thresholdGLM-4.7 Flash6.1674/96 — below thresholdGLM-5.25.26Qwen3.5 Flash4.77Claude Sonnet 4.63.21GPT-5.6 Sol1.91Claude Opus 4.61.86Kimi K2.51.36GLM-51.23
Why the overall leader differs from the vertical leaders

gpt-oss 120B ranks first overall because its lead in Customer Service, Medical, Software Engineering outweighs DeepSeek V3.2's lead in Banking & Finance under the benchmark's equal-weight methodology — each vertical counts the same regardless of how many tasks or attempts it holds. Weight the verticals to your own workload in the calculator above and the ranking can change.

Observations
  • The choice of model matters most in Software Engineering, where measured cost per successful task spans 132x.
  • GPT-5.6 Luna Pro completed every benchmark run (96/96).

By vertical

Every panel shares one axis, so they are directly comparable. This is the cut that matters — cost-effectiveness is a property of a model on a kind of work, not of a model alone.

A hollow, struck-through dot marks a model that recorded below 80% observed success on that vertical. It is excluded from the recommendation, not judged unusable — cost per successful task already prices retries, but it assumes a failure is free to detect and that retrying eventually works. Neither held in these measurements: most model-task pairs returned byte-identical text on every repetition, so a failed attempt repeated rather than varied.

Banking & Finance
0.101.0010.01001,000DeepSeek V3.2100GPT-5.6 Luna94.4gpt-oss 120B86.4Gemini 2.5 Flash40.1GPT-5.4 mini23.9Claude Haiku 4.517.0GPT-5.6 Terra13.0GLM-4.7 Flash12.1GPT-5.6 Luna Pro11.0Llama 4 Maverick8.75Qwen3.5 Flash6.57GLM-5.26.29Claude Sonnet 4.63.86GPT-5.6 Sol3.68Kimi K2.52.01GLM-51.89Claude Opus 4.61.07
  • DeepSeek V3.2 recorded the lowest measured cost among models above the threshold, at 17/18 successful · 94% observed.
  • Measured cost per successful task spans 94x across the field, from DeepSeek V3.2 to Claude Opus 4.6.
Customer Service
0.101.0010.01001,000gpt-oss 120B100GPT-5.6 Luna94.9DeepSeek V3.274.4Gemini 2.5 Flash52.9GPT-5.4 mini19.4Llama 4 Maverick17.3GLM-4.7 Flash16.3GPT-5.6 Terra12.1GPT-5.6 Luna Pro11.5GLM-5.27.39Qwen3.5 Flash6.37Kimi K2.52.54GPT-5.6 Sol2.33GLM-52.21Claude Haiku 4.52.16Claude Sonnet 4.61.79Claude Opus 4.61.16
  • gpt-oss 120B recorded the lowest measured cost among models above the threshold, at 15/15 successful · 100% observed.
  • Measured cost per successful task spans 86x across the field, from gpt-oss 120B to Claude Opus 4.6.
Legal
0.101.0010.01001,000GPT-5.6 Luna100gpt-oss 120B77.0Gemini 2.5 Flash25.5GPT-5.6 Luna Pro16.6DeepSeek V3.215.8GPT-5.4 mini14.9Llama 4 Maverick14.0GPT-5.6 Terra12.6GLM-5.210.4Claude Haiku 4.59.42GLM-4.7 Flash4.60Claude Sonnet 4.62.91Claude Opus 4.62.47GPT-5.6 Sol2.19Qwen3.5 Flash1.82GLM-51.41Kimi K2.50.90
  • GPT-5.6 Luna recorded the lowest measured cost among models above the threshold, at 14/15 successful · 93% observed.
  • Measured cost per successful task spans 111x across the field, from GPT-5.6 Luna to Kimi K2.5.
Medical
0.101.0010.01001,000gpt-oss 120B100GPT-5.6 Luna85.9DeepSeek V3.261.0Gemini 2.5 Flash30.4GPT-5.4 mini21.0Qwen3.5 Flash15.1GLM-4.7 Flash13.7GPT-5.6 Luna Pro10.7GPT-5.6 Terra10.4Llama 4 Maverick7.72Claude Haiku 4.57.44GLM-5.27.27Claude Sonnet 4.62.74GPT-5.6 Sol2.17Kimi K2.51.95GLM-51.82Claude Opus 4.61.64
  • gpt-oss 120B recorded the lowest measured cost among models above the threshold, at 27/27 successful · 100% observed.
  • Measured cost per successful task spans 61x across the field, from gpt-oss 120B to Claude Opus 4.6.
Software Engineering
0.101.0010.01001,000gpt-oss 120B100DeepSeek V3.290.5Llama 4 Maverick71.0GPT-5.6 Luna56.2GPT-5.4 mini15.3Gemini 2.5 Flash12.5GPT-5.6 Luna Pro9.74Claude Haiku 4.59.27Qwen3.5 Flash8.85GPT-5.6 Terra5.79GLM-4.7 Flash4.14Claude Sonnet 4.63.37GLM-5.22.99Claude Opus 4.61.92GPT-5.6 Sol1.24Kimi K2.51.11GLM-50.75
  • gpt-oss 120B recorded the lowest measured cost among models above the threshold, at 21/21 successful · 100% observed.
  • Measured cost per successful task spans 132x across the field, from gpt-oss 120B to GLM-5.

The frontier

Observed success against cost per successful task. Bottom-right is the good corner: high observed success at low cost. Position carries the reading, not colour — every point is directly labelled.

$0.00001$0.00010$0.00100$0.01000%25%50%75%100%success rate → bettercost per success ↓ bettergpt-oss 120BGPT-5.6 LunaDeepSeek V3.2Gemini 2.5 FlashLlama 4 MaverickGPT-5.4 miniGPT-5.6 Luna ProGPT-5.6 TerraClaude Haiku 4.5GLM-4.7 FlashGLM-5.2Qwen3.5 FlashClaude Sonnet 4.6GPT-5.6 SolClaude Opus 4.6Kimi K2.5GLM-5

Methodology

Cost per Successful Task = Total Billed API Cost Across All Attempts ÷ Successful Outputs

Total billed cost includes input tokens, output tokens, cached tokens where applicable, reasoning tokens where separately billed, failed attempts, and retries where benchmarked. Work you paid for and could not use stays in the numerator.

Full methodology, grader definitions, confidence rules and limitations →

What we test

Every task behind these numbers, with how it is graded and what it weighs. Full prompts are deliberately unpublished — printed prompts end up in training data and the benchmark decays — but per-task results are in the CSV and JSON exports, so any figure on this page traces to graded attempts.

Banking & Finance 6 tasks

TaskGraded byWeightVersion
Accrue interest under a stated day-count conventionexact answer2v1
Convert a nominal rate to an effective annual ratestructured output, field-checked1.5v1
Refuse to settle an FX conversion at the wrong date's rateexact answer2v1
Assign a KYC risk tierexact answer1v1
Allocate a partial payment down a contractual waterfallstructured output, field-checked1.5v1
Extract structured fields from a bank statement linestructured output, field-checked1v1

Customer Service 5 tasks

TaskGraded byWeightVersion
Apply a return policy where the exception overrides the deadlineexact answer2v1
Apply a multi-condition escalation ruleexact answer1.5v1
Compute a prorated refund on cancellationstructured output, field-checked1.5v1
Decide refund eligibility against a policystructured output, field-checked1v1
Classify a support ticket into a routing intentexact answer1v1

Legal 5 tasks

TaskGraded byWeightVersion
Resolve conflicting governing-law clauses via a precedence rulestructured output, field-checked1.5v1
Extract governing law and venue from a contract clausestructured output, field-checked1v1
Apply a liability cap with a carve-out and an absolute exclusionexact answer2v1
Compute a notice expiry with exclusion and weekend roll-forwardstructured output, field-checked1.5v1
Earliest termination date under an initial-term barexact answer2v1

Medical 9 tasks

TaskGraded byWeightVersion
Apply a drug interaction only when its stated condition holdsexact answer2v1
Identify the ICD-10 code for a described conditionexact answer1v2
Convert a weight-based infusion order into a pump rateexact answer2v1
Extract a drug interaction into structured formstructured output, field-checked1v1
Apply a maximum-dose ceiling before converting to volumeexact answer2v1
Refuse to fill a weight-based order when the weight is missingexact answer2v1
Convert a weight-based dose into a dispensing volumestructured output, field-checked1.5v1
Interpret a lab value against a sex-specific reference rangeexact answer1.5v1
Adjust a dose for renal function using Cockcroft-Gaultexact answer2v1

Software Engineering 7 tasks

TaskGraded byWeightVersion
Add months to a date, clamping to the end of the monthcode executed against tests2v1
Implement order-preserving deduplicationcode executed against tests1.5v1
Predict the output of an order-preserving deduplicationexact answer1v1
Merge overlapping intervals, including touching endpointscode executed against tests1.5v1
Count divisible pairs at a scale that punishes brute forcecode executed against tests2v1
Parse durations, honouring a buried error-handling contractcode executed against tests1.5v1
Subtract one set of half-open intervals from anothercode executed against tests2v1

Benchmark your own workload

This public basket is 23 short, single-shot tasks. Your work is not these tasks. The same machinery runs against representative samples of your own workload — extraction from your documents, your support tickets, your code — and returns model comparison, total-cost analysis under your review and error costs, observed success by task type, routing recommendations, and the caveats that apply.

Useful for model selection, vendor evaluation, cost optimisation, routing design and procurement support. Get in touch to benchmark your workload →

All figures

Task-level rows are omitted here for length. The CSV and JSON exports carry them in full.

ScopeVerticalModelObserved successCost per successful taskCost 95% CIOctaneOctane 95% CITotal billed cost
overallgpt-oss 120B93/96 · 97%$0.000062$0.0000312–$0.00011100$0.00522
overallGPT-5.6 Luna95/96 · 99%$0.00007681.554.1–145$0.00698
overallDeepSeek V3.280/96 · 83%$0.0001443.411.6–120$0.00705
overallGemini 2.5 Flash86/96 · 90%$0.0002921.513.0–55.4$0.0243
overallLlama 4 Maverick55/96 · 57%$0.0003318.85.08–47.6$0.0132
overallGPT-5.4 mini83/96 · 86%$0.0003418.412.1–25.1$0.0256
overallGPT-5.6 Luna Pro96/96 · 100%$0.0005012.58.33–21.0$0.0466
overallGPT-5.6 Terra93/96 · 97%$0.000679.235.59–17.3$0.0608
overallClaude Haiku 4.575/96 · 78%$0.000827.572.73–14.7$0.0469
overallGLM-4.7 Flash74/96 · 77%$0.001016.161.94–15.7$0.0439
overallGLM-5.290/96 · 94%$0.001185.262.24–12.6$0.0949
overallQwen3.5 Flash88/96 · 92%$0.001304.771.14–14.4$0.0857
overallClaude Sonnet 4.692/96 · 96%$0.001933.212.00–4.92$0.1619
overallGPT-5.6 Sol88/91 · 97%$0.003241.911.23–3.43$0.2731
overallClaude Opus 4.692/96 · 96%$0.003331.861.11–3.07$0.3002
overallKimi K2.590/96 · 94%$0.004561.360.95–2.52$0.3351
overallGLM-579/96 · 82%$0.005021.230.41–2.48$0.2844
verticalbankingClaude Haiku 4.518/18 · 100%$0.0002017.011.6–25.1$0.00364
verticalbankingClaude Opus 4.615/18 · 83%$0.003101.070.43–2.08$0.0538
verticalbankingClaude Sonnet 4.618/18 · 100%$0.000863.863.12–5.40$0.0158
verticalbankingDeepSeek V3.217/18 · 94%$0.000033110067.3–167$0.00055
verticalbankingGemini 2.5 Flash18/18 · 100%$0.000082740.123.9–69.4$0.00160
verticalbankingGLM-4.7 Flash18/18 · 100%$0.0002712.18.73–15.5$0.00514
verticalbankingGLM-518/18 · 100%$0.001751.891.24–2.87$0.0326
verticalbankingGLM-5.218/18 · 100%$0.000536.293.92–11.9$0.0102
verticalbankingGPT-5.4 mini18/18 · 100%$0.0001423.917.2–35.1$0.00254
verticalbankingGPT-5.6 Luna18/18 · 100%$0.000035194.474.8–118$0.00064
verticalbankingGPT-5.6 Luna Pro18/18 · 100%$0.0003011.08.51–14.2$0.00547
verticalbankingGPT-5.6 Sol18/18 · 100%$0.000903.682.51–5.71$0.0167
verticalbankingGPT-5.6 Terra18/18 · 100%$0.0002513.011.0–15.2$0.00453
verticalbankinggpt-oss 120B18/18 · 100%$0.0000384$0.0000269–$0.00005486.4$0.00069
verticalbankingKimi K2.518/18 · 100%$0.001642.011.38–2.97$0.0294
verticalbankingLlama 4 Maverick9/18 · 50%$0.000388.751.63–39.0$0.00231
verticalbankingQwen3.5 Flash18/18 · 100%$0.000506.574.24–12.1$0.00950
verticalcodingClaude Haiku 4.521/21 · 100%$0.001229.276.74–12.0$0.0235
verticalcodingClaude Opus 4.621/21 · 100%$0.005921.921.47–2.41$0.1131
verticalcodingClaude Sonnet 4.621/21 · 100%$0.003373.372.55–4.38$0.0641
verticalcodingDeepSeek V3.218/21 · 86%$0.0001390.548.2–149$0.00210
verticalcodingGemini 2.5 Flash21/21 · 100%$0.0009112.58.46–17.9$0.0175
verticalcodingGLM-4.7 Flash10/21 · 48%$0.002744.141.11–8.81$0.0214
verticalcodingGLM-511/21 · 52%$0.01500.750.20–1.75$0.1344
verticalcodingGLM-5.218/21 · 86%$0.003802.991.42–8.35$0.0585
verticalcodingGPT-5.4 mini21/21 · 100%$0.0007415.311.5–19.5$0.0143
verticalcodingGPT-5.6 Luna21/21 · 100%$0.0002056.236.9–86.4$0.00393
verticalcodingGPT-5.6 Luna Pro21/21 · 100%$0.001179.746.01–15.5$0.0231
verticalcodingGPT-5.6 Sol20/20 · 100%$0.009171.240.90–2.37$0.1669
verticalcodingGPT-5.6 Terra21/21 · 100%$0.001965.793.76–12.3$0.0380
verticalcodinggpt-oss 120B21/21 · 100%$0.00011$0.0000594–$0.00017100$0.00215
verticalcodingKimi K2.518/21 · 86%$0.01021.110.83–1.41$0.1608
verticalcodingLlama 4 Maverick15/21 · 71%$0.0001671.034.0–110$0.00232
verticalcodingQwen3.5 Flash19/21 · 90%$0.001288.853.42–21.3$0.0237
verticalcustomer_serviceClaude Haiku 4.56/15 · 40%$0.001482.160.60–8.36$0.00610
verticalcustomer_serviceClaude Opus 4.615/15 · 100%$0.002761.160.80–2.33$0.0392
verticalcustomer_serviceClaude Sonnet 4.614/15 · 93%$0.001781.791.27–2.84$0.0227
verticalcustomer_serviceDeepSeek V3.214/15 · 93%$0.000042974.434.1–154$0.00057
verticalcustomer_serviceGemini 2.5 Flash15/15 · 100%$0.000060452.934.3–69.7$0.00090
verticalcustomer_serviceGLM-4.7 Flash15/15 · 100%$0.0002016.311.0–21.9$0.00284
verticalcustomer_serviceGLM-514/15 · 93%$0.001442.211.36–3.28$0.0198
verticalcustomer_serviceGLM-5.215/15 · 100%$0.000437.394.77–10.7$0.00626
verticalcustomer_serviceGPT-5.4 mini13/15 · 87%$0.0001619.414.9–23.6$0.00207
verticalcustomer_serviceGPT-5.6 Luna15/15 · 100%$0.000033694.974.5–139$0.00049
verticalcustomer_serviceGPT-5.6 Luna Pro15/15 · 100%$0.0002811.58.72–13.4$0.00409
verticalcustomer_serviceGPT-5.6 Sol15/15 · 100%$0.001372.331.99–2.82$0.0198
verticalcustomer_serviceGPT-5.6 Terra15/15 · 100%$0.0002612.110.6–14.1$0.00383
verticalcustomer_servicegpt-oss 120B15/15 · 100%$0.0000319$0.0000186–$0.0000481100$0.00045
verticalcustomer_serviceKimi K2.515/15 · 100%$0.001262.541.92–3.00$0.0183
verticalcustomer_serviceLlama 4 Maverick12/15 · 80%$0.0001917.35.31–40.0$0.00218
verticalcustomer_serviceQwen3.5 Flash15/15 · 100%$0.000506.374.26–14.0$0.00711
verticallegalClaude Haiku 4.56/15 · 40%$0.000779.423.35–15.0$0.00448
verticallegalClaude Opus 4.614/15 · 93%$0.002952.471.40–4.93$0.0444
verticallegalClaude Sonnet 4.612/15 · 80%$0.002502.911.08–6.73$0.0306
verticallegalDeepSeek V3.27/15 · 47%$0.0004615.84.54–88.7$0.00268
verticallegalGemini 2.5 Flash6/15 · 40%$0.0002925.59.21–41.9$0.00169
verticallegalGLM-4.7 Flash5/15 · 33%$0.001584.601.26–8.60$0.00870
verticallegalGLM-59/15 · 60%$0.005151.410.48–2.26$0.0527
verticallegalGLM-5.212/15 · 80%$0.0007010.46.66–14.8$0.00840
verticallegalGPT-5.4 mini7/15 · 47%$0.0004914.96.07–22.3$0.00324
verticallegalGPT-5.6 Luna14/15 · 93%$0.000072810068.1–161$0.00097
verticallegalGPT-5.6 Luna Pro15/15 · 100%$0.0004416.610.2–26.3$0.00634
verticallegalGPT-5.6 Sol11/14 · 79%$0.003322.191.50–3.05$0.0365
verticallegalGPT-5.6 Terra12/15 · 80%$0.0005812.68.31–18.5$0.00674
verticallegalgpt-oss 120B12/15 · 80%$0.0000946$0.0000464–$0.0001777.0$0.00111
verticallegalKimi K2.512/15 · 80%$0.008060.900.55–1.95$0.0869
verticallegalLlama 4 Maverick7/15 · 47%$0.0005214.03.01–44.6$0.00281
verticallegalQwen3.5 Flash9/15 · 60%$0.004011.820.41–14.3$0.0400
verticalmedicalClaude Haiku 4.524/27 · 89%$0.000427.443.88–17.1$0.00919
verticalmedicalClaude Opus 4.627/27 · 100%$0.001921.640.90–3.95$0.0497
verticalmedicalClaude Sonnet 4.627/27 · 100%$0.001152.741.86–4.43$0.0287
verticalmedicalDeepSeek V3.224/27 · 89%$0.000051661.034.2–115$0.00114
verticalmedicalGemini 2.5 Flash26/27 · 96%$0.0001030.414.6–76.8$0.00258
verticalmedicalGLM-4.7 Flash26/27 · 96%$0.0002313.710.3–17.6$0.00587
verticalmedicalGLM-527/27 · 100%$0.001731.821.30–2.45$0.0449
verticalmedicalGLM-5.227/27 · 100%$0.000437.274.71–11.7$0.0116
verticalmedicalGPT-5.4 mini24/27 · 89%$0.0001521.012.1–32.1$0.00349
verticalmedicalGPT-5.6 Luna27/27 · 100%$0.000036685.964.8–113$0.00094
verticalmedicalGPT-5.6 Luna Pro27/27 · 100%$0.0002910.77.62–14.7$0.00765
verticalmedicalGPT-5.6 Sol24/24 · 100%$0.001452.171.68–2.83$0.0331
verticalmedicalGPT-5.6 Terra27/27 · 100%$0.0003010.47.73–13.8$0.00770
verticalmedicalgpt-oss 120B27/27 · 100%$0.0000315$0.0000206–$0.0000474100$0.00081
verticalmedicalKimi K2.527/27 · 100%$0.001621.951.09–3.89$0.0397
verticalmedicalLlama 4 Maverick12/27 · 44%$0.000417.721.54–28.0$0.00354
verticalmedicalQwen3.5 Flash27/27 · 100%$0.0002115.111.6–19.9$0.00545