DataChiBenchmark
Pricing
Loading…
DataChi

The benchmark for real-world AI. Built and hosted in the EU.

Product
AI GatewayLLM APIEU Sovereign GatewayObservabilityCompare modelsModel race
Benchmarks
LeaderboardLiveBenchModelsTasksMethodologyBest AI for…
Resources
DocumentationAPI referenceBlogPricing
Company
AboutEU AICloud ActPrivacyTermsContact
Newsletter

Get notified when new models are added to the leaderboard.

All systems normalEU sovereign
© 2026 DataChi · Made in Europe with ♥
PrivacyTermsCompliance

LiveBench subsets

Quote boards, one exam at a time

Honest ranking is the same exam. These tables are WorkChi pins of public LiveBench items, scored with official generators and official scorers. They are not livebench.ai Overall, not Artificial Analysis Intelligence, and the coding table is not site Coding 83.9.

Snapshot 2026-08-31 · methodology · business-task leaderboard

HS-48Math + instruction following

Isolated protocol · n=48

HS-48

livebench-hs-v1|workchi-livebench-hs-2026|official-gen|temp=yaml:1|effort=default|provider=*

  • Same 48-item LiveBench HS pin, official-gen yaml/floor applied (temp 1, ≥32k; yaml where shipped, floor otherwise). Default effort — not Max/xHigh.
  • Headline is item-pooled pass rate. Language / reasoning / data analysis are category slices, not a second ranking.
  • Not livebench.ai Overall. Ties at n=48 are common — read the tier, and use mathif to separate the cheap-model cluster.
  • Read `tier`, not the row order: rows sharing a tier are not separable on this pin (exact paired sign test over the shared items, alpha 0.05). At n=48 one standard error is ~7pp.
  • `ci95` is Wilson on the pooled pass rate; `runCount` is how many replicate seats were pooled and `spreadPp` their max-min. Repeat seats pool item-wise — the board never publishes the best of k.
  • A second t0/4k block exists on disk. It is not this exam and is not shown here.
  • A model id without a date or version is a floating alias and can move under you: `deepseek/deepseek-v4-flash` and the pinned `deepseek/deepseek-v4-flash-0731` sit in different tiers on identical items, 60.4% against 70.8%. Prefer the pinned id when quoting a row.
  • Rows that lost more than 10% of their cells to the transport (DNS, dropped sockets, provider timeouts) are off this board — a lost cell scores zero, so they understate the model. A model's own empty or truncated answers still count against it.
Tier#ModelPass %PassLanguageReasoningDatap50Cost
T11
Claude Opus 5
anthropic/claude-opus-5
77.1
37/48
63–87%
81.393.856.39819ms$1.52
T1=
Claude Sonnet 5
anthropic/claude-sonnet-5
77.1
37/48
63–87%
81.393.856.31890ms$1.44
T1=
Gemini 3.7 Flash
google/gemini-3.7-flash
77.1
37/48
63–87%
81.393.856.32132ms$0.26
T14
fusion:lightning
fusion:lightning
75.0
36/48
61–85%
75.093.856.32988ms$0.35
T15
Grok 4.6
x-ai/grok-4.6
72.9
35/48
59–83%
81.393.843.839.8s$0.10
T16
deepseek-v4-flash-0731
deepseek/deepseek-v4-flash-0731
70.8
34/48
57–82%
75.093.843.8569ms$0.13
T1=
Gemini 3.6 Flash
google/gemini-3.6-flash
70.8
68/96
61–79% · k=2
78.193.840.61512ms$2.17
T18
glm-5.3-flash
@cf/zai-org/glm-5.3-flash
68.8
33/48
55–80%
68.887.550.027.9s$0.06
T1=
fusion:field
fusion:field
68.8
33/48
55–80%
62.581.362.52105ms$1.40
T1=
Gemma 4 31B
google/gemma-4-31b-it
68.8
33/48
55–80%
75.081.350.0228ms$0.02
T1=
GPT-5.6 Sol
openai/gpt-5.6-sol
68.8
33/48
55–80%
75.087.543.811.3s$1.07
T1=
GPT-5.6 Terra
openai/gpt-5.6-terra
68.8
33/48
55–80%
81.387.537.57401ms$0.26
T1=
glm-5.3
z-ai/glm-5.3
68.8
33/48
55–80%
68.887.550.02182ms$1.40
T114
MiniMax M3
minimax/minimax-m3
67.7
65/96
58–76% · k=2
71.987.543.822.9s$0.58
T215
DeepSeek V4 Flash
deepseek/deepseek-v4-flash
60.4
29/48
46–73%
75.056.350.0347ms$0.02
T2=
GLM 5.2
z-ai/glm-5.2
60.4
29/48
46–73%
62.575.043.812.2s$0.83
T217
Gemini 3.5 Flash-Lite
google/gemini-3.5-flash-lite
52.1
25/48
38–66%
62.550.043.8435ms$0.12
T318
Ling 3.0 Flash
inclusionai/ling-3.0-flash
27.1
13/48
17–41%
12.518.850.0507ms$0.03

Isolated protocol · n=32

Math + instruction following

livebench-mathif-v1|workchi-livebench-mathif-2026|official-gen|temp=yaml:1|effort=high|provider=*

  • 32-item pin: mathematics + instruction following. Ranked by cat-mean (task-mean, then category-mean).
  • effort=high is not livebench.ai Max / xHigh. Not official LiveBench Overall.
  • Do not average this table with HS-48.
  • Rows that lost more than 10% of their cells to the transport (DNS, dropped sockets, provider timeouts) are off this board — a lost cell scores zero, so they understate the model. A model's own empty or truncated answers still count against it.
  • Read `tier`, not the row order. `ci95` is Wilson on pass rate; repeat seats of one serving stack are pooled item-wise, not best-of-k.
Tier#ModelCat-meanPassMathIFp50Cost
T11
Gemini 3.7 Flash
google/gemini-3.7-flash
93.0
160/182
82–92% · k=2
90.095.92241ms$2.23
T12
GPT-5.6 Sol
openai/gpt-5.6-sol
92.0
157/182
81–91% · k=2
90.393.7614ms$2.89
T13
Gemini 3.6 Flash
google/gemini-3.6-flash
91.6
155/182
79–90% · k=2
90.992.41558ms$4.04
T14
Claude Opus 5
anthropic/claude-opus-5
90.0
27/32
68–93%
88.191.91650ms$1.27
T25
GPT-5.6 Terra
openai/gpt-5.6-terra
88.9
145/182
73–85% · k=2
90.587.2598ms$2.20
T26
fusion:lightning
fusion:lightning
86.7
25/32
61–89%
73.5100.02479ms$0.60
T27
fusion:balanced
fusion:balanced
83.4
112/150
67–81%
81.585.417.3s$3.33
T28
Gemini 3.5 Flash-Lite
google/gemini-3.5-flash-lite
82.8
25/32
61–89%
65.5100.01209ms$0.61
T29
fusion:fast
fusion:fast
81.9
110/150
66–80%
74.089.81619ms$1.65
T210
MiniMax M2.7
minimax/minimax-m2.7
81.9
25/32
61–89%
67.696.142.7s$0.27
T211
GLM 5.2
z-ai/glm-5.2
81.5
25/32
61–89%
75.387.8472ms$0.33
T312
MiniMax M3
minimax/minimax-m3
81.5
230/332
64–74% · k=3
86.776.312.6s$2.17
T313
DeepSeek V4 Flash
deepseek/deepseek-v4-flash
79.0
24/32
58–87%
62.295.81220ms$0.05
T314
Gemma 4 31B
google/gemma-4-31b-it
75.1
22/32
51–82%
66.683.6422ms$0.14
T315
Grok 4.6
x-ai/grok-4.6
74.8
24/32
58–87%
57.292.353.6s$0.09
T316
Claude Sonnet 5
anthropic/claude-sonnet-5
74.6
22/32
51–82%
73.675.71934ms$1.37
T317
Ling 3.0 Flash
inclusionai/ling-3.0-flash
74.6
24/32
58–87%
49.1100.0600ms$0.04
T318
fusion:field
fusion:field
69.3
18/32
39–72%
58.280.460.4s$2.26

What these numbers are

HS-48 is language + reasoning + data analysis (no math, no coding, no IF), default effort. Math+IF and coding-v2 use effort=high, which is still not livebench.ai Max / xHigh. Coding-v2 is a 2024 public slice (removal_date=2025-04-02); Opus and Gemini 3.6/3.7 all score 100 on it. Rank that pair on math+IF, not coding-v2.

Fusion rows are ensembles, not single models. Serving is mixed: MiniMax and xAI first-party when available, everyone else OpenRouter. Do not average a row with a different protocol string.

A separate residual coding pin (n=3, not shown as a ranking table) keeps the leftover items Opus and Gemini 3.7 both fail. On that exam Sol scores 1/3; Opus, Gemini 3.6, and Gemini 3.7 score 0/3. n is too small to quote a ranking. It is not coding-v2 and not site Coding.