DataChiBenchmark
Pricing
Loading…
DataChi

The benchmark for real-world AI. Built and hosted in the EU.

Product
AI GatewayLLM APIEU Sovereign GatewayObservabilityCompare modelsModel race
Benchmarks
LeaderboardLiveBenchModelsTasksMethodologyBest AI for…
Resources
DocumentationAPI referenceBlogPricing
Company
AboutEU AICloud ActPrivacyTermsContact
Newsletter

Get notified when new models are added to the leaderboard.

All systems normalEU sovereign
© 2026 DataChi · Made in Europe with ♥
PrivacyTermsCompliance

LiveBench subsets

Quote boards, one exam at a time

Honest ranking is the same exam. These tables are WorkChi pins of public LiveBench items, scored with official generators and official scorers. They are not livebench.ai Overall, not Artificial Analysis Intelligence, and the coding table is not site Coding 83.9.

Snapshot 2026-08-17 · methodology · business-task leaderboard

HS-48Math + instruction followingCoding v2

Isolated protocol · n=48

HS-48

livebench-hs-v1|workchi-livebench-hs-2026|official-gen|temp=yaml:1|effort=default|provider=*

  • Same 48-item LiveBench HS pin, official-gen yaml/floor applied (temp 1, ≥64k). Default effort — not Max/xHigh.
  • Headline is item-pooled pass rate. Language / reasoning / data analysis are category slices, not a second ranking.
  • Not livebench.ai Overall. Ties at n=48 are common; mathif and coding-v2 break the cheap-model cluster.
  • A second t0/4k block exists on disk. It is not this exam and is not shown here.
#ModelPass %PassLanguageReasoningDatap50Cost
1
Claude Opus 5
anthropic/claude-opus-5
77.137/4881.393.856.39819ms$1.52
=
Claude Sonnet 5
anthropic/claude-sonnet-5
77.137/4881.393.856.31890ms$1.44
3
Grok 4.6
x-ai/grok-4.6
72.935/4881.393.843.839.8s$0.10
4
Gemini 3.6 Flash
google/gemini-3.6-flash
70.834/4881.393.837.51605ms$1.08
=
MiniMax M3
minimax/minimax-m3
70.834/4881.387.543.821.9s$0.33
6
fusion:field
fusion:field
68.833/4862.581.362.52105ms$1.40
=
Gemma 4 31B
google/gemma-4-31b-it
68.833/4875.081.350.0228ms$0.02
=
GPT-5.6 Sol
openai/gpt-5.6-sol
68.833/4875.087.543.811.3s$1.07
=
GPT-5.6 Terra
openai/gpt-5.6-terra
68.833/4881.387.537.57401ms$0.26
10
fusion:lightning
fusion:lightning
66.732/4881.368.850.01051ms—
11
MiniMax M2.7
minimax/minimax-m2.7
62.530/4843.893.850.087.1s$0.34
12
DeepSeek V4 Flash
deepseek/deepseek-v4-flash
60.429/4875.056.350.0347ms$0.02
=
GLM 5.2
z-ai/glm-5.2
60.429/4862.575.043.812.2s$0.83
14
Gemini 3.5 Flash-Lite
google/gemini-3.5-flash-lite
52.125/4862.550.043.8435ms$0.12

Isolated protocol · n=32

Math + instruction following

livebench-mathif-v1|workchi-livebench-mathif-2026|official-gen|temp=yaml:1|effort=high|provider=*

  • 32-item pin: mathematics + instruction following. Ranked by cat-mean (task-mean, then category-mean).
  • effort=high is not livebench.ai Max / xHigh. Not official LiveBench Overall.
  • Do not average this table with HS-48 or coding-v2.
#ModelCat-meanPassMathIFp50Cost
1
Claude Opus 5
anthropic/claude-opus-5
90.027/3288.191.91650ms$1.27
2
Gemini 3.6 Flash
google/gemini-3.6-flash
87.628/3275.2100.01467ms$0.85
3
Gemini 3.7 Flash
google/gemini-3.7-flash
87.427/3274.8100.02209ms$0.23
4
fusion:lightning
fusion:lightning
86.725/3273.5100.02479ms$0.60
5
GPT-5.6 Sol
openai/gpt-5.6-sol
83.625/3274.692.6581ms$1.21
6
GPT-5.6 Terra
openai/gpt-5.6-terra
83.524/3274.692.3561ms$0.25
7
Gemini 3.5 Flash-Lite
google/gemini-3.5-flash-lite
82.825/3265.5100.01209ms$0.61
8
MiniMax M2.7
minimax/minimax-m2.7
81.925/3267.696.142.7s$0.27
9
GLM 5.2
z-ai/glm-5.2
81.525/3275.387.8472ms$0.33
10
DeepSeek V4 Flash
deepseek/deepseek-v4-flash
79.024/3262.295.81220ms$0.05
11
Gemma 4 31B
google/gemma-4-31b-it
75.122/3266.683.6422ms$0.14
12
Grok 4.6
x-ai/grok-4.6
74.824/3257.292.353.6s$0.09
13
Claude Sonnet 5
anthropic/claude-sonnet-5
74.622/3273.675.71934ms$1.37
14
Ling 3.0 Flash
inclusionai/ling-3.0-flash
74.624/3249.1100.0600ms$0.04
15
MiniMax M3
minimax/minimax-m3
73.820/3271.975.823.5s$0.25
16
fusion:field
fusion:field
69.318/3258.280.460.4s$2.26

Isolated protocol · n=32

Coding v2

livebench-coding-v2|workchi-livebench-coding-2026-v2|official-gen|temp=yaml:1|effort=high|provider=*

  • 32-item public 2024 pin. Official tests in Docker (runc + caps), not string-match and not gVisor.
  • Every public HF livebench/coding row has livebench_removal_date=2025-04-02. This is not livebench.ai site Coding (83.9).
  • effort=high is not site Max/xHigh. Opus and Gemini 3.6/3.7 saturate this pin (32/32) — the exam cannot rank that pair.
  • Do not pool with coding-v1 or leftover/residual probes.
#ModelCat-meanPassLCBCompletionp50Cost
1
Claude Opus 5
anthropic/claude-opus-5
100.032/32100.0100.01736ms$1.44
=
Gemini 3.6 Flash
google/gemini-3.6-flash
100.032/32100.0100.01440ms$0.57
=
Gemini 3.7 Flash
google/gemini-3.7-flash
100.032/32100.0100.01806ms$0.19
4
fusion:lightning
fusion:lightning
96.931/32100.093.82270ms$0.52
5
GPT-5.6 Sol
openai/gpt-5.6-sol
93.830/3293.893.8494ms$0.67
6
fusion:field
fusion:field
90.629/3293.887.548.9s$2.58
=
GPT-5.6 Terra
openai/gpt-5.6-terra
90.629/3293.887.5489ms$0.18
8
Grok 4.6
x-ai/grok-4.6
87.528/3293.881.345.8s$0.07
9
GLM 5.2
z-ai/glm-5.2
84.427/3287.581.3755ms$1.61
10
Claude Sonnet 5
anthropic/claude-sonnet-5
78.125/3293.862.51899ms$0.42
11
MiniMax M3
minimax/minimax-m3
75.024/3281.368.848.6s$0.27
12
Gemma 4 31B
google/gemma-4-31b-it
71.923/3293.850.0350ms$0.16
13
DeepSeek V4 Flash
deepseek/deepseek-v4-flash
68.822/3293.843.8983ms$0.03
14
Gemini 3.5 Flash-Lite
google/gemini-3.5-flash-lite
62.520/3287.537.51151ms$0.39
15
Ling 3.0 Flash
inclusionai/ling-3.0-flash
53.117/3275.031.3685ms$0.02
=
MiniMax M2.7
minimax/minimax-m2.7
53.117/3281.325.027.4s$0.11

What these numbers are

HS-48 is language + reasoning + data analysis (no math, no coding, no IF), default effort. Math+IF and coding-v2 use effort=high, which is still not livebench.ai Max / xHigh. Coding-v2 is a 2024 public slice (removal_date=2025-04-02); Opus and Gemini 3.6/3.7 all score 100 on it. Rank that pair on math+IF, not coding-v2.

Fusion rows are ensembles, not single models. Serving is mixed: MiniMax and xAI first-party when available, everyone else OpenRouter. Do not average a row with a different protocol string.

A separate residual coding pin (n=3, not shown as a ranking table) keeps the leftover items Opus and Gemini 3.7 both fail. On that exam Sol scores 1/3; Opus, Gemini 3.6, and Gemini 3.7 score 0/3. n is too small to quote a ranking. It is not coding-v2 and not site Coding.