LiveBench subsets
Honest ranking is the same exam. These tables are WorkChi pins of public LiveBench items, scored with official generators and official scorers. They are not livebench.ai Overall, not Artificial Analysis Intelligence, and the coding table is not site Coding 83.9.
Snapshot 2026-08-17 · methodology · business-task leaderboard
Isolated protocol · n=48
livebench-hs-v1|workchi-livebench-hs-2026|official-gen|temp=yaml:1|effort=default|provider=*
| # | Model | Pass % | Pass |
|---|---|---|---|
| 1 | Claude Opus 5 anthropic/claude-opus-5 | 77.1 | 37/48 |
| = | Claude Sonnet 5 anthropic/claude-sonnet-5 | 77.1 | 37/48 |
| 3 | Grok 4.6 x-ai/grok-4.6 | 72.9 | 35/48 |
| 4 | Gemini 3.6 Flash google/gemini-3.6-flash | 70.8 | 34/48 |
| = | MiniMax M3 minimax/minimax-m3 | 70.8 | 34/48 |
| 6 | fusion:field fusion:field | 68.8 | 33/48 |
| = | Gemma 4 31B google/gemma-4-31b-it | 68.8 | 33/48 |
| = | GPT-5.6 Sol openai/gpt-5.6-sol | 68.8 | 33/48 |
| = | GPT-5.6 Terra openai/gpt-5.6-terra | 68.8 | 33/48 |
| 10 | fusion:lightning fusion:lightning | 66.7 | 32/48 |
| 11 | MiniMax M2.7 minimax/minimax-m2.7 | 62.5 | 30/48 |
| 12 | DeepSeek V4 Flash deepseek/deepseek-v4-flash | 60.4 | 29/48 |
| = | GLM 5.2 z-ai/glm-5.2 | 60.4 | 29/48 |
| 14 | Gemini 3.5 Flash-Lite google/gemini-3.5-flash-lite | 52.1 | 25/48 |
Isolated protocol · n=32
livebench-mathif-v1|workchi-livebench-mathif-2026|official-gen|temp=yaml:1|effort=high|provider=*
| # | Model | Cat-mean | Pass |
|---|---|---|---|
| 1 | Claude Opus 5 anthropic/claude-opus-5 | 90.0 | 27/32 |
| 2 | Gemini 3.6 Flash google/gemini-3.6-flash | 87.6 | 28/32 |
| 3 | Gemini 3.7 Flash google/gemini-3.7-flash | 87.4 | 27/32 |
| 4 | fusion:lightning fusion:lightning | 86.7 | 25/32 |
| 5 | GPT-5.6 Sol openai/gpt-5.6-sol | 83.6 | 25/32 |
| 6 | GPT-5.6 Terra openai/gpt-5.6-terra | 83.5 | 24/32 |
| 7 | Gemini 3.5 Flash-Lite google/gemini-3.5-flash-lite | 82.8 | 25/32 |
| 8 | MiniMax M2.7 minimax/minimax-m2.7 | 81.9 | 25/32 |
| 9 | GLM 5.2 z-ai/glm-5.2 | 81.5 | 25/32 |
| 10 | DeepSeek V4 Flash deepseek/deepseek-v4-flash | 79.0 | 24/32 |
| 11 | Gemma 4 31B google/gemma-4-31b-it | 75.1 | 22/32 |
| 12 | Grok 4.6 x-ai/grok-4.6 | 74.8 | 24/32 |
| 13 | Claude Sonnet 5 anthropic/claude-sonnet-5 | 74.6 | 22/32 |
| 14 | Ling 3.0 Flash inclusionai/ling-3.0-flash | 74.6 | 24/32 |
| 15 | MiniMax M3 minimax/minimax-m3 | 73.8 | 20/32 |
| 16 | fusion:field fusion:field | 69.3 | 18/32 |
Isolated protocol · n=32
livebench-coding-v2|workchi-livebench-coding-2026-v2|official-gen|temp=yaml:1|effort=high|provider=*
| # | Model | Cat-mean | Pass |
|---|---|---|---|
| 1 | Claude Opus 5 anthropic/claude-opus-5 | 100.0 | 32/32 |
| = | Gemini 3.6 Flash google/gemini-3.6-flash | 100.0 | 32/32 |
| = | Gemini 3.7 Flash google/gemini-3.7-flash | 100.0 | 32/32 |
| 4 | fusion:lightning fusion:lightning | 96.9 | 31/32 |
| 5 | GPT-5.6 Sol openai/gpt-5.6-sol | 93.8 | 30/32 |
| 6 | fusion:field fusion:field | 90.6 | 29/32 |
| = | GPT-5.6 Terra openai/gpt-5.6-terra | 90.6 | 29/32 |
| 8 | Grok 4.6 x-ai/grok-4.6 | 87.5 | 28/32 |
| 9 | GLM 5.2 z-ai/glm-5.2 | 84.4 | 27/32 |
| 10 | Claude Sonnet 5 anthropic/claude-sonnet-5 | 78.1 | 25/32 |
| 11 | MiniMax M3 minimax/minimax-m3 | 75.0 | 24/32 |
| 12 | Gemma 4 31B google/gemma-4-31b-it | 71.9 | 23/32 |
| 13 | DeepSeek V4 Flash deepseek/deepseek-v4-flash | 68.8 | 22/32 |
| 14 | Gemini 3.5 Flash-Lite google/gemini-3.5-flash-lite | 62.5 | 20/32 |
| 15 | Ling 3.0 Flash inclusionai/ling-3.0-flash | 53.1 | 17/32 |
| = | MiniMax M2.7 minimax/minimax-m2.7 | 53.1 | 17/32 |
HS-48 is language + reasoning + data analysis (no math, no coding, no IF), default effort. Math+IF and coding-v2 use effort=high, which is still not livebench.ai Max / xHigh. Coding-v2 is a 2024 public slice (removal_date=2025-04-02); Opus and Gemini 3.6/3.7 all score 100 on it. Rank that pair on math+IF, not coding-v2.
Fusion rows are ensembles, not single models. Serving is mixed: MiniMax and xAI first-party when available, everyone else OpenRouter. Do not average a row with a different protocol string.
A separate residual coding pin (n=3, not shown as a ranking table) keeps the leftover items Opus and Gemini 3.7 both fail. On that exam Sol scores 1/3; Opus, Gemini 3.6, and Gemini 3.7 score 0/3. n is too small to quote a ranking. It is not coding-v2 and not site Coding.