LiveBench subsets
Honest ranking is the same exam. These tables are WorkChi pins of public LiveBench items, scored with official generators and official scorers. They are not livebench.ai Overall, not Artificial Analysis Intelligence, and the coding table is not site Coding 83.9.
Snapshot 2026-08-31 · methodology · business-task leaderboard
Isolated protocol · n=48
livebench-hs-v1|workchi-livebench-hs-2026|official-gen|temp=yaml:1|effort=default|provider=*
| Tier | # | Model | Pass % | Pass |
|---|---|---|---|---|
| T1 | 1 | Claude Opus 5 anthropic/claude-opus-5 | 77.1 | 37/48 63–87% |
| T1 | = | Claude Sonnet 5 anthropic/claude-sonnet-5 | 77.1 | 37/48 63–87% |
| T1 | = | Gemini 3.7 Flash google/gemini-3.7-flash | 77.1 | 37/48 63–87% |
| T1 | 4 | fusion:lightning fusion:lightning | 75.0 | 36/48 61–85% |
| T1 | 5 | Grok 4.6 x-ai/grok-4.6 | 72.9 | 35/48 59–83% |
| T1 | 6 | deepseek-v4-flash-0731 deepseek/deepseek-v4-flash-0731 | 70.8 | 34/48 57–82% |
| T1 | = | Gemini 3.6 Flash google/gemini-3.6-flash | 70.8 | 68/96 61–79% · k=2 |
| T1 | 8 | glm-5.3-flash @cf/zai-org/glm-5.3-flash | 68.8 | 33/48 55–80% |
| T1 | = | fusion:field fusion:field | 68.8 | 33/48 55–80% |
| T1 | = | Gemma 4 31B google/gemma-4-31b-it | 68.8 | 33/48 55–80% |
| T1 | = | GPT-5.6 Sol openai/gpt-5.6-sol | 68.8 | 33/48 55–80% |
| T1 | = | GPT-5.6 Terra openai/gpt-5.6-terra | 68.8 | 33/48 55–80% |
| T1 | = | glm-5.3 z-ai/glm-5.3 | 68.8 | 33/48 55–80% |
| T1 | 14 | MiniMax M3 minimax/minimax-m3 | 67.7 | 65/96 58–76% · k=2 |
| T2 | 15 | DeepSeek V4 Flash deepseek/deepseek-v4-flash | 60.4 | 29/48 46–73% |
| T2 | = | GLM 5.2 z-ai/glm-5.2 | 60.4 | 29/48 46–73% |
| T2 | 17 | Gemini 3.5 Flash-Lite google/gemini-3.5-flash-lite | 52.1 | 25/48 38–66% |
| T3 | 18 | Ling 3.0 Flash inclusionai/ling-3.0-flash | 27.1 | 13/48 17–41% |
Isolated protocol · n=32
livebench-mathif-v1|workchi-livebench-mathif-2026|official-gen|temp=yaml:1|effort=high|provider=*
| Tier | # | Model | Cat-mean | Pass |
|---|---|---|---|---|
| T1 | 1 | Gemini 3.7 Flash google/gemini-3.7-flash | 93.0 | 160/182 82–92% · k=2 |
| T1 | 2 | GPT-5.6 Sol openai/gpt-5.6-sol | 92.0 | 157/182 81–91% · k=2 |
| T1 | 3 | Gemini 3.6 Flash google/gemini-3.6-flash | 91.6 | 155/182 79–90% · k=2 |
| T1 | 4 | Claude Opus 5 anthropic/claude-opus-5 | 90.0 | 27/32 68–93% |
| T2 | 5 | GPT-5.6 Terra openai/gpt-5.6-terra | 88.9 | 145/182 73–85% · k=2 |
| T2 | 6 | fusion:lightning fusion:lightning | 86.7 | 25/32 61–89% |
| T2 | 7 | fusion:balanced fusion:balanced | 83.4 | 112/150 67–81% |
| T2 | 8 | Gemini 3.5 Flash-Lite google/gemini-3.5-flash-lite | 82.8 | 25/32 61–89% |
| T2 | 9 | fusion:fast fusion:fast | 81.9 | 110/150 66–80% |
| T2 | 10 | MiniMax M2.7 minimax/minimax-m2.7 | 81.9 | 25/32 61–89% |
| T2 | 11 | GLM 5.2 z-ai/glm-5.2 | 81.5 | 25/32 61–89% |
| T3 | 12 | MiniMax M3 minimax/minimax-m3 | 81.5 | 230/332 64–74% · k=3 |
| T3 | 13 | DeepSeek V4 Flash deepseek/deepseek-v4-flash | 79.0 | 24/32 58–87% |
| T3 | 14 | Gemma 4 31B google/gemma-4-31b-it | 75.1 | 22/32 51–82% |
| T3 | 15 | Grok 4.6 x-ai/grok-4.6 | 74.8 | 24/32 58–87% |
| T3 | 16 | Claude Sonnet 5 anthropic/claude-sonnet-5 | 74.6 | 22/32 51–82% |
| T3 | 17 | Ling 3.0 Flash inclusionai/ling-3.0-flash | 74.6 | 24/32 58–87% |
| T3 | 18 | fusion:field fusion:field | 69.3 | 18/32 39–72% |
HS-48 is language + reasoning + data analysis (no math, no coding, no IF), default effort. Math+IF and coding-v2 use effort=high, which is still not livebench.ai Max / xHigh. Coding-v2 is a 2024 public slice (removal_date=2025-04-02); Opus and Gemini 3.6/3.7 all score 100 on it. Rank that pair on math+IF, not coding-v2.
Fusion rows are ensembles, not single models. Serving is mixed: MiniMax and xAI first-party when available, everyone else OpenRouter. Do not average a row with a different protocol string.
A separate residual coding pin (n=3, not shown as a ranking table) keeps the leftover items Opus and Gemini 3.7 both fail. On that exam Sol scores 1/3; Opus, Gemini 3.6, and Gemini 3.7 score 0/3. n is too small to quote a ranking. It is not coding-v2 and not site Coding.