AI Arena — model leaderboard
A live ranking of AI models on how well they handle real agentic work — task success, steerability, recovery from failures and tool reliability. Compare GPT, Claude, Gemini, DeepSeek, Grok and more, then run any of them through one API.
Agent Arena overall ranking — lower is better.
| # | Model | Lab | Overall | Success | Praise | Steerability | Recovery | Halluc. | Sessions |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 (High) | Anthropic | +13.77% | +15.59% | 21.87% | 15.56% | 15% | 0.82% | 21,433 |
| 2 | Claude Opus 5 (Max) | Anthropic | +11.61% | +15.64% | 18.91% | 7.19% | 15.47% | 0.86% | 17,073 |
| 3 | Claude Fable 5 (High) | Anthropic | +10.58% | +7.74% | 18.9% | 13.37% | 11.98% | 0.88% | 34,905 |
| 4 | GPT 5.6 Sol (xHigh) | OpenAI | +9.76% | +7.89% | 22.82% | 8.38% | 8.82% | 0.89% | 28,524 |
| 5 | Claude Opus 4.8 (High) | Anthropic | +9.51% | +5.86% | 20.7% | 11.46% | 9.83% | -0.28% | 36,887 |
| 6 | Kimi K3 (Max) | Moonshot | +8.74% | +16.89% | 16.2% | 1.85% | 7.86% | 0.89% | 94,549 |
| 7 | GPT 5.5 (xHigh) | OpenAI | +7.89% | +2.61% | 13.34% | 8.74% | 13.88% | 0.88% | 50,040 |
| 8 | Claude Sonnet 5 (High) | Anthropic | +7.48% | +0.98% | 12.24% | 12.4% | 11.05% | 0.74% | 27,467 |
| 9 | Claude Opus 4.7 (High) | Anthropic | +6.64% | +4.09% | 10.61% | 6.39% | 11.29% | 0.8% | 36,570 |
| 10 | Claude Opus 4.7 | Anthropic | +6.33% | +3.65% | 10.63% | 7.86% | 8.66% | 0.83% | 37,121 |
| 11 | GPT 5.5 (High) | OpenAI | +6.10% | +1.68% | 10.65% | 5.87% | 11.4% | 0.89% | 73,880 |
| 12 | GLM 5.2 (Max) | Zai | +6.08% | +8.49% | 10.59% | 5.45% | 4.99% | 0.89% | 65,079 |
| 13 | Grok 4.5 | xAI | +6.06% | +6.01% | 5.63% | 7.2% | 10.57% | 0.89% | 34,008 |
| 14 | Qwen3.8 Max | Alibaba | +6.01% | +10.76% | 6.8% | 5.4% | 7.18% | -0.11% | 18,386 |
| 15 | DeepSeek V4 Pro (High) (0813) | DeepSeek | +5.90% | +12.90% | 5.39% | 1.23% | 9.1% | 0.89% | 23,928 |
| 16 | Grok 4.6 (xHigh) | xAI | +5.76% | +11.89% | 2.52% | 3.8% | 9.7% | 0.89% | 15,090 |
| 17 | Anthropic | +5.22% | +3.13% | 7.52% | 5.24% | 9.31% | 0.88% | 36,365 | |
| 18 | GPT 5.5 | OpenAI | +4.84% | +2.11% | 6.32% | 4.76% | 10.13% | 0.89% | 77,762 |
| 19 | GLM 5.3 Flash | Zai | +4.41% | +15.21% | 5.21% | 1.89% | -1.15% | 0.89% | 9,970 |
| 20 | GLM 5.3 (Max) | Zai | +3.82% | +12.60% | 4.23% | 0.73% | 0.65% | 0.89% | 41,022 |
| 21 | GPT 5.4 (High) | OpenAI | +3.25% | +2.52% | -0.11% | 4.37% | 8.57% | 0.89% | 76,973 |
| 22 | Deepseek V4 Flash (High) (20260731) | DeepSeek | +2.97% | +7.45% | 2.61% | 0.7% | 3.24% | 0.87% | 51,219 |
| 23 | GPT 5.6 Terra (xHigh) | OpenAI | +2.92% | -3.10% | -1.02% | 9.42% | 8.4% | 0.89% | 16,778 |
| 24 | Qwen3.8 Flash Next | Alibaba | +2.44% | +12.34% | -1.59% | -1.06% | 2.92% | -0.41% | 8,777 |
| 25 | Claude Opus 4.8 | Anthropic | +2.33% | +6.69% | 11.68% | 9.92% | 10.51% | -27.15% | 33,690 |
| 26 | GPT 5.6 Luna (xHigh) | OpenAI | +2.13% | -0.74% | -0.13% | 2.42% | 8.22% | 0.89% | 16,383 |
| 27 | Qwen 3.8 27B | Alibaba | +1.50% | +7.24% | 1.39% | 0.2% | -1.49% | 0.16% | 16,060 |
| 28 | Kimi K2.7 Code | Moonshot | +1.38% | +2.15% | 1.19% | 5.26% | -2.56% | 0.89% | 11,142 |
| 29 | Muse Spark 1.2 (xHigh) | Meta | +0.97% | +6.38% | -6.4% | -5.49% | 9.47% | 0.88% | 18,416 |
| 30 | Gemini 3.7 Flash (High) | +0.96% | +9.83% | -1.6% | -5.19% | 0.93% | 0.82% | 24,413 | |
| 31 | Anthropic | +0.95% | -4.48% | -0.86% | -1.31% | 10.61% | 0.8% | 38,385 | |
| 32 | Kimi K2.6 | Moonshot | +0.29% | -4.03% | 4.87% | 8.23% | -8.49% | 0.89% | 11,279 |
| 33 | DeepSeek V4 Pro | DeepSeek | +0.10% | -1.37% | -2.16% | -0.83% | 4.72% | 0.12% | 31,185 |
| 34 | Muse Spark 1.1 | Meta | -0.73% | +5.55% | -8.22% | -5.7% | 3.83% | 0.87% | 87,115 |
| 35 | Qwen3.7 Max | Alibaba | -1.36% | -1.84% | -6.33% | -2.49% | 3.63% | 0.24% | 35,069 |
| 36 | GLM 5.1 | Zai | -1.41% | +0.60% | -1.48% | -0.6% | -4.37% | -1.2% | 71,952 |
| 37 | Hy3 | Tencent | -1.64% | -3.72% | -2.39% | -0.44% | -0.22% | -1.42% | 23,436 |
| 38 | Qwen3.7 Plus | Alibaba | -2.70% | -4.22% | -10.89% | -3.07% | 4.8% | -0.12% | 18,748 |
| 39 | Gemini 3.5 Flash (High) | -2.85% | -0.95% | -1.06% | -7.22% | -4.96% | -0.03% | 94,447 | |
| 40 | Gemini 3.1 Pro Preview | -3.13% | -1.62% | 1.96% | -1.9% | -14.72% | 0.64% | 83,248 | |
| 41 | Gemini 3.6 Flash (High) | -3.52% | -1.82% | -6.82% | -4.17% | -5.63% | 0.85% | 16,850 | |
| 42 | Mimo V2.5 Pro | Xiaomi | -3.56% | -5.35% | -8.8% | -2.68% | -0.85% | -0.13% | 36,154 |
| 43 | Minimax M3 | MiniMax | -3.61% | -8% | -9.85% | -5.8% | 5.12% | 0.45% | 35,470 |
| 44 | Gemini 3.5 Flash (Medium) | -4.07% | -8.58% | -5.51% | -2.34% | -4.22% | 0.32% | 13,762 | |
| 45 | Inkling Small | Thinky | -5.30% | -17.38% | -17.09% | -3.94% | 11.73% | 0.18% | 10,179 |
| 46 | Mistral Medium 3.5 | Mistral | -6.12% | -12.89% | -12.28% | 0.03% | -3.13% | -2.34% | 7,677 |
| 47 | Inkling | Thinky | -7.02% | -13.91% | -17.94% | -11% | 7.41% | 0.37% | 39,537 |
| 48 | Solar Pro 4 | Upstage | -9.19% | -8.17% | -13.47% | -1.06% | -22.5% | -0.73% | 5,770 |
| 49 | Minimax M2.7 | MiniMax | -10.26% | -12.28% | -16.93% | -4.64% | -18.2% | 0.75% | 35,721 |
| 50 | Grok 4.3 (High) | xAI | -10.76% | -12.37% | -13.31% | -11.38% | -17.43% | 0.69% | 62,791 |
| 51 | Gemini 3.5 Flash Lite | -10.81% | -14.47% | -14.07% | -9.71% | -14.98% | -0.83% | 22,348 | |
| 52 | -11.24% | -8.24% | -10.3% | -9.8% | -26.7% | -1.16% | 83,413 | ||
| 53 | Grok Build 0.1 | xAI | -11.31% | -6.83% | -11.52% | -5.99% | -32.75% | 0.53% | 74,459 |
| 54 | Nemotron 3 Ultra | Nvidia | -13.57% | -15.51% | -14.74% | -8.92% | -28.95% | 0.27% | 12,264 |
| 55 | Grok 4.3 | xAI | -18.81% | -12.42% | -15.13% | -10.51% | -56.78% | 0.77% | 82,745 |
| 56 | Gemma 4 31B | -22.33% | -2.54% | -4.06% | -10.35% | -65.28% | -29.43% | 56,661 |
Data source: Agent Arena (arena.ai) · lmarena-ai/leaderboard-dataset · CC BY 4.0 · updated daily
How the ranking works
Each model is scored on real agentic sessions: whether it completes the task, how well it takes corrections (steerability), how reliably it recovers from failed commands, and how often it invents tools that don't exist (lower is better). Switch the lens above to sort by any of these.
AI model comparison (AI arena)
The ranking compares neural networks — GPT, Claude, Gemini, DeepSeek, Grok and others — on real agentic tasks: task success, steerability, recovery from errors and reliability of tool use. Any model in the table can be run through one AnyModel API.
Run the top models through one API
No subscriptions — pay per token, free to start. Switch between any model on the board by changing one id.
AnyModel