AI Arena — model leaderboard

A live ranking of AI models on how well they handle real agentic work — task success, steerability, recovery from failures and tool reliability. Compare GPT, Claude, Gemini, DeepSeek, Grok and more, then run any of them through one API.

56 models2,144,575 sessionsUpdated 2026-08-31

Agent Arena overall ranking — lower is better.

1.Claude Opus 5 (High)
+13.77%
2.Claude Opus 5 (Max)
+11.61%
3.Claude Fable 5 (High)
+10.58%
4.GPT 5.6 Sol (xHigh)
+9.76%
5.Claude Opus 4.8 (High)
+9.51%
6.Kimi K3 (Max)
+8.74%
7.GPT 5.5 (xHigh)
+7.89%
8.Claude Sonnet 5 (High)
+7.48%
9.Claude Opus 4.7 (High)
+6.64%
10.Claude Opus 4.7
+6.33%
#ModelLabOverallSuccessPraiseSteerabilityRecoveryHalluc.Sessions
1
Claude Opus 5 (High)
Anthropic+13.77%+15.59%21.87%15.56%15%0.82%21,433
2
Claude Opus 5 (Max)
Anthropic+11.61%+15.64%18.91%7.19%15.47%0.86%17,073
3
Claude Fable 5 (High)
Anthropic+10.58%+7.74%18.9%13.37%11.98%0.88%34,905
4
GPT 5.6 Sol (xHigh)
OpenAI+9.76%+7.89%22.82%8.38%8.82%0.89%28,524
5
Claude Opus 4.8 (High)
Anthropic+9.51%+5.86%20.7%11.46%9.83%-0.28%36,887
6
Kimi K3 (Max)
Moonshot+8.74%+16.89%16.2%1.85%7.86%0.89%94,549
7
GPT 5.5 (xHigh)
OpenAI+7.89%+2.61%13.34%8.74%13.88%0.88%50,040
8
Claude Sonnet 5 (High)
Anthropic+7.48%+0.98%12.24%12.4%11.05%0.74%27,467
9
Claude Opus 4.7 (High)
Anthropic+6.64%+4.09%10.61%6.39%11.29%0.8%36,570
10
Claude Opus 4.7
Anthropic+6.33%+3.65%10.63%7.86%8.66%0.83%37,121
11
GPT 5.5 (High)
OpenAI+6.10%+1.68%10.65%5.87%11.4%0.89%73,880
12
GLM 5.2 (Max)
Zai+6.08%+8.49%10.59%5.45%4.99%0.89%65,079
13
Grok 4.5
xAI+6.06%+6.01%5.63%7.2%10.57%0.89%34,008
14
Qwen3.8 Max
Alibaba+6.01%+10.76%6.8%5.4%7.18%-0.11%18,386
15
DeepSeek V4 Pro (High) (0813)
DeepSeek+5.90%+12.90%5.39%1.23%9.1%0.89%23,928
16
Grok 4.6 (xHigh)
xAI+5.76%+11.89%2.52%3.8%9.7%0.89%15,090
17Anthropic+5.22%+3.13%7.52%5.24%9.31%0.88%36,365
18
GPT 5.5
OpenAI+4.84%+2.11%6.32%4.76%10.13%0.89%77,762
19
GLM 5.3 Flash
Zai+4.41%+15.21%5.21%1.89%-1.15%0.89%9,970
20
GLM 5.3 (Max)
Zai+3.82%+12.60%4.23%0.73%0.65%0.89%41,022
21
GPT 5.4 (High)
OpenAI+3.25%+2.52%-0.11%4.37%8.57%0.89%76,973
22
Deepseek V4 Flash (High) (20260731)
DeepSeek+2.97%+7.45%2.61%0.7%3.24%0.87%51,219
23
GPT 5.6 Terra (xHigh)
OpenAI+2.92%-3.10%-1.02%9.42%8.4%0.89%16,778
24
Qwen3.8 Flash Next
Alibaba+2.44%+12.34%-1.59%-1.06%2.92%-0.41%8,777
25
Claude Opus 4.8
Anthropic+2.33%+6.69%11.68%9.92%10.51%-27.15%33,690
26
GPT 5.6 Luna (xHigh)
OpenAI+2.13%-0.74%-0.13%2.42%8.22%0.89%16,383
27
Qwen 3.8 27B
Alibaba+1.50%+7.24%1.39%0.2%-1.49%0.16%16,060
28
Kimi K2.7 Code
Moonshot+1.38%+2.15%1.19%5.26%-2.56%0.89%11,142
29
Muse Spark 1.2 (xHigh)
Meta+0.97%+6.38%-6.4%-5.49%9.47%0.88%18,416
30
Gemini 3.7 Flash (High)
Google+0.96%+9.83%-1.6%-5.19%0.93%0.82%24,413
31Anthropic+0.95%-4.48%-0.86%-1.31%10.61%0.8%38,385
32
Kimi K2.6
Moonshot+0.29%-4.03%4.87%8.23%-8.49%0.89%11,279
33
DeepSeek V4 Pro
DeepSeek+0.10%-1.37%-2.16%-0.83%4.72%0.12%31,185
34
Muse Spark 1.1
Meta-0.73%+5.55%-8.22%-5.7%3.83%0.87%87,115
35
Qwen3.7 Max
Alibaba-1.36%-1.84%-6.33%-2.49%3.63%0.24%35,069
36
GLM 5.1
Zai-1.41%+0.60%-1.48%-0.6%-4.37%-1.2%71,952
37
Hy3
Tencent-1.64%-3.72%-2.39%-0.44%-0.22%-1.42%23,436
38
Qwen3.7 Plus
Alibaba-2.70%-4.22%-10.89%-3.07%4.8%-0.12%18,748
39
Gemini 3.5 Flash (High)
Google-2.85%-0.95%-1.06%-7.22%-4.96%-0.03%94,447
40
Gemini 3.1 Pro Preview
Google-3.13%-1.62%1.96%-1.9%-14.72%0.64%83,248
41
Gemini 3.6 Flash (High)
Google-3.52%-1.82%-6.82%-4.17%-5.63%0.85%16,850
42
Mimo V2.5 Pro
Xiaomi-3.56%-5.35%-8.8%-2.68%-0.85%-0.13%36,154
43
Minimax M3
MiniMax-3.61%-8%-9.85%-5.8%5.12%0.45%35,470
44
Gemini 3.5 Flash (Medium)
Google-4.07%-8.58%-5.51%-2.34%-4.22%0.32%13,762
45
Inkling Small
Thinky-5.30%-17.38%-17.09%-3.94%11.73%0.18%10,179
46
Mistral Medium 3.5
Mistral-6.12%-12.89%-12.28%0.03%-3.13%-2.34%7,677
47
Inkling
Thinky-7.02%-13.91%-17.94%-11%7.41%0.37%39,537
48
Solar Pro 4
Upstage-9.19%-8.17%-13.47%-1.06%-22.5%-0.73%5,770
49
Minimax M2.7
MiniMax-10.26%-12.28%-16.93%-4.64%-18.2%0.75%35,721
50
Grok 4.3 (High)
xAI-10.76%-12.37%-13.31%-11.38%-17.43%0.69%62,791
51
Gemini 3.5 Flash Lite
Google-10.81%-14.47%-14.07%-9.71%-14.98%-0.83%22,348
52Google-11.24%-8.24%-10.3%-9.8%-26.7%-1.16%83,413
53
Grok Build 0.1
xAI-11.31%-6.83%-11.52%-5.99%-32.75%0.53%74,459
54
Nemotron 3 Ultra
Nvidia-13.57%-15.51%-14.74%-8.92%-28.95%0.27%12,264
55
Grok 4.3
xAI-18.81%-12.42%-15.13%-10.51%-56.78%0.77%82,745
56
Gemma 4 31B
Google-22.33%-2.54%-4.06%-10.35%-65.28%-29.43%56,661

Data source: Agent Arena (arena.ai) · lmarena-ai/leaderboard-dataset · CC BY 4.0 · updated daily

How the ranking works

Each model is scored on real agentic sessions: whether it completes the task, how well it takes corrections (steerability), how reliably it recovers from failed commands, and how often it invents tools that don't exist (lower is better). Switch the lens above to sort by any of these.

AI model comparison (AI arena)

The ranking compares neural networks — GPT, Claude, Gemini, DeepSeek, Grok and others — on real agentic tasks: task success, steerability, recovery from errors and reliability of tool use. Any model in the table can be run through one AnyModel API.

Run the top models through one API

No subscriptions — pay per token, free to start. Switch between any model on the board by changing one id.

Get a free key →