AI Model Leaderboard

Style control ?
Every vote is a pairwise match, and a Bradley–Terry regression over all matches estimates each model's strength on an Elo-like scale: the reference model is pinned at 1500, and a 400-point lead means roughly 10× higher odds of being preferred. The small figure below each score is a 95% confidence range. With Style control on, the regression also holds formatting and length constant, so a model can't climb by looking polished or writing more; off ranks on the raw votes.
13 models · 415 total votes
Last updated Aug 2, 2026
RankModelRatingWin RateBattles
1
Claude Opus 4.6
Anthropic
2002
95% CI 1584 – 2419
67.6%17
2
GPT-5.4
OpenAI
1886
95% CI 1645 – 2126
66.7%24
3
Claude Sonnet 4.6
Anthropic
1736
95% CI 1482 – 1991
50.0%21
4
Grok 4.3
xAI
1695
95% CI 1261 – 2130
40.0%5
5
Claude Haiku 4.5
Anthropic
1637
95% CI 1406 – 1868
47.8%23
6
Mistral Large 3
Mistral
1619
95% CI 1374 – 1864
53.6%14
7
Llama 3.3 70B
Meta
1560
95% CI 1308 – 1811
32.4%17
8
DeepSeek V4 Pro
DeepSeek
1528
95% CI 1059 – 1996
21.4%7
9
GPT-5.2 Chat
OpenAI
1500
50.0%38
10
Claude Opus 4.8
Anthropic
1499
95% CI 1089 – 1909
18.2%11
11
Claude Opus 4.7
Anthropic
1456
95% CI 978 – 1934
33.3%6
12
GPT-4o
OpenAI
1448
95% CI 1279 – 1618
36.4%33
13
DeepSeek V4 Flash
DeepSeek
988
95% CI 2 – 1975
10.0%5