AI Model Leaderboard
Every vote is a pairwise match, and a Bradley–Terry regression over all matches estimates each model's strength on an Elo-like scale: the reference model is pinned at 1500, and a 400-point lead means roughly 10× higher odds of being preferred. The small figure below each score is a 95% confidence range. With Style control on, the regression also holds formatting and length constant, so a model can't climb by looking polished or writing more; off ranks on the raw votes.
The same Bradley–Terry regression, but a "win" is holding your attention longer than another anonymous answer in the same turn — with an answer's position in the list and its length always regressed out. With Style control on, formatting is held constant as well. Ratings sit on an Elo-like scale: the reference model is pinned at 1500, a 400-point lead means roughly 10× higher odds, and the small figure below each score is a 95% confidence range.
Last updated Aug 2, 2026
| Rank | Model | Rating | Win Rate | Battles |
|---|---|---|---|---|
| 1 | Claude Opus 4.6 Anthropic | 2002 95% CI 1584 – 2419 | 67.6% | 17 |
| 2 | GPT-5.4 OpenAI | 1886 95% CI 1645 – 2126 | 66.7% | 24 |
| 3 | Claude Sonnet 4.6 Anthropic | 1736 95% CI 1482 – 1991 | 50.0% | 21 |
| 4 | Grok 4.3 xAI | 1695 95% CI 1261 – 2130 | 40.0% | 5 |
| 5 | Claude Haiku 4.5 Anthropic | 1637 95% CI 1406 – 1868 | 47.8% | 23 |
| 6 | Mistral Large 3 Mistral | 1619 95% CI 1374 – 1864 | 53.6% | 14 |
| 7 | Llama 3.3 70B Meta | 1560 95% CI 1308 – 1811 | 32.4% | 17 |
| 8 | DeepSeek V4 Pro DeepSeek | 1528 95% CI 1059 – 1996 | 21.4% | 7 |
| 9 | GPT-5.2 Chat OpenAI | 1500 | 50.0% | 38 |
| 10 | Claude Opus 4.8 Anthropic | 1499 95% CI 1089 – 1909 | 18.2% | 11 |
| 11 | Claude Opus 4.7 Anthropic | 1456 95% CI 978 – 1934 | 33.3% | 6 |
| 12 | GPT-4o OpenAI | 1448 95% CI 1279 – 1618 | 36.4% | 33 |
| 13 | DeepSeek V4 Flash DeepSeek | 988 95% CI 2 – 1975 | 10.0% | 5 |
?