LLM Leaderboard 2026: AI Model Rankings
All 540 models ranked by Artificial Analysis Intelligence
Index, refreshed daily. Reasoning-effort tiers (low/medium/high/…) of the same model are
merged into one row showing its highest-scoring tier — see the +N tiers badge for the
full spread. Ranked #1 is currently the top scorer in our dataset; scroll for the full table or
sort by any column.
| Rank | Model | Creator | Released | Intelligence | Coding | Math | Tok/s | Blended $/1M |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)NEW+4 tiers | Anthropic | 2026-09-01 | 53.4 | 81.6 | — | 72 | 20.00 |
| 2 | GPT-6 Astra (max)NEW+5 tiers | OpenAI | 2026-09-03 | 52.7 | 76.9 | — | 66 | 20.00 |
| 3 | Claude Opus 5 (Adaptive Reasoning, Max Effort)+4 tiers | Anthropic | 2026-07-24 | 50.8 | 78.0 | — | 57 | 10.00 |
| 4 | Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) | Anthropic | 2026-06-09 | 49.6 | 76.5 | — | 0 | 20.00 |
| 5 | Muse Spark 1.3 (max)NEW+1 tiers | Meta | 2026-09-02 | 48.1 | 75.8 | — | 290 | 2.00 |
| 6 | GPT-5.6 Sol (max)+5 tiers | OpenAI | 2026-07-09 | 47.0 | 77.4 | — | 75 | 8.00 |
| 7 | Qwen3.8 Max (0902)NEW | Alibaba | 2026-09-02 | 45.4 | 76.2 | — | 41 | 3.00 |
| 8 | GLM-5.3 (max) | Z AI | 2026-08-18 | 44.8 | 74.8 | — | 73 | 2.15 |
| 9 | Grok 4.6 (high)+3 tiers | SpaceXAI | 2026-08-12 | 44.3 | 76.8 | — | 70 | 3.00 |
| 10 | Step 5 PreviewNEW | StepFun | 2026-09-18 | 43.7 | — | — | 88 | 1.43 |
| 11 | Kimi K3 (max)+1 tiers | Kimi | 2026-07-16 | 43.6 | 76.2 | — | 40 | 6.00 |
| 12 | GPT-5.6 Terra (max)+5 tiers | OpenAI | 2026-07-09 | 42.1 | 76.7 | — | 107 | 4.50 |
| 13 | Claude Opus 4.8 (Adaptive Reasoning, Max Effort) | Anthropic | 2026-05-28 | 41.8 | 74.3 | — | 0 | 10.00 |
| 14 | GLM 5.3 FlashNEW | Z AI | 2026-08-26 | 41.8 | 71.5 | — | 90 | 0.24 |
| 15 | Gemini 3.8 Flash (high)NEW+2 tiers | 2026-09-02 | 40.9 | 76.3 | — | 354 | 1.50 | |
| 16 | Claude Opus 4.7 (Adaptive Reasoning, Max Effort)+1 tiers | Anthropic | 2026-04-16 | 40.7 | 73.6 | — | 0 | 10.00 |
| 17 | Qwen3.8 Max | Alibaba | 2026-08-03 | 40.2 | 71.8 | — | 0 | 3.00 |
| 18 | Qwen3.8 2.4T A95B | Alibaba | 2026-08-12 | 39.9 | 71.9 | — | 41 | 3.00 |
| 19 | Qwen3.8-Flash-NextNEW | Alibaba | 2026-08-26 | 39.8 | 73.1 | — | 54 | 0.23 |
| 20 | Gemini 3.7 Flash (medium)+2 tiers | 2026-08-13 | 39.6 | 71.5 | — | 0 | 1.50 | |
| 21 | Muse Spark 1.2 (xhigh) | Meta | 2026-08-05 | 39.6 | 72.2 | — | 0 | 2.00 |
| 22 | DeepSeek V4.1 Flash (Reasoning, Max Effort)NEW | DeepSeek | 2026-09-10 | 39.5 | — | — | 231 | 0.53 |
| 23 | GPT-5.4 (xhigh)+2 tiers | OpenAI | 2026-03-05 | 39.0 | 71.1 | — | 0 | 5.62 |
| 24 | Grok 4.5 (high) | SpaceXAI | 2026-07-08 | 38.8 | 72.4 | — | 0 | 3.00 |
| 25 | GPT-5.5 (xhigh)+4 tiers | OpenAI | 2026-04-23 | 38.4 | 74.9 | — | 0 | 11.25 |
| 26 | Claude Sonnet 5 (Adaptive Reasoning, Max Effort)+5 tiers | Anthropic | 2026-06-30 | 38.2 | 71.5 | — | 87 | 4.00 |
| 27 | GPT-5.6 Luna (max)+5 tiers | OpenAI | 2026-07-09 | 37.3 | 71.4 | — | 154 | 0.45 |
| 28 | DeepSeek V4 Pro 0813 (Reasoning, Max Effort) | DeepSeek | 2026-08-13 | 36.0 | 68.8 | — | 94 | 1.98 |
| 29 | Agnes 3.0 FlashNEW | Sapiens AI | 2026-09-11 | 35.5 | — | — | 0 | 0.07 |
| 30 | Agnes 2.5 Pro BetaNEW | Sapiens AI | 2026-08-26 | 35.2 | 62.3 | — | 0 | 0.15 |
| 31 | DeepSeek V4 Flash Vision (Reasoning, Max Effort) | DeepSeek | 2026-08-21 | 34.8 | 65.0 | — | 237 | 0.66 |
| 32 | DeepSeek V4 Flash 0731 (Reasoning, Max Effort) | DeepSeek | 2026-07-31 | 34.3 | 69.1 | — | 0 | 0.66 |
| 33 | Gemini 3.6 Flash (high) | 2026-07-21 | 34.0 | 69.2 | — | 0 | 1.50 | |
| 34 | GLM-5.2 (max)+1 tiers | Z AI | 2026-06-16 | 33.7 | 68.8 | — | 0 | 2.15 |
| 35 | Muse Spark 1.1 (xhigh) | Meta | 2026-07-09 | 33.7 | 71.3 | — | 0 | 2.00 |
| 36 | Qwen3.8 27B (xhigh)+3 tiers | Alibaba | 2026-08-14 | 33.7 | 68.1 | — | 47 | 1.12 |
| 37 | Gemini 3.5 Flash (medium)+2 tiers | 2026-05-19 | 33.6 | — | — | 0 | 3.38 | |
| 38 | Motif 3 | Motif Technologies | 2026-08-12 | 33.6 | 63.5 | — | 0 | — |
| 39 | GPT-5.3 Codex (xhigh) | OpenAI | 2026-02-05 | 32.5 | — | — | 167 | 4.81 |
| 40 | Motif 3 (Beta) | Motif Technologies | 2026-07-14 | 32.3 | 62.0 | — | 0 | — |
| 41 | Claude Opus 4.6 (Adaptive Reasoning, Max Effort) | Anthropic | 2026-02-05 | 31.9 | — | — | 0 | 10.00 |
| 42 | Muse Spark | Meta | 2026-04-08 | 31.3 | 58.6 | — | 0 | — |
| 43 | K2 Horizon 375B A23BNEW | Institute of Foundation Models | 2026-09-03 | 30.5 | 61.5 | — | 0 | — |
| 44 | Apodex 1.1NEW | Apodex | 2026-08-30 | 30.4 | 60.8 | — | 0 | 0.97 |
| 45 | DeepSeek V4 Pro 0424 (Reasoning, Max Effort)+2 tiers | DeepSeek | 2026-04-24 | 30.4 | 59.4 | — | 0 | 0.54 |
| 46 | GPT-5.2 (xhigh)+2 tiers | OpenAI | 2025-12-11 | 30.4 | — | 99.0 | 0 | 4.81 |
| 47 | Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) | Anthropic | 2026-02-17 | 30.1 | 63.0 | — | 0 | 6.00 |
| 48 | Gemini 3.1 Pro Preview | 2026-02-19 | 29.7 | 68.8 | — | 137 | 4.50 | |
| 49 | Qwen3.7 Max | Alibaba | 2026-05-19 | 29.5 | 66.0 | — | 0 | 3.75 |
| 50 | MiniMax-M3 | MiniMax | 2026-06-01 | 29.2 | 58.6 | — | 129 | 0.53 |
| 51 | Claude Opus 4.5 (Reasoning) | Anthropic | 2025-11-24 | 29.1 | — | 91.3 | 0 | 10.00 |
| 52 | MiMo-V2-Pro | Xiaomi | 2026-03-18 | 28.6 | — | — | 0 | — |
| 53 | GPT-5.2 Codex (xhigh) | OpenAI | 2025-12-11 | 28.5 | — | — | 0 | 4.81 |
| 54 | Qwen3.6 Max Preview | Alibaba | 2026-04-20 | 28.4 | — | — | 0 | 2.92 |
| 55 | Nex-N2-Pro (based on Qwen3.5-397B-A17B) | Nex AGI | 2026-06-02 | 28.2 | 59.1 | — | 0 | — |
| 56 | Solar Pro 4 | Upstage | 2026-08-06 | 28.2 | 52.7 | — | 55 | 0.53 |
| 57 | Gemini 3 Pro Preview (high)+1 tiers | 2025-11-18 | 28.0 | — | 95.7 | 0 | 4.50 | |
| 58 | GLM-5 (Reasoning)+1 tiers | Z AI | 2026-02-11 | 27.9 | — | — | 0 | 1.55 |
| 59 | Inkling Small | Thinking Machines | 2026-07-30 | 27.8 | 52.9 | — | 219 | 0.53 |
| 60 | JT-4.1 Flash 236B A21B | China Mobile | 2026-07-09 | 27.3 | 52.4 | — | 0 | — |
| 61 | Grok Build 0.1 0616 | SpaceXAI | 2026-06-16 | 27.2 | 51.5 | — | 0 | 1.25 |
| 62 | Kimi K2.6+1 tiers | Kimi | 2026-04-20 | 27.0 | 61.8 | — | 0 | 1.71 |
| 63 | Qwen3.6 Plus | Alibaba | 2026-04-02 | 27.0 | 54.5 | — | 0 | 1.12 |
| 64 | Agnes 2.5 Pro Alpha | Sapiens AI | 2026-07-24 | 26.8 | 58.8 | — | 192 | 0.56 |
| 65 | Quasar 438B (max, based on GLM-5.2) | Multiverse Computing | 2026-08-10 | 26.7 | 61.2 | — | 163 | 0.90 |
| 66 | GLM-5-Turbo | Z AI | 2026-03-15 | 26.6 | — | — | 0 | — |
| 67 | Claude Opus 4.6 (Non-reasoning, High Effort) | Anthropic | 2026-02-05 | 26.4 | — | — | 0 | 10.00 |
| 68 | Gemini 3 Flash Preview (Reasoning) | 2025-12-17 | 26.3 | — | 97.0 | 0 | 1.12 | |
| 69 | GLM-5.1 (Reasoning)+1 tiers | Z AI | 2026-04-07 | 26.1 | 55.8 | — | 0 | 2.00 |
| 70 | DeepSeek V4 Flash 0420 (Reasoning, High Effort)+2 tiers | DeepSeek | 2026-04-24 | 26.0 | 52.0 | — | 0 | 0.17 |
| 71 | GPT-5.5 Instant (June 2026) | OpenAI | 2026-06-25 | 26.0 | 39.4 | — | 143 | 11.25 |
| 72 | MiMo-V2.5-Pro+1 tiers | Xiaomi | 2026-04-22 | 26.0 | 60.2 | — | 50 | 0.54 |
| 73 | Kimi K2.7 Code | Kimi | 2026-06-12 | 25.8 | 60.8 | — | 52 | 1.71 |
| 74 | Grok 4.20 0309 v2 (Reasoning)+1 tiers | SpaceXAI | 2026-04-07 | 25.7 | — | — | 0 | 1.56 |
| 75 | Hy3+1 tiers | Tencent | 2026-07-06 | 25.3 | 58.8 | — | 92 | 0.25 |
| 76 | K2 Horizon MoVA 36B A4BNEW | Institute of Foundation Models | 2026-09-03 | 25.3 | — | — | 0 | — |
| 77 | Grok 4.20 0309 (Reasoning)+1 tiers | SpaceXAI | 2026-03-10 | 25.2 | — | — | 0 | 3.00 |
| 78 | MiMo-V2.5 | Xiaomi | 2026-04-22 | 25.2 | 56.8 | — | 32 | 0.17 |
| 79 | Qwen3.7 Plus | Alibaba | 2026-06-01 | 25.2 | 55.9 | — | 68 | 0.70 |
| 80 | MiMo-V2-Omni-0327 | Xiaomi | 2026-03-27 | 25.1 | — | — | 0 | — |
| 81 | Inkling (xhigh) | Thinking Machines | 2026-07-15 | 25.0 | 52.1 | — | 100 | 1.76 |
| 82 | GPT-5 Codex (high) | OpenAI | 2025-09-23 | 24.9 | — | 98.7 | 0 | 3.44 |
| 83 | Grok 4.3 (high)+3 tiers | SpaceXAI | 2026-04-30 | 24.9 | 42.2 | — | 0 | 1.56 |
| 84 | Ling 3.0 Flash | InclusionAI | 2026-08-04 | 24.9 | 50.6 | — | 382 | 0.11 |
| 85 | Claude Sonnet 4.6 (Non-reasoning, High Effort) | Anthropic | 2026-02-17 | 24.7 | — | — | 0 | 6.00 |
| 86 | GPT-5.1 (high)+1 tiers | OpenAI | 2025-11-13 | 24.7 | 49.4 | 94.0 | 0 | 3.44 |
| 87 | Solar Open2 250B | Upstage | 2026-08-12 | 24.7 | 45.0 | — | 0 | — |
| 88 | Ling-3.0-flash-VLNEW | InclusionAI | 2026-09-10 | 24.6 | 57.0 | — | 141 | 0.11 |
| 89 | GPT-5.4 mini (xhigh)+2 tiers | OpenAI | 2026-03-17 | 24.1 | 56.1 | — | 0 | 1.69 |
| 90 | MiMo-V2-Omni | Xiaomi | 2026-03-19 | 23.9 | — | — | 0 | — |
| 91 | Claude Opus 4.5 (Non-reasoning) | Anthropic | 2025-11-24 | 23.7 | — | 62.7 | 0 | 10.00 |
| 92 | GPT-5.1 Codex (high) | OpenAI | 2025-11-13 | 23.7 | — | 95.7 | 0 | 3.44 |
| 93 | GLM 5V Turbo (Reasoning) | Z AI | 2026-04-01 | 23.5 | — | — | 0 | — |
| 94 | Kimi K2.5 (Reasoning)+1 tiers | Kimi | 2026-01-27 | 23.5 | 46.8 | — | 0 | 1.14 |
| 95 | Claude Sonnet 4.6 (Non-reasoning, Low Effort) | Anthropic | 2026-02-17 | 23.3 | — | — | 0 | 6.00 |
| 96 | GPT-5 (high)+3 tiers | OpenAI | 2025-08-07 | 23.0 | 37.8 | 94.3 | 0 | 3.44 |
| 97 | Nemotron 3 Ultra 550B A55B (Reasoning) | NVIDIA | 2026-06-04 | 22.9 | 49.3 | — | 186 | 1.05 |
| 98 | Qwen3.5 27B (Reasoning)+1 tiers | Alibaba | 2026-02-24 | 22.9 | — | — | 0 | 0.82 |
| 99 | Claude 4.1 Opus (Reasoning) | Anthropic | 2025-08-05 | 22.8 | — | 80.3 | 0 | 30.00 |
| 100 | MiniMax-M2.5 | MiniMax | 2026-02-12 | 22.8 | — | — | 0 | 0.53 |
| 101 | MiniMax-M2.7 | MiniMax | 2026-03-18 | 22.8 | 52.6 | — | 0 | 0.53 |
| 102 | A.X-K2 | SK Telecom | 2026-08-12 | 22.7 | 38.8 | — | 0 | — |
| 103 | GPT-5.5 Instant (May 2026) | OpenAI | 2026-05-05 | 22.7 | — | — | 0 | 11.25 |
| 104 | Hy3-preview (Reasoning) | Tencent | 2026-04-23 | 22.7 | — | — | 0 | 0.10 |
| 105 | Ling-3.0-flash-FinNEW | InclusionAI | 2026-09-11 | 22.6 | 55.6 | — | 162 | — |
| 106 | Grok 4 | SpaceXAI | 2025-07-10 | 22.5 | — | 92.7 | 0 | 6.00 |
| 107 | MiMo-V2-Flash (Feb 2026) | Xiaomi | 2025-12-16 | 22.4 | — | — | 0 | — |
| 108 | GLM-4.7 (Reasoning)+1 tiers | Z AI | 2025-12-22 | 22.2 | 45.3 | 95.0 | 0 | 1.00 |
| 109 | Gemini 3.5 Flash-Lite | 2026-07-21 | 22.2 | 49.3 | — | 352 | 0.85 | |
| 110 | Kimi K2 Thinking | Kimi | 2025-11-06 | 22.0 | — | 94.7 | 0 | 1.07 |
| 111 | o3-pro | OpenAI | 2025-06-10 | 21.9 | — | — | 0 | 35.00 |
| 112 | G9v3-39A5B | AI9Stars | 2026-08-20 | 21.8 | 33.1 | — | 0 | — |
| 113 | KAT Coder Pro V2 | KwaiKAT | 2026-03-27 | 21.7 | 59.5 | — | 0 | 0.53 |
| 114 | DeepSeek V3.2 (Reasoning) | DeepSeek | 2025-12-01 | 21.5 | 44.2 | 92.0 | 0 | 0.32 |
| 115 | Qwen3.5 397B A17B (Non-reasoning)+1 tiers | Alibaba | 2026-02-16 | 21.4 | — | — | 85 | 1.35 |
| 116 | Qwen3.6 27B (Reasoning)+1 tiers | Alibaba | 2026-04-22 | 21.4 | 53.7 | — | 0 | 1.35 |
| 117 | Qwen3 Max Thinking | Alibaba | 2026-01-26 | 21.3 | — | — | 0 | — |
| 118 | MiniMax-M2.1 | MiniMax | 2025-12-23 | 20.9 | — | 82.7 | 0 | 0.53 |
| 119 | MiMo-V2-Flash (Reasoning) | Xiaomi | 2025-12-16 | 20.8 | — | 96.3 | 0 | 0.15 |
| 120 | Claude 4.5 Sonnet (Reasoning) | Anthropic | 2025-09-29 | 20.7 | 52.1 | 88.0 | 0 | 6.00 |
| 121 | GPT-5.4 nano (xhigh)+2 tiers | OpenAI | 2026-03-17 | 20.7 | 56.1 | — | 0 | 0.46 |
| 122 | Claude 4 Opus (Reasoning) | Anthropic | 2025-05-22 | 20.6 | — | 73.3 | 0 | 30.00 |
| 123 | GPT-5 mini (medium)+2 tiers | OpenAI | 2025-08-07 | 20.6 | — | 85.0 | 0 | 0.69 |
| 124 | K2 Horizon 7BNEW | Institute of Foundation Models | 2026-09-03 | 20.6 | 38.6 | — | 0 | — |
| 125 | GPT-5.1 Codex mini (high) | OpenAI | 2025-11-13 | 20.4 | — | 91.7 | 0 | 0.69 |
| 126 | Grok 4.1 Fast (Reasoning) | SpaceXAI | 2025-11-19 | 20.4 | — | 89.3 | 0 | — |
| 127 | Qwen3.5 Omni Plus | Alibaba | 2026-03-30 | 20.4 | — | — | 90 | 1.50 |
| 128 | o3 | OpenAI | 2025-04-16 | 20.2 | — | 88.3 | 159 | 3.50 |
| 129 | K-EXAONE 2.0 0803 | LG AI Research | 2026-08-12 | 19.7 | 40.6 | — | 0 | — |
| 130 | Step 3.7 Flash | StepFun | 2026-05-29 | 19.5 | 39.6 | — | 204 | 0.44 |
| 131 | Claude 4.5 Sonnet (Non-reasoning) | Anthropic | 2025-09-29 | 19.3 | — | 37.0 | 0 | 6.00 |
| 132 | Qwen3.5 35B A3B (Reasoning)+1 tiers | Alibaba | 2026-02-24 | 19.3 | — | — | 0 | 0.69 |
| 133 | LongCat 2.0 | LongCat | 2026-06-29 | 19.1 | 45.3 | — | 0 | 0.53 |
| 134 | Gemma 4 31B (Reasoning)+1 tiers | 2026-04-02 | 19.0 | 43.4 | — | 35 | — | |
| 135 | Claude 4 Sonnet (Reasoning) | Anthropic | 2025-05-22 | 18.9 | 37.6 | 74.3 | 0 | — |
| 136 | JT-35B-Flash | China Mobile | 2026-05-14 | 18.7 | — | — | 0 | — |
| 137 | Claude 4.1 Opus (Non-reasoning) | Anthropic | 2025-08-05 | 18.6 | — | — | 0 | 30.00 |
| 138 | KAT-Coder-Pro V1 | KwaiKAT | 2025-11-11 | 18.6 | — | 94.7 | 0 | — |
| 139 | MiniMax-M2 | MiniMax | 2025-10-26 | 18.6 | — | 78.3 | 0 | 0.53 |
| 140 | GLM-4.6 (Reasoning) | Z AI | 2025-09-30 | 18.5 | 45.8 | 86.0 | 0 | 0.96 |
| 141 | Qwen3.6 35B A3B (Reasoning)+1 tiers | Alibaba | 2026-04-16 | 18.2 | 41.9 | — | 119 | 0.84 |
| 142 | Gemini 3 Flash Preview (Non-reasoning) | 2025-12-17 | 17.9 | — | 55.7 | 0 | 1.12 | |
| 143 | Grok 4 Fast (Reasoning) | SpaceXAI | 2025-09-19 | 17.9 | — | 89.7 | 0 | 0.28 |
| 144 | Claude 3.7 Sonnet (Reasoning) | Anthropic | 2025-02-24 | 17.7 | 36.4 | 56.3 | 0 | — |
| 145 | Qwen3.5 122B A10B (Non-reasoning)+1 tiers | Alibaba | 2026-02-24 | 17.7 | 43.3 | — | 143 | 1.10 |
| 146 | Muse Glimmer (high) | Meta | 2026-08-10 | 17.5 | 49.0 | — | 91 | 0.64 |
| 147 | Ling-2.6-1T | InclusionAI | 2026-04-23 | 17.0 | — | — | 0 | 0.85 |
| 148 | Step 3.5 Flash 2603 | StepFun | 2026-04-02 | 17.0 | — | — | 0 | 0.15 |
| 149 | Claude 4.5 Haiku (Reasoning) | Anthropic | 2025-10-15 | 16.9 | 43.9 | 83.7 | 148 | 2.00 |
| 150 | Doubao Seed Code | ByteDance Seed | 2025-11-11 | 16.9 | — | 79.3 | 0 | — |
| 151 | Gemma 4 26B A4B (Reasoning)+1 tiers | 2026-04-02 | 16.7 | 39.3 | — | 0 | 0.18 | |
| 152 | o4-mini (high) | OpenAI | 2025-04-16 | 16.7 | — | 90.7 | 0 | 1.93 |
| 153 | Claude 4 Opus (Non-reasoning) | Anthropic | 2025-05-22 | 16.6 | — | 36.3 | 0 | 30.00 |
| 154 | Claude 4 Sonnet (Non-reasoning) | Anthropic | 2025-05-22 | 16.6 | — | 38.0 | 0 | — |
| 155 | DeepSeek V3.2 Exp (Reasoning) | DeepSeek | 2025-09-29 | 16.6 | — | 87.7 | 0 | 0.32 |
| 156 | Ring-2.6-1T | InclusionAI | 2026-05-08 | 16.6 | 42.8 | — | 125 | 0.85 |
| 157 | Step 3.5 Flash | StepFun | 2026-02-02 | 16.6 | — | — | 0 | 0.15 |
| 158 | Qwen3 Max Thinking (Preview) | Alibaba | 2025-11-03 | 16.3 | — | 82.3 | 0 | 2.40 |
| 159 | Gemini 2.5 Pro | 2025-06-05 | 16.1 | 33.3 | 87.7 | 0 | 3.44 | |
| 160 | DeepSeek V3.2 (Non-reasoning) | DeepSeek | 2025-12-01 | 16.0 | — | 59.0 | 0 | 0.32 |
| 161 | MiMo-V2-Flash (Non-reasoning) | Xiaomi | 2025-12-16 | 16.0 | 49.8 | 67.7 | 0 | — |
| 162 | Gemini 3.1 Flash-Lite | 2026-03-03 | 15.6 | 34.7 | — | 0 | 0.56 | |
| 163 | K2 Horizon 3.7BNEW | Institute of Foundation Models | 2026-09-03 | 15.6 | 26.1 | — | 0 | — |
| 164 | Qwen3 Max | Alibaba | 2025-09-23 | 15.6 | — | 80.7 | 0 | 2.40 |
| 165 | Gemini 2.5 Flash Preview (Sep '25) (Reasoning) | 2025-09-25 | 15.5 | — | 78.3 | 0 | — | |
| 166 | Claude 4.5 Haiku (Non-reasoning) | Anthropic | 2025-10-15 | 15.4 | — | 39.0 | 107 | 2.00 |
| 167 | Claude 3.7 Sonnet (Non-reasoning) | Anthropic | 2025-02-24 | 15.3 | — | 21.0 | 0 | 6.00 |
| 168 | Kimi K2 0905 | Kimi | 2025-09-05 | 15.3 | — | 57.3 | 0 | 1.07 |
| 169 | Ling 3.0 Tiny | InclusionAI | 2026-08-06 | 15.3 | 26.5 | — | 58 | — |
| 170 | o1 | OpenAI | 2024-12-05 | 15.2 | 39.7 | — | 0 | 26.25 |
| 171 | Gemini 2.5 Pro Preview (Mar' 25) | 2025-03-25 | 15.0 | 46.7 | — | 0 | — | |
| 172 | GLM-4.6 (Non-reasoning) | Z AI | 2025-09-30 | 14.9 | — | 44.3 | 0 | 0.98 |
| 173 | GLM-4.7-Flash (Reasoning)+1 tiers | Z AI | 2026-01-19 | 14.9 | — | — | 0 | 0.15 |
| 174 | DeepSeek V3.1 Terminus (Reasoning) | DeepSeek | 2025-09-22 | 14.8 | 43.5 | 89.7 | 0 | 1.91 |
| 175 | Granite 4.2 30BNEW | IBM | 2026-08-25 | 14.8 | 29.9 | — | 77 | 0.28 |
| 176 | Grok 3 mini Reasoning (high) | SpaceXAI | 2025-02-19 | 14.6 | — | 84.7 | 0 | 0.35 |
| 177 | DeepSeek V3.2 Speciale | DeepSeek | 2025-12-01 | 14.5 | — | 96.7 | 0 | — |
| 178 | Gemini 2.5 Pro Preview (May' 25) | 2025-05-06 | 14.5 | — | — | 0 | 3.44 | |
| 179 | K-EXAONE (Reasoning)+1 tiers | LG AI Research | 2025-12-31 | 14.4 | 32.1 | 90.3 | 0 | — |
| 180 | ERNIE 5.0 Thinking Preview | Baidu | 2025-11-13 | 14.3 | — | 85.0 | 0 | — |
| 181 | Gemma 4 12B (Reasoning)+1 tiers | 2026-06-03 | 14.2 | 31.0 | — | 143 | 0.15 | |
| 182 | Mistral Medium 3.5 | Mistral | 2026-04-29 | 14.2 | 46.9 | — | 163 | 3.00 |
| 183 | Nova 2.0 Pro Preview (medium)+1 tiers | Amazon | 2025-11-27 | 14.2 | 34.0 | 89.0 | 125 | 3.44 |
| 184 | Grok Code Fast 1 | SpaceXAI | 2025-08-28 | 14.1 | — | 43.3 | 0 | — |
| 185 | DeepSeek V3.1 Terminus (Non-reasoning) | DeepSeek | 2025-09-22 | 13.9 | — | 53.7 | 0 | 0.45 |
| 186 | DeepSeek V3.2 Exp (Non-reasoning) | DeepSeek | 2025-09-29 | 13.9 | — | 57.7 | 0 | 0.32 |
| 187 | Apriel-v1.5-15B-Thinker | ServiceNow | 2025-09-30 | 13.8 | — | 87.5 | 0 | — |
| 188 | Mercury 2 | Inception | 2026-02-20 | 13.8 | 31.1 | — | 734 | 0.38 |
| 189 | DeepSeek V3.1 (Non-reasoning) | DeepSeek | 2025-08-21 | 13.7 | — | 49.7 | 0 | 0.85 |
| 190 | Qwen3.5 9B (Reasoning)+1 tiers | Alibaba | 2026-03-02 | 13.7 | 28.7 | — | 71 | 0.15 |
| 191 | Nova 2.0 Omni (medium)+1 tiers | Amazon | 2025-11-26 | 13.6 | — | 89.7 | 0 | 0.85 |
| 192 | DeepSeek V3.1 (Reasoning) | DeepSeek | 2025-08-21 | 13.5 | — | 89.7 | 0 | 0.86 |
| 193 | Apriel-v1.6-15B-Thinker | ServiceNow | 2025-11-25 | 13.4 | — | 88.0 | 0 | — |
| 194 | Nova 2.0 Lite (high)+2 tiers | Amazon | 2025-10-29 | 13.4 | 23.0 | 94.3 | 185 | 0.85 |
| 195 | Qwen3 VL 235B A22B (Reasoning) | Alibaba | 2025-09-23 | 13.4 | — | 88.3 | 0 | 1.30 |
| 196 | EXAONE 4.5 33B+1 tiers | LG AI Research | 2026-04-09 | 13.2 | 23.6 | — | 0 | — |
| 197 | Command A+ | Cohere | 2026-05-20 | 13.1 | 27.8 | — | 233 | — |
| 198 | DeepSeek R1 0528 (May '25) | DeepSeek | 2025-05-28 | 13.1 | — | 76.0 | 0 | 1.76 |
| 199 | Gemini 2.5 Flash (Reasoning) | 2025-05-20 | 13.1 | — | 73.3 | 0 | 0.85 | |
| 200 | Qwen3.5 4B (Reasoning)+1 tiers | Alibaba | 2026-03-02 | 13.1 | 22.6 | — | 17 | 0.06 |
| 201 | GPT-5 nano (high)+2 tiers | OpenAI | 2025-08-07 | 13.0 | — | 83.7 | 0 | 0.14 |
| 202 | Nemotron 3.5 Lightning | NVIDIA | 2026-08-11 | 12.9 | 26.8 | — | 291 | 0.10 |
| 203 | GLM-4.5 (Reasoning) | Z AI | 2025-07-28 | 12.8 | — | 73.7 | 0 | — |
| 204 | Nemotron 3 Super 120B A12B (Reasoning) | NVIDIA | 2026-03-11 | 12.8 | 37.7 | — | 210 | 0.31 |
| 205 | GPT-4.1 | OpenAI | 2025-04-14 | 12.7 | — | 34.7 | 0 | 3.50 |
| 206 | Kimi K2 | Kimi | 2025-07-11 | 12.7 | — | 57.0 | 0 | 1.00 |
| 207 | Qwen3 235B A22B 2507 (Reasoning) | Alibaba | 2025-07-25 | 12.7 | 22.1 | 91.0 | 0 | 0.75 |
| 208 | Qwen3 Max (Preview) | Alibaba | 2025-09-05 | 12.6 | — | 75.0 | 0 | 2.40 |
| 209 | MiniCPM5-2BNEW | OpenBMB | 2026-09-07 | 12.5 | 14.5 | — | 0 | — |
| 210 | Qwen3.5 Omni Flash | Alibaba | 2026-03-30 | 12.5 | — | — | 235 | 0.28 |
| 211 | o3-mini+1 tiers | OpenAI | 2025-01-31 | 12.5 | — | — | 0 | 1.93 |
| 212 | Gemini 2.5 Flash Preview (Sep '25) (Non-reasoning) | 2025-09-25 | 12.4 | — | 56.7 | 0 | — | |
| 213 | o1-pro | OpenAI | 2025-03-19 | 12.4 | — | — | 0 | 262.50 |
| 214 | JT-MINI | China Mobile | 2026-04-15 | 12.2 | — | — | 0 | — |
| 215 | Grok 3 | SpaceXAI | 2025-02-19 | 12.1 | — | 58.0 | 0 | 8.00 |
| 216 | Seed-OSS-36B-Instruct | ByteDance Seed | 2025-08-20 | 12.1 | — | 84.7 | 0 | 0.30 |
| 217 | Qwen3 235B A22B 2507 Instruct | Alibaba | 2025-07-21 | 12.0 | — | 71.7 | 0 | 0.40 |
| 218 | Qwen3 Coder 480B A35B Instruct | Alibaba | 2025-07-22 | 11.9 | — | 39.3 | 0 | 3.00 |
| 219 | Qwen3 VL 32B (Reasoning) | Alibaba | 2025-10-21 | 11.9 | — | 84.7 | 0 | 0.28 |
| 220 | Magistral Medium 1.2 | Mistral | 2025-09-18 | 11.8 | 21.3 | 82.0 | 0 | — |
| 221 | Sonar Reasoning Pro | Perplexity | 2025-01-28 | 11.8 | — | — | 0 | — |
| 222 | Gemini 2.5 Flash Preview (Reasoning) | 2025-04-17 | 11.7 | — | — | 0 | — | |
| 223 | HyperNova 60B 2605 (high, based on gpt-oss-120b) | Multiverse Computing | 2026-05-26 | 11.7 | 23.2 | — | 0 | — |
| 224 | MiniMax M1 80k | MiniMax | 2025-06-17 | 11.7 | — | 61.0 | 0 | 0.96 |
| 225 | Nemotron Cascade 2 30B A3B | NVIDIA | 2026-03-19 | 11.7 | 25.3 | — | 0 | — |
| 226 | gpt-oss-120b (high)+1 tiers | OpenAI | 2025-08-05 | 11.6 | 30.4 | 93.4 | 218 | 0.26 |
| 227 | K2 Think V2 | Institute of Foundation Models | 2025-12-15 | 11.5 | 21.0 | — | 0 | — |
| 228 | LongCat Flash Lite | LongCat | 2026-01-28 | 11.5 | — | — | 0 | — |
| 229 | DeepSeek R1 (Jan '25) | DeepSeek | 2025-01-20 | 11.4 | 24.6 | 68.0 | 0 | 2.50 |
| 230 | HyperCLOVA X SEED Think (32B) | Naver | 2025-12-26 | 11.4 | — | 59.0 | 0 | — |
| 231 | o1-preview | OpenAI | 2024-09-12 | 11.4 | 34.0 | 79.7 | 0 | 28.88 |
| 232 | Grok 4.1 Fast (Non-reasoning) | SpaceXAI | 2025-11-19 | 11.3 | — | 34.3 | 0 | — |
| 233 | Mistral Small 4 (Reasoning)+1 tiers | Mistral | 2026-03-16 | 11.3 | 26.6 | — | 169 | 0.26 |
| 234 | GLM-4.6V (Reasoning) | Z AI | 2025-12-08 | 11.2 | — | 85.3 | 0 | 0.45 |
| 235 | Qwen3 Next 80B A3B (Reasoning) | Alibaba | 2025-09-11 | 11.2 | 17.4 | 84.3 | 168 | 0.41 |
| 236 | GLM-4.5-Air | Z AI | 2025-07-28 | 11.1 | — | 80.7 | 0 | 0.37 |
| 237 | Granite 4.2 8BNEW | IBM | 2026-08-25 | 11.1 | 22.4 | — | 96 | 0.11 |
| 238 | Grok 4 Fast (Non-reasoning) | SpaceXAI | 2025-09-19 | 11.1 | — | 41.3 | 0 | 0.28 |
| 239 | Mi:dm K 2.5 Pro | Korea Telecom | 2025-12-11 | 11.0 | — | 76.7 | 0 | — |
| 240 | Ring-1T | InclusionAI | 2025-10-13 | 10.9 | — | 89.3 | 0 | — |
| 241 | G9v3-3B | AI9Stars | 2026-07-23 | 10.8 | 9.9 | — | 0 | — |
| 242 | Trinity Large Thinking | Arcee AI | 2026-04-01 | 10.8 | 25.8 | — | 324 | 0.41 |
| 243 | INTELLECT-3 | Prime Intellect | 2025-11-27 | 10.6 | — | 88.0 | 0 | — |
| 244 | GPT-5 (ChatGPT) | OpenAI | 2025-08-07 | 10.4 | — | 48.3 | 0 | — |
| 245 | Gemini 2.5 Flash-Lite Preview (Sep '25) (Reasoning) | 2025-09-25 | 10.4 | — | 68.7 | 0 | 0.17 | |
| 246 | Grok 3 Reasoning Beta | SpaceXAI | 2025-02-19 | 10.4 | — | — | 0 | — |
| 247 | Solar Open 100B (Reasoning) | Upstage | 2025-12-17 | 10.4 | — | — | 0 | — |
| 248 | Nemotron 3 Nano Omni 30B A3B Reasoning | NVIDIA | 2026-04-29 | 10.3 | 13.8 | — | 0 | 0.16 |
| 249 | GPT-4.1 mini | OpenAI | 2025-04-14 | 10.2 | 20.2 | 46.3 | 0 | 0.70 |
| 250 | Llama 4 Maverick | Meta | 2025-04-05 | 10.0 | 16.3 | 19.3 | 107 | 0.42 |
| 251 | MiniMax M1 40k | MiniMax | 2025-06-17 | 10.0 | — | 13.7 | 0 | — |
| 252 | Nova 2.0 Pro Preview (Non-reasoning) | Amazon | 2025-11-27 | 10.0 | 20.9 | 30.7 | 136 | 3.44 |
| 253 | gpt-oss-20b (low)+1 tiers | OpenAI | 2025-08-05 | 10.0 | — | 62.3 | 252 | 0.10 |
| 254 | Gemini 2.5 Flash (Non-reasoning) | 2025-05-20 | 9.9 | — | 60.3 | 0 | 0.85 | |
| 255 | K2-V2 (high)+2 tiers | Institute of Foundation Models | 2025-12-05 | 9.9 | — | 78.3 | 0 | — |
| 256 | North Mini Code | Cohere | 2026-06-09 | 9.9 | 36.5 | — | 89 | — |
| 257 | Qwen3 VL 235B A22B Instruct | Alibaba | 2025-09-23 | 9.9 | — | 70.7 | 0 | 0.70 |
| 258 | Qwen3 30B A3B 2507 (Reasoning) | Alibaba | 2025-07-30 | 9.8 | 12.1 | 56.3 | 0 | 0.75 |
| 259 | o1-mini | OpenAI | 2024-09-12 | 9.8 | — | — | 0 | — |
| 260 | DeepSeek V3 0324 | DeepSeek | 2025-03-25 | 9.7 | 21.2 | 41.0 | 0 | 0.93 |
| 261 | Ling 2.6 Flash | InclusionAI | 2026-04-21 | 9.7 | 25.3 | — | 0 | — |
| 262 | GPT-4.5 (Preview) | OpenAI | 2025-02-27 | 9.6 | — | — | 0 | — |
| 263 | Qwen3 Coder 30B A3B Instruct | Alibaba | 2025-07-31 | 9.6 | — | 29.0 | 0 | 0.90 |
| 264 | Qwen3 Next 80B A3B Instruct | Alibaba | 2025-09-11 | 9.6 | — | 66.3 | 173 | 0.41 |
| 265 | Tri-21B-think Preview | Trillion Labs | 2026-02-10 | 9.6 | — | — | 0 | — |
| 266 | DiffusionGemma 26B A4B | 2026-06-10 | 9.5 | 19.7 | — | 0 | — | |
| 267 | QwQ 32B | Alibaba | 2025-03-05 | 9.5 | — | 29.0 | 0 | 0.74 |
| 268 | Qwen3 235B A22B (Reasoning) | Alibaba | 2025-04-28 | 9.5 | — | 82.0 | 0 | 2.62 |
| 269 | Qwen3 VL 30B A3B (Reasoning) | Alibaba | 2025-10-03 | 9.5 | — | 82.3 | 0 | 0.75 |
| 270 | Gemini 2.0 Flash Thinking Experimental (Jan '25) | 2025-01-21 | 9.4 | 24.1 | — | 0 | — | |
| 271 | Gemini 2.5 Flash-Lite Preview (Sep '25) (Non-reasoning) | 2025-09-25 | 9.3 | — | 46.7 | 0 | 0.17 | |
| 272 | Mistral Large 3 | Mistral | 2025-12-02 | 9.3 | 20.1 | 38.0 | 77 | 0.75 |
| 273 | Ling-1T | InclusionAI | 2025-10-08 | 9.2 | — | 71.3 | 0 | — |
| 274 | Mistral Medium 3.1 | Mistral | 2025-08-12 | 9.2 | 20.5 | 38.3 | 0 | 0.80 |
| 275 | Motif-2-12.7B-Reasoning | Motif Technologies | 2025-12-04 | 9.2 | — | 80.3 | 0 | — |
| 276 | Nova Premier | Amazon | 2025-04-30 | 9.2 | — | 17.3 | 0 | 5.00 |
| 277 | Qwen3 Coder Next | Alibaba | 2026-02-03 | 9.2 | 36.2 | — | 118 | 0.56 |
| 278 | Granite 4.2 3BNEW | IBM | 2026-08-25 | 9.1 | 17.5 | — | 231 | 0.05 |
| 279 | Magistral Medium 1 | Mistral | 2025-06-10 | 9.1 | — | 40.3 | 0 | — |
| 280 | Solar Pro 2 (Preview) (Reasoning) | Upstage | 2025-05-20 | 9.1 | — | — | 0 | — |
| 281 | Devstral Medium | Mistral | 2025-07-10 | 9.0 | — | 4.7 | 0 | — |
| 282 | GPT-4o (March 2025, chatgpt-4o-latest) | OpenAI | 2025-03-27 | 9.0 | — | 25.7 | 0 | — |
| 283 | Llama Nemotron Super 49B v1.5 (Reasoning) | NVIDIA | 2025-07-25 | 9.0 | — | 76.7 | 150 | 0.40 |
| 284 | Mistral Medium 3 | Mistral | 2025-05-07 | 9.0 | — | 30.3 | 0 | 0.80 |
| 285 | Tri-21B-Think | Trillion Labs | 2026-02-10 | 9.0 | — | — | 0 | — |
| 286 | Claude 3.5 Haiku | Anthropic | 2024-10-22 | 8.9 | 15.9 | — | 0 | — |
| 287 | Gemini 2.0 Flash (Feb '25) | 2025-02-05 | 8.9 | — | 21.7 | 0 | — | |
| 288 | Gemma 4 E4B (Reasoning)+1 tiers | 2026-04-03 | 8.9 | 9.4 | — | 43 | 0.04 | |
| 289 | Llama 3.3 Nemotron Super 49B v1 (Reasoning) | NVIDIA | 2025-03-18 | 8.9 | — | 54.7 | 0 | — |
| 290 | NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) | NVIDIA | 2025-12-15 | 8.9 | 14.4 | 91.0 | 131 | 0.09 |
| 291 | MiniCPM5-1B (Reasoning)+1 tiers | OpenBMB | 2026-05-25 | 8.8 | — | — | 0 | — |
| 292 | Qwen3 4B 2507 (Reasoning) | Alibaba | 2025-08-06 | 8.8 | — | 82.7 | 0 | — |
| 293 | Sarvam 105B (high) | Sarvam | 2026-03-06 | 8.8 | — | — | 0 | 0.07 |
| 294 | Claude 3 Opus | Anthropic | 2024-03-04 | 8.7 | 19.5 | — | 0 | 30.00 |
| 295 | Devstral Small (May '25) | Mistral | 2025-05-21 | 8.7 | — | — | 0 | — |
| 296 | Gemini 2.0 Pro Experimental (Feb '25) | 2025-02-05 | 8.7 | 25.5 | — | 0 | — | |
| 297 | Gemini 2.5 Flash Preview (Non-reasoning) | 2025-04-17 | 8.7 | — | — | 0 | — | |
| 298 | Nova 2.0 Lite (Non-reasoning) | Amazon | 2025-10-29 | 8.7 | — | 33.7 | 222 | 0.85 |
| 299 | Sonar Reasoning | Perplexity | 2025-01-28 | 8.7 | — | — | 0 | — |
| 300 | Devstral 2 | Mistral | 2025-12-09 | 8.6 | 31.3 | 36.7 | 143 | — |
| 301 | Magistral Small 1.2 | Mistral | 2025-09-17 | 8.6 | 14.7 | 80.3 | 0 | 0.75 |
| 302 | Qwen3 32B (Reasoning) | Alibaba | 2025-04-28 | 8.6 | 15.3 | 73.0 | 0 | 0.28 |
| 303 | DeepSeek V3 (Dec '24) | DeepSeek | 2024-12-26 | 8.5 | 23.0 | 26.0 | 0 | 0.46 |
| 304 | Gemini 2.5 Flash-Lite (Reasoning) | 2025-06-17 | 8.5 | — | 53.3 | 0 | 0.17 | |
| 305 | DeepSeek R1 Distill Qwen 32B | DeepSeek | 2025-01-20 | 8.4 | — | 63.0 | 0 | — |
| 306 | GLM-4.6V (Non-reasoning) | Z AI | 2025-12-08 | 8.4 | — | 26.3 | 0 | 0.45 |
| 307 | GPT-4o (Nov '24) | OpenAI | 2024-11-20 | 8.4 | — | 6.0 | 0 | 4.38 |
| 308 | LFM2.5-2.6B | Liquid AI | 2026-08-04 | 8.4 | 7.7 | — | 0 | — |
| 309 | Nanbeige4.1-3B | Nanbeige | 2026-02-11 | 8.4 | 9.6 | — | 0 | — |
| 310 | Qwen3 VL 32B Instruct | Alibaba | 2025-10-21 | 8.4 | — | 68.3 | 0 | 0.28 |
| 311 | Qwen3 235B A22B (Non-reasoning) | Alibaba | 2025-04-28 | 8.3 | — | 23.7 | 0 | 1.23 |
| 312 | EXAONE 4.0 32B (Reasoning) | LG AI Research | 2025-07-15 | 8.2 | — | 80.0 | 0 | — |
| 313 | Gemini 2.0 Flash (experimental) | 2024-12-11 | 8.2 | — | — | 0 | — | |
| 314 | Magistral Small 1 | Mistral | 2025-06-10 | 8.2 | — | 41.3 | 0 | — |
| 315 | Mistral Small 3.2 | Mistral | 2025-06-20 | 8.2 | 12.5 | 27.0 | 0 | 0.15 |
| 316 | Nova 2.0 Omni (Non-reasoning) | Amazon | 2025-11-26 | 8.2 | — | 37.0 | 0 | 0.85 |
| 317 | Qwen3 14B (Reasoning) | Alibaba | 2025-04-28 | 8.2 | 13.8 | 55.7 | 0 | 1.31 |
| 318 | Qwen3 VL 8B (Reasoning) | Alibaba | 2025-10-14 | 8.2 | — | 30.7 | 0 | 0.66 |
| 319 | DeepSeek R1 0528 Qwen3 8B | DeepSeek | 2025-05-29 | 8.1 | — | 63.7 | 0 | — |
| 320 | Llama 4 Scout | Meta | 2025-04-05 | 8.1 | 8.2 | 14.0 | 104 | 0.31 |
| 321 | Qwen2.5 Max | Alibaba | 2025-01-28 | 8.0 | — | — | 0 | — |
| 322 | Claude 3.5 Sonnet (Oct '24) | Anthropic | 2024-10-22 | 7.9 | 30.2 | — | 0 | 6.00 |
| 323 | DeepSeek R1 Distill Llama 70B | DeepSeek | 2025-01-20 | 7.9 | — | 53.7 | 0 | 0.80 |
| 324 | Gemini 1.5 Pro (Sep '24) | 2024-09-24 | 7.9 | 23.6 | — | 0 | — | |
| 325 | Hermes 4 - Llama-3.1 70B (Reasoning) | Nous Research | 2025-08-27 | 7.9 | — | 68.7 | 0 | — |
| 326 | Qwen3 VL 30B A3B Instruct | Alibaba | 2025-10-03 | 7.9 | — | 72.3 | 0 | 0.35 |
| 327 | Solar Pro 2 (Preview) (Non-reasoning) | Upstage | 2025-05-20 | 7.9 | — | — | 0 | — |
| 328 | DeepSeek R1 Distill Qwen 14B | DeepSeek | 2025-01-20 | 7.8 | — | 55.7 | 0 | — |
| 329 | Falcon-H1R-7B | TII UAE | 2026-01-04 | 7.8 | — | 80.0 | 0 | — |
| 330 | GPT-4.1 nano | OpenAI | 2025-04-14 | 7.8 | 11.1 | 24.0 | 0 | 0.17 |
| 331 | Gemma 4 E2B (Reasoning)+1 tiers | 2026-04-02 | 7.8 | 7.2 | — | 0 | — | |
| 332 | Ling-flash-2.0 | InclusionAI | 2025-09-17 | 7.8 | — | 65.3 | 0 | 0.25 |
| 333 | Qwen3 Omni 30B A3B (Reasoning) | Alibaba | 2025-09-22 | 7.8 | — | 74.0 | 108 | 0.43 |
| 334 | Solar Pro 3 | Upstage | 2026-04-06 | 7.8 | 16.2 | — | 0 | 0.26 |
| 335 | GPT-4o (Aug '24) | OpenAI | 2024-08-06 | 7.7 | — | — | 0 | 4.38 |
| 336 | Llama 3.3 Instruct 70B | Meta | 2024-12-06 | 7.7 | 11.9 | 7.7 | 89 | 0.71 |
| 337 | Qwen2.5 Instruct 72B | Alibaba | 2024-09-19 | 7.7 | — | 14.0 | 0 | 0.48 |
| 338 | Sonar | Perplexity | 2025-01-21 | 7.7 | — | — | 0 | — |
| 339 | Step3 VL 10B | StepFun | 2026-01-20 | 7.7 | — | — | 0 | — |
| 340 | Devstral Small (Jul '25) | Mistral | 2025-07-10 | 7.6 | — | 29.3 | 0 | — |
| 341 | GLM-4.5V (Reasoning) | Z AI | 2025-08-11 | 7.6 | — | 73.0 | 0 | 0.90 |
| 342 | Mistral Large 2 (Nov '24) | Mistral | 2024-11-18 | 7.6 | — | 14.0 | 0 | — |
| 343 | QwQ 32B-Preview | Alibaba | 2024-11-27 | 7.6 | — | — | 0 | — |
| 344 | Qwen3 30B A3B (Reasoning) | Alibaba | 2025-04-28 | 7.6 | — | 72.3 | 0 | 0.75 |
| 345 | Sonar Pro | Perplexity | 2025-01-21 | 7.6 | — | — | 0 | — |
| 346 | Devstral Small 2 | Mistral | 2025-12-09 | 7.5 | 29.3 | 34.3 | 145 | — |
| 347 | ERNIE 4.5 300B A47B | Baidu | 2025-06-30 | 7.5 | — | 41.3 | 0 | 0.48 |
| 348 | Hermes 4 - Llama-3.1 405B (Reasoning) | Nous Research | 2025-08-27 | 7.5 | — | 69.7 | 43 | 1.50 |
| 349 | Llama 3.1 Nemotron Ultra 253B v1 (Reasoning) | NVIDIA | 2025-04-07 | 7.5 | — | 63.7 | 0 | — |
| 350 | NVIDIA Nemotron Nano 12B v2 VL (Reasoning) | NVIDIA | 2025-10-28 | 7.5 | — | 75.0 | 131 | 0.30 |
| 351 | Qwen3 30B A3B 2507 Instruct | Alibaba | 2025-07-29 | 7.5 | — | 66.3 | 0 | 0.35 |
| 352 | Solar Pro 2 (Reasoning) | Upstage | 2025-07-09 | 7.5 | — | 61.3 | 0 | — |
| 353 | Gemini 2.0 Flash-Lite (Feb '25) | 2025-02-25 | 7.4 | — | — | 0 | — | |
| 354 | Granite 4.1 30B | IBM | 2026-04-29 | 7.4 | 10.4 | — | 0 | — |
| 355 | Hermes 4 - Llama-3.1 405B (Non-reasoning) | Nous Research | 2025-08-27 | 7.4 | — | 15.3 | 44 | 1.50 |
| 356 | Llama Nemotron Super 49B v1.5 (Non-reasoning) | NVIDIA | 2025-07-25 | 7.4 | — | 8.0 | 140 | 0.40 |
| 357 | NVIDIA Nemotron 3 Nano 4B | NVIDIA | 2026-03-16 | 7.4 | 8.0 | — | 0 | — |
| 358 | NVIDIA Nemotron Nano 9B V2 (Reasoning) | NVIDIA | 2025-08-18 | 7.4 | — | 69.7 | 104 | 0.07 |
| 359 | GPT-4o (May '24) | OpenAI | 2024-05-13 | 7.3 | 24.2 | — | 0 | 7.50 |
| 360 | Gemini 2.0 Flash-Lite (Preview) | 2025-02-05 | 7.3 | — | — | 0 | — | |
| 361 | Kimi Linear 48B A3B Instruct | Kimi | 2025-10-30 | 7.3 | — | 36.3 | 0 | — |
| 362 | Llama 3.1 Instruct 405B | Meta | 2024-07-23 | 7.3 | — | 3.0 | 0 | — |
| 363 | Llama 3.1 Nemotron Nano 4B v1.1 (Reasoning) | NVIDIA | 2025-05-20 | 7.3 | — | 50.0 | 0 | — |
| 364 | Llama 3.3 Nemotron Super 49B v1 (Non-reasoning) | NVIDIA | 2025-03-18 | 7.3 | — | 7.7 | 0 | — |
| 365 | Qwen3 32B (Non-reasoning) | Alibaba | 2025-04-28 | 7.3 | — | 19.7 | 0 | 0.28 |
| 366 | Qwen3 8B (Reasoning) | Alibaba | 2025-04-28 | 7.3 | 9.0 | 19.0 | 0 | 0.66 |
| 367 | Qwen3 VL 8B Instruct | Alibaba | 2025-10-14 | 7.3 | — | 27.3 | 0 | 0.31 |
| 368 | Claude 3.5 Sonnet (June '24) | Anthropic | 2024-06-21 | 7.2 | 26.0 | — | 0 | 6.00 |
| 369 | GPT-4o (ChatGPT) | OpenAI | 2025-02-15 | 7.2 | — | — | 0 | — |
| 370 | LFM2.5-8B-A1B | Liquid AI | 2026-05-28 | 7.2 | — | — | 0 | — |
| 371 | Llama 3.1 Tulu3 405B | Allen Institute for AI | 2025-01-30 | 7.2 | — | — | 0 | — |
| 372 | Qwen3 4B (Reasoning) | Alibaba | 2025-04-28 | 7.2 | — | 22.3 | 0 | — |
| 373 | Ring-flash-2.0 | InclusionAI | 2025-09-19 | 7.2 | — | 83.7 | 0 | 0.25 |
| 374 | Gemini 1.5 Flash (Sep '24) | 2024-09-24 | 7.1 | — | — | 0 | — | |
| 375 | Grok 2 (Dec '24) | SpaceXAI | 2024-12-12 | 7.1 | — | — | 0 | — |
| 376 | Mistral Small 3.1 | Mistral | 2025-03-17 | 7.1 | 26.3 | 3.7 | 0 | 0.15 |
| 377 | Olmo 3.1 32B Think | Allen Institute for AI | 2025-12-12 | 7.1 | — | 77.3 | 0 | — |
| 378 | Pixtral Large | Mistral | 2024-11-18 | 7.1 | — | 2.3 | 0 | — |
| 379 | Command A | Cohere | 2025-03-13 | 7.0 | — | 13.0 | 60 | 4.38 |
| 380 | GPT-4 Turbo | OpenAI | 2023-11-06 | 7.0 | 21.5 | — | 0 | 15.00 |
| 381 | Nova Pro | Amazon | 2024-12-03 | 7.0 | — | 7.0 | 0 | 1.40 |
| 382 | Qwen3 VL 4B (Reasoning) | Alibaba | 2025-10-14 | 7.0 | — | 25.7 | 0 | — |
| 383 | Solar Pro 2 (Non-reasoning) | Upstage | 2025-07-09 | 7.0 | — | 30.0 | 0 | — |
| 384 | Grok Beta | SpaceXAI | 2024-08-13 | 6.9 | — | — | 0 | — |
| 385 | Llama 3.1 Instruct 8B | Meta | 2024-07-23 | 6.9 | 5.4 | 4.3 | 0 | 0.03 |
| 386 | Llama 3.1 Nemotron Instruct 70B | NVIDIA | 2024-10-15 | 6.9 | — | 11.0 | 168 | 1.20 |
| 387 | Qwen2.5 Instruct 32B | Alibaba | 2024-09-19 | 6.9 | — | — | 0 | — |
| 388 | Qwen3.5 2B (Reasoning)+1 tiers | Alibaba | 2026-03-02 | 6.9 | 2.9 | — | 0 | — |
| 389 | Mistral Large 2 (Jul '24) | Mistral | 2024-07-24 | 6.8 | — | 0.0 | 0 | 3.00 |
| 390 | NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) | NVIDIA | 2025-12-15 | 6.8 | — | 13.3 | 177 | 0.09 |
| 391 | NVIDIA Nemotron Nano 9B V2 (Non-reasoning) | NVIDIA | 2025-08-18 | 6.8 | — | 62.3 | 153 | 0.09 |
| 392 | GLM-4.5V (Non-reasoning) | Z AI | 2025-08-11 | 6.7 | — | 15.3 | 0 | 0.90 |
| 393 | GPT-4 | OpenAI | 2023-03-14 | 6.7 | 13.1 | — | 0 | 37.50 |
| 394 | GPT-4o mini | OpenAI | 2024-07-18 | 6.7 | 11.4 | 14.7 | 0 | 0.26 |
| 395 | Gemini 2.5 Flash-Lite (Non-reasoning) | 2025-06-17 | 6.7 | — | 35.3 | 0 | 0.17 | |
| 396 | Hermes 4 - Llama-3.1 70B (Non-reasoning) | Nous Research | 2025-08-27 | 6.7 | — | 11.3 | 0 | — |
| 397 | Mistral Small 3 | Mistral | 2025-01-30 | 6.7 | — | 4.3 | 0 | 0.15 |
| 398 | Nova Lite | Amazon | 2024-12-03 | 6.7 | — | 7.0 | 0 | 0.10 |
| 399 | Qwen2.5 Coder Instruct 32B | Alibaba | 2024-11-11 | 6.7 | — | — | 0 | — |
| 400 | Qwen3 14B (Non-reasoning) | Alibaba | 2025-04-28 | 6.7 | — | 58.0 | 0 | 0.61 |
| 401 | Qwen3 4B 2507 Instruct | Alibaba | 2025-08-06 | 6.7 | — | 52.3 | 0 | — |
| 402 | DeepSeek-V2.5 | DeepSeek | 2024-09-06 | 6.6 | — | — | 0 | — |
| 403 | DeepSeek-V2.5 (Dec '24) | DeepSeek | 2024-12-10 | 6.6 | — | — | 0 | — |
| 404 | Gemini 2.0 Flash Thinking Experimental (Dec '24) | 2024-12-19 | 6.6 | — | — | 0 | — | |
| 405 | Granite 4.1 8B | IBM | 2026-04-29 | 6.6 | 9.5 | — | 0 | 0.06 |
| 406 | Llama 3.1 Instruct 70B | Meta | 2024-07-23 | 6.6 | — | 4.0 | 0 | 0.56 |
| 407 | Qwen3 30B A3B (Non-reasoning) | Alibaba | 2025-04-28 | 6.6 | — | 21.7 | 0 | 0.35 |
| 408 | Qwen3 4B (Non-reasoning) | Alibaba | 2025-04-28 | 6.6 | — | — | 0 | — |
| 409 | Sarvam 30B (high) | Sarvam | 2026-03-06 | 6.6 | — | — | 0 | 0.05 |
| 410 | DeepSeek R1 Distill Llama 8B | DeepSeek | 2025-01-20 | 6.5 | — | 41.3 | 0 | — |
| 411 | Mistral Saba | Mistral | 2025-02-17 | 6.5 | — | — | 0 | — |
| 412 | Olmo 3 32B Think | Allen Institute for AI | 2025-11-20 | 6.5 | — | 73.7 | 0 | — |
| 413 | Olmo 3.1 32B Instruct | Allen Institute for AI | 2026-01-13 | 6.5 | — | — | 0 | — |
| 414 | Gemini 1.5 Pro (May '24) | 2024-05-15 | 6.4 | 19.8 | — | 0 | — | |
| 415 | Llama 3.2 Instruct 90B (Vision) | Meta | 2024-09-25 | 6.4 | — | — | 0 | — |
| 416 | Qwen2.5 Turbo | Alibaba | 2024-11-18 | 6.4 | — | — | 0 | 0.09 |
| 417 | R1 1776 | Perplexity | 2025-02-18 | 6.4 | — | — | 0 | — |
| 418 | Reka Flash (Sep '24) | Reka AI | 2024-10-04 | 6.4 | — | — | 0 | 0.35 |
| 419 | Solar Mini | Upstage | 2024-01-25 | 6.4 | — | — | 0 | 0.15 |
| 420 | Celeris-1 | Celeris | 2026-07-24 | 6.3 | 14.4 | — | 2011 | 0.33 |
| 421 | EXAONE 4.0 32B (Non-reasoning) | LG AI Research | 2025-07-15 | 6.3 | — | 39.3 | 0 | — |
| 422 | Grok-1 | SpaceXAI | 2024-03-17 | 6.3 | — | — | 0 | — |
| 423 | Phi-4 Mini Instruct | Microsoft | 2024-02-26 | 6.3 | 3.8 | 6.7 | 46 | — |
| 424 | Qwen2 Instruct 72B | Alibaba | 2024-06-07 | 6.3 | — | — | 0 | — |
| 425 | Gemini 1.5 Flash-8B | 2024-10-03 | 6.2 | — | — | 0 | — | |
| 426 | DeepHermes 3 - Mistral 24B Preview (Non-reasoning) | Nous Research | 2025-03-13 | 6.1 | — | — | 0 | — |
| 427 | Jamba 1.7 Large | AI21 Labs | 2025-07-07 | 6.1 | — | 2.3 | 0 | — |
| 428 | Qwen3.5 0.8B (Reasoning)+1 tiers | Alibaba | 2026-03-02 | 6.1 | 0.0 | — | 0 | — |
| 429 | DeepSeek-Coder-V2 | DeepSeek | 2024-06-17 | 6.0 | — | — | 0 | — |
| 430 | Granite 4.0 H Small | IBM | 2025-09-22 | 6.0 | — | 13.7 | 13 | 0.11 |
| 431 | Hermes 3 - Llama-3.1 70B | Nous Research | 2024-08-15 | 6.0 | — | — | 0 | 0.70 |
| 432 | Jamba 1.5 Large | AI21 Labs | 2024-08-22 | 6.0 | — | — | 0 | 3.50 |
| 433 | Jamba 1.6 Large | AI21 Labs | 2025-03-06 | 6.0 | — | — | 0 | — |
| 434 | Ministral 3 14B | Mistral | 2025-12-02 | 6.0 | 14.4 | 30.0 | 92 | 0.20 |
| 435 | OLMo 2 32B | Allen Institute for AI | 2025-03-13 | 6.0 | — | 3.3 | 0 | — |
| 436 | Qwen3 8B (Non-reasoning) | Alibaba | 2025-04-28 | 6.0 | — | 24.3 | 0 | 0.31 |
| 437 | Qwen3 Omni 30B A3B Instruct | Alibaba | 2025-09-22 | 6.0 | — | 52.3 | 107 | 0.43 |
| 438 | Claude 3 Sonnet | Anthropic | 2024-03-04 | 5.9 | — | — | 0 | — |
| 439 | Gemini 1.5 Flash (May '24) | 2024-05-14 | 5.9 | — | — | 0 | — | |
| 440 | Granite 4.1 3B | IBM | 2026-04-29 | 5.9 | 4.7 | — | 0 | — |
| 441 | LFM2 24B A2B | Liquid AI | 2026-02-25 | 5.9 | — | — | 0 | — |
| 442 | Nova Micro | Amazon | 2024-12-03 | 5.9 | — | 6.0 | 224 | 0.06 |
| 443 | Phi-4 | Microsoft | 2024-12-12 | 5.9 | — | 18.0 | 35 | 0.22 |
| 444 | Gemini 1.0 Ultra | 2023-12-06 | 5.8 | 17.6 | — | 0 | — | |
| 445 | Gemma 3n E4B Instruct Preview (May '25) | 2025-05-20 | 5.8 | — | — | 0 | — | |
| 446 | Mistral Large (Feb '24) | Mistral | 2024-02-26 | 5.8 | — | — | 0 | 6.00 |
| 447 | Mistral Small (Sep '24) | Mistral | 2024-09-17 | 5.8 | — | — | 0 | 0.30 |
| 448 | NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) | NVIDIA | 2025-10-28 | 5.8 | — | 26.7 | 198 | 0.30 |
| 449 | Phi-3 Mini Instruct 3.8B | Microsoft | 2024-04-23 | 5.8 | — | 0.3 | 0 | — |
| 450 | Phi-4 Multimodal Instruct | Microsoft | 2025-02-26 | 5.8 | — | — | 17 | — |
| 451 | Qwen2.5 Coder Instruct 7B | Alibaba | 2024-09-19 | 5.8 | — | — | 0 | — |
| 452 | Jamba Reasoning 3B | AI21 Labs | 2025-10-08 | 5.7 | — | 10.7 | 0 | — |
| 453 | Llama 2 Chat 7B | Meta | 2023-07-18 | 5.7 | — | — | 0 | 0.10 |
| 454 | Llama 3.2 Instruct 3B | Meta | 2024-09-25 | 5.7 | — | 3.3 | 0 | — |
| 455 | MiniCPM-V 4.6 1.3B | OpenBMB | 2026-05-11 | 5.7 | 0.7 | — | 0 | — |
| 456 | Mixtral 8x22B Instruct | Mistral | 2024-04-17 | 5.7 | — | — | 0 | — |
| 457 | Qwen1.5 Chat 110B | Alibaba | 2024-04-25 | 5.7 | — | — | 0 | — |
| 458 | Qwen3 VL 4B Instruct | Alibaba | 2025-10-14 | 5.7 | — | 37.0 | 0 | — |
| 459 | Claude 2.1 | Anthropic | 2023-11-21 | 5.6 | 14.0 | — | 0 | — |
| 460 | Claude 3 Haiku | Anthropic | 2024-03-04 | 5.6 | — | — | 0 | 0.50 |
| 461 | Molmo 7B-D | Allen Institute for AI | 2024-09-25 | 5.6 | — | 0.0 | 0 | — |
| 462 | OLMo 2 7B | Allen Institute for AI | 2024-11-26 | 5.6 | — | 0.7 | 0 | — |
| 463 | Olmo 3 7B Think | Allen Institute for AI | 2025-11-20 | 5.6 | — | 70.7 | 0 | — |
| 464 | Reka Flash 3 | Reka AI | 2025-03-10 | 5.6 | — | 33.7 | 92 | 0.35 |
| 465 | Claude 2.0 | Anthropic | 2023-07-11 | 5.5 | 12.9 | — | 0 | — |
| 466 | DeepSeek R1 Distill Qwen 1.5B | DeepSeek | 2025-01-20 | 5.5 | — | 22.0 | 0 | — |
| 467 | DeepSeek-V2-Chat | DeepSeek | 2024-05-06 | 5.5 | — | — | 0 | — |
| 468 | GPT-3.5 Turbo | OpenAI | 2022-11-30 | 5.5 | 10.7 | — | 0 | 0.75 |
| 469 | Ling-mini-2.0 | InclusionAI | 2025-09-09 | 5.5 | — | 49.3 | 0 | — |
| 470 | Llama 3 Instruct 70B | Meta | 2024-04-18 | 5.5 | — | — | 0 | 1.18 |
| 471 | Ministral 3 8B | Mistral | 2025-12-02 | 5.5 | 9.7 | 31.7 | 103 | 0.15 |
| 472 | Mistral Medium | Mistral | 2023-12-11 | 5.5 | — | — | 0 | 3.00 |
| 473 | Mistral Small (Feb '24) | Mistral | 2024-02-26 | 5.5 | — | — | 0 | 0.26 |
| 474 | Arctic Instruct | Snowflake | 2024-04-24 | 5.4 | — | — | 0 | — |
| 475 | LFM 40B | Liquid AI | 2024-09-30 | 5.4 | — | — | 0 | — |
| 476 | Llama 3.2 Instruct 11B (Vision) | Meta | 2024-09-25 | 5.4 | — | 1.7 | 20 | 0.34 |
| 477 | PALM-2 | 2023-05-10 | 5.4 | 4.6 | — | 0 | — | |
| 478 | Qwen Chat 72B | Alibaba | 2023-11-30 | 5.4 | — | — | 0 | — |
| 479 | Command-R+ (Apr '24) | Cohere | 2024-04-04 | 5.3 | — | — | 0 | — |
| 480 | DBRX Instruct | Databricks | 2024-03-27 | 5.3 | — | — | 0 | — |
| 481 | DeepSeek Coder V2 Lite Instruct | DeepSeek | 2024-06-17 | 5.3 | — | — | 0 | — |
| 482 | DeepSeek LLM 67B Chat (V1) | DeepSeek | 2023-11-29 | 5.3 | — | — | 0 | — |
| 483 | Exaone 4.0 1.2B (Reasoning) | LG AI Research | 2025-07-15 | 5.3 | — | 50.3 | 0 | — |
| 484 | Gemini 1.0 Pro | 2023-12-06 | 5.3 | — | — | 0 | — | |
| 485 | Llama 2 Chat 13B | Meta | 2023-07-18 | 5.3 | — | — | 0 | — |
| 486 | Llama 2 Chat 70B | Meta | 2023-07-18 | 5.3 | — | — | 0 | — |
| 487 | OpenChat 3.5 (1210) | OpenChat | 2023-12-18 | 5.3 | — | — | 0 | — |
| 488 | Sarvam M (Reasoning) | Sarvam | 2025-05-23 | 5.3 | — | — | 0 | — |
| 489 | Exaone 4.0 1.2B (Non-reasoning) | LG AI Research | 2025-07-15 | 5.2 | — | 24.0 | 0 | — |
| 490 | Granite 4.0 H 1B | IBM | 2025-10-28 | 5.2 | — | 6.3 | 0 | — |
| 491 | Jamba 1.5 Mini | AI21 Labs | 2024-08-22 | 5.2 | — | — | 0 | 0.25 |
| 492 | Jamba 1.6 Mini | AI21 Labs | 2025-03-06 | 5.2 | — | — | 0 | — |
| 493 | Jamba 1.7 Mini | AI21 Labs | 2025-07-07 | 5.2 | — | 0.3 | 0 | — |
| 494 | LFM2 2.6B | Liquid AI | 2025-09-23 | 5.2 | — | 8.3 | 0 | — |
| 495 | LFM2.5-1.2B-Instruct | Liquid AI | 2026-01-05 | 5.2 | — | — | 0 | — |
| 496 | LFM2.5-1.2B-Thinking | Liquid AI | 2026-01-20 | 5.2 | — | — | 0 | — |
| 497 | Olmo 3 7B Instruct | Allen Institute for AI | 2025-11-20 | 5.2 | — | 41.3 | 0 | 0.12 |
| 498 | Qwen3 1.7B (Reasoning) | Alibaba | 2025-04-28 | 5.2 | — | 38.7 | 0 | — |
| 499 | Apertus 70B Instruct | Swiss AI Initiative | 2025-09-02 | 5.1 | — | — | 0 | 1.34 |
| 500 | DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning) | Nous Research | 2025-02-13 | 5.1 | — | — | 0 | — |
| 501 | Gemma 3 270M | 2025-08-14 | 5.1 | — | 2.3 | 0 | — | |
| 502 | Granite 4.0 Micro | IBM | 2025-09-22 | 5.1 | — | 6.0 | 0 | — |
| 503 | Mixtral 8x7B Instruct | Mistral | 2023-12-11 | 5.1 | — | — | 0 | 0.51 |
| 504 | Claude Instant | Anthropic | 2023-03-14 | 5.0 | 7.8 | — | 0 | — |
| 505 | Command-R (Mar '24) | Cohere | 2024-03-12 | 5.0 | — | — | 0 | — |
| 506 | Granite 4.0 1B | IBM | 2025-10-28 | 5.0 | — | 6.3 | 0 | — |
| 507 | Llama 65B | Meta | 2023-02-24 | 5.0 | — | — | 0 | — |
| 508 | Mistral 7B Instruct | Mistral | 2023-09-27 | 5.0 | — | — | 0 | 0.25 |
| 509 | Molmo2-8B | Allen Institute for AI | 2025-12-11 | 5.0 | — | — | 0 | — |
| 510 | Qwen Chat 14B | Alibaba | 2023-09-25 | 5.0 | — | — | 0 | — |
| 511 | Gemma 3 27B Instruct | 2025-03-12 | 4.9 | 10.1 | 20.7 | 0 | — | |
| 512 | Granite 3.3 8B (Non-reasoning) | IBM | 2025-04-16 | 4.9 | — | 6.7 | 0 | 0.09 |
| 513 | LFM2 8B A1B | Liquid AI | 2025-10-07 | 4.9 | — | 25.3 | 0 | — |
| 514 | Qwen3 1.7B (Non-reasoning) | Alibaba | 2025-04-28 | 4.9 | — | 7.3 | 0 | — |
| 515 | Apertus 8B Instruct | Swiss AI Initiative | 2025-09-02 | 4.8 | — | — | 0 | 0.12 |
| 516 | Gemma 3 1B Instruct | 2025-03-13 | 4.8 | — | 3.3 | 0 | — | |
| 517 | Gemma 3 4B Instruct | 2025-03-12 | 4.8 | 2.7 | 12.7 | 0 | — | |
| 518 | Gemma 3n E2B Instruct | 2025-06-26 | 4.8 | — | 10.3 | 0 | — | |
| 519 | Gemma 3n E4B Instruct | 2025-06-26 | 4.8 | 3.2 | 14.3 | 0 | — | |
| 520 | Granite 4.0 350M | IBM | 2025-10-28 | 4.8 | — | 0.0 | 0 | — |
| 521 | Granite 4.0 H 350M | IBM | 2025-10-28 | 4.8 | — | 1.3 | 0 | — |
| 522 | LFM2 1.2B | Liquid AI | 2025-07-10 | 4.8 | — | 3.3 | 0 | — |
| 523 | LFM2.5-VL-1.6B | Liquid AI | 2026-01-05 | 4.8 | — | — | 0 | — |
| 524 | Llama 3 Instruct 8B | Meta | 2024-04-18 | 4.8 | — | — | 0 | 0.07 |
| 525 | Llama 3.2 Instruct 1B | Meta | 2024-09-25 | 4.8 | — | 0.0 | 0 | — |
| 526 | Ministral 3 3B | Mistral | 2025-12-02 | 4.8 | 4.8 | 22.0 | 237 | 0.10 |
| 527 | Qwen3 0.6B (Non-reasoning) | Alibaba | 2025-04-28 | 4.8 | — | 10.3 | 0 | — |
| 528 | Qwen3 0.6B (Reasoning) | Alibaba | 2025-04-28 | 4.8 | — | 18.0 | 0 | — |
| 529 | Tiny Aya Global | Cohere | 2026-02-17 | 4.8 | — | — | 132 | — |
| 530 | Gemma 3 12B Instruct | 2025-03-12 | 3.8 | 5.8 | 18.3 | 0 | — | |
| 531 | K2 Horizon 0.9BNEW | Institute of Foundation Models | 2026-09-03 | 3.0 | 3.4 | — | 0 | — |
| 532 | Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 5 Fallback) | Anthropic | — | — | — | — | 0 | 20.00 |
| 533 | Cogito v2.1 (Reasoning) | Deep Cogito | 2025-11-18 | — | — | 72.7 | 0 | 1.25 |
| 534 | GPT-3.5 Turbo (0613) | OpenAI | 2023-06-13 | — | — | — | 0 | — |
| 535 | GPT-4o Realtime (Dec '24) | OpenAI | 2024-12-17 | — | — | — | 0 | — |
| 536 | GPT-4o mini Realtime (Dec '24) | OpenAI | 2024-12-17 | — | — | — | 0 | — |
| 537 | GPT-5.4 Pro (xhigh) | OpenAI | 2026-03-05 | — | — | — | 0 | 67.50 |
| 538 | GPT-5.5 Pro (xhigh) | OpenAI | 2026-04-23 | — | — | — | 0 | — |
| 539 | Gemini 3 Deep Think | 2026-02-05 | — | — | — | 0 | — | |
| 540 | Mi:dm K 2.5 Pro Preview | Korea Telecom | 2025-12-11 | — | — | 78.7 | 0 | — |
FAQ
What is the Artificial Analysis Intelligence Index?
It's a composite benchmark score published by Artificial Analysis that blends results across multiple public reasoning, knowledge and instruction-following evaluations (e.g. MMLU-Pro, GPQA Diamond, AIME-style math sets) into one number per model, so models can be ranked on general capability rather than any single test.
How often is this leaderboard updated?
Daily. The underlying data is synced from Artificial Analysis and this page is rebuilt automatically, so all 540 rankings reflect the latest published benchmark runs.
Why are some models missing from individual model rows even though they exist?
They aren't missing — a model released at several reasoning-effort tiers (low/medium/high/xhigh/…) is merged into a single ranked row showing its best-scoring tier, with a "+N tiers" badge noting how many effort levels were combined. This avoids one model occupying multiple leaderboard slots.
How is this different from the model pricing table?
This page ranks purely by capability (intelligence, coding, math, speed). The /models/ page is built for comparing and filtering by price — sort by input, output or blended $/1M tokens, filter by creator or a price ceiling. Use this page to find the strongest model; use that page to find the cheapest one that's good enough.
Why do some rows show "—" instead of a price?
Artificial Analysis has not published pricing for that model (commonly a newly benchmarked or not-yet-generally-available release). We never guess a price or display an unpublished price as free — it's shown as "—" instead.
What do the Coding and Math columns measure?
They're the same Artificial Analysis composite scores, but restricted to coding benchmarks (agentic and single-turn code generation/repair tasks) and math benchmarks (competition-style problem sets), respectively — useful when your workload leans heavily toward one skill rather than general intelligence.
Want the full price breakdown instead of a ranking? See the model pricing table. Want to know what a specific prompt actually costs? Use the token counter.
Need a cheaper way to call these APIs?
Sponsored — third-party API resellers, not official providers. Synced from howtok.net.-
NovaAPI
WeChat PayAlipayUSDTVISA
NovaAPI is a global AI token relay platform offering unified access to OpenAI, Claude, Gemini, and leading Chinese LLMs through low-latency, always-on routing. Pricing is fully transparent, with API calls discounted up to 15% off list rates, and referring a friend earns ongoing rewards.
-
Apimart
WeChat PayAlipayUSDTVISA
A full-modality AI API supply layer aggregating 500+ models — Sora 2, Veo 3.1, Kling 3.0, Seedream 5.0, GPT-5, Claude, and more. Supports prepaid top-ups, dedicated enterprise channels, and revenue-share partnerships, serving relay platforms, AI app developers, and teams building for global markets.
-
DanceHorse
USDTWeChat Pay
An AI video API supply layer connecting to Seedance, HappyHorse, and other video-generation models. Supports prepaid top-ups, distribution codes, white-label API access, and revenue-share partnerships — built for relay platforms, video tool builders, and content teams.
These are paid placements synced from howtok.net's relay-station directory. LLMAPIComparison has no affiliation with the listed services — evaluate them yourself before committing spend.