MATH: Benchmark Detailed Guide
Detailed guide to MATH: category, metrics, sources and applicable models.
Overview
12,500 道竞赛数学题(含 7 个学科:代数、几何、数论、微积分、概率等,5 个难度级别),评估模型的复杂数学推理能力。
Metrics
| Metric | Unit | Direction |
|---|---|---|
| accuracy | % | ↑ Higher is better |
Sources
Model Score Ranking
| # | Model | Vendor | Score |
|---|---|---|---|
| 1 | Qwen2.5 72B | alibaba | 83.1 |
| 2 | Grok-2 | xai | 76.8 |
| 3 | GPT-4o | openai | 76.6 |
| 4 | Claude 3.5 Sonnet (2024-10-22) | anthropic | 75.1 |
| 5 | Llama 3.1 405B | meta | 73.8 |
| 6 | Llama 3.3 70B | meta | 73.8 |
| 7 | Qwen2 72B | alibaba | 71.9 |
| 8 | o1 Preview | openai | 71.5 |
| 9 | Gemini 2.0 Flash | 71.3 | |
| 10 | Claude 3.5 Sonnet | anthropic | 71.1 |
| 11 | Mistral Large 2 | mistral | 71 |
| 12 | GPT-4o (2024-08-06) | openai | 68.2 |
| 13 | GPT-4o (2024-05-13) | openai | 67.1 |
| 14 | Microsoft WizardMath 7B v1 | other | 65.8 |
| 15 | Hermes 3 Llama 3.1 405B | other | 65.7 |
| 16 | DeepSeek V3 | deepseek | 61.6 |
| 17 | Mathstral 7B | mistral | 61 |
| 18 | Jamba 1.5 Large | other | 60 |
| 19 | Yi Vision | other | 58.7 |
| 20 | Gemini 1.5 Pro | 58.5 | |
| 21 | Claude 3 Opus (2024-02-29) | anthropic | 58.2 |
| 22 | Gemini 1.0 Pro | 58.2 | |
| 23 | Gemini 1.0 Ultra | 57.1 | |
| 24 | Gemini 2.0 Flash Thinking | 56.6 | |
| 25 | Gemini 1.5 Pro 002 | 55.9 | |
| 26 | Yi Large | other | 55.9 |
| 27 | o1 | openai | 55.7 |
| 28 | Mistral Medium | mistral | 54.9 |
| 29 | Qwen2 57B | alibaba | 54.8 |
| 30 | Hermes 3 Llama 3.1 70B | other | 54.4 |
| 31 | Qwen1.5 110B | alibaba | 54.1 |
| 32 | DeepSeek Math 7B | deepseek | 53.8 |
| 33 | GLM-4 Plus | other | 53.3 |
| 34 | Gemini 1.5 Flash | 53.2 | |
| 35 | Claude 3.5 Haiku | anthropic | 53 |
| 36 | Llama 3.2 90B Vision | meta | 52.5 |
| 37 | Claude 3 Haiku | anthropic | 52.2 |
| 38 | GPT-4 1106 Preview | openai | 52.1 |
| 39 | GPT-4o mini | openai | 51.8 |
| 40 | Sonar Large | other | 51.6 |
| 41 | GPT-4 Turbo | openai | 50.9 |
| 42 | Sonar Reasoning | other | 50.7 |
| 43 | Yi 1.5 34B | other | 50.4 |
| 44 | DBRX Instruct | other | 50.1 |
| 45 | Llama 3 70B | meta | 50.1 |
| 46 | GPT-4 Vision Preview | openai | 49.7 |
| 47 | DeepSeek V2 Chat | deepseek | 49.3 |
| 48 | Mistral Large | mistral | 49.3 |
| 49 | Command Nightly | cohere | 49.2 |
| 50 | Jamba Instruct | other | 49 |
| 51 | GPT-4 Vision | openai | 48.3 |
| 52 | Jamba 1.5 | other | 48.1 |
| 53 | Llemma 7B | other | 47.5 |
| 54 | Code Llama 7B | meta | 47.1 |
| 55 | Command R (08-2024) | cohere | 47.1 |
| 56 | Sonar Huge | other | 46.8 |
| 57 | Claude 3 Sonnet (2024-02-29) | anthropic | 46.4 |
| 58 | Qwen1.5 32B | alibaba | 46.3 |
| 59 | Qwen1.5 14B | alibaba | 46.2 |
| 60 | StarCoder2 15B | other | 46.2 |
| 61 | Mixtral 8x22B | mistral | 46 |
| 62 | Mistral Small | mistral | 45.8 |
| 63 | StarChat2 15B v0.1 | other | 45.3 |
| 64 | Grok-2 Mini | xai | 45.1 |
| 65 | Nous Hermes 2 Yi 34B | other | 44.9 |
| 66 | Falcon 180B | other | 44.8 |
| 67 | Claude 3 Haiku (2024-03-07) | anthropic | 44.7 |
| 68 | Code Llama 34B | meta | 44.7 |
| 69 | Yi 1.5 6B | other | 44.5 |
| 70 | o1 mini | openai | 44.2 |
| 71 | Command R | cohere | 43.9 |
| 72 | GPT-4 | openai | 43.9 |
| 73 | Mistral Nemo | mistral | 43.6 |
| 74 | Phi-3 Small | other | 43.5 |
| 75 | Command R+ (08-2024) | cohere | 43.3 |
| 76 | Sonar Small | other | 43.3 |
| 77 | OLMo 7B SFT | other | 43.1 |
| 78 | Gemini 1.5 Flash-8B 002 | 43 | |
| 79 | Claude 3 Opus | anthropic | 42.7 |
| 80 | Nous Hermes 2 Mixtral 8x7B | other | 42.6 |
| 81 | Jamba 1.5 Mini | other | 42.5 |
| 82 | Claude 3 Sonnet | anthropic | 42.4 |
| 83 | GPT-4 0125 Preview | openai | 42.3 |
| 84 | Orca 2 13B | other | 42.3 |
| 85 | Yi Large Turbo | other | 42.2 |
| 86 | Qwen2.5 7B | alibaba | 41.9 |
| 87 | Grok-2 Vision | xai | 41.3 |
| 88 | Qwen2 7B | alibaba | 40.7 |
| 89 | Command R7B | cohere | 40.6 |
| 90 | GLM-4 Air | other | 40.6 |
| 91 | Microsoft WizardCoder Python 34B | other | 40.6 |
| 92 | Gemini 1.5 Flash 002 | 40.5 | |
| 93 | NVIDIA Llama 3.1 Nemotron 70B | other | 40.3 |
| 94 | GPT-4 32K | openai | 40.1 |
| 95 | Llama 3.1 Nemotron 70B | meta | 40 |
| 96 | Llama 3 8B | meta | 40 |
| 97 | Gemini 1.0 Flash | 39.9 | |
| 98 | Ministral 8B | mistral | 39.9 |
| 99 | Microsoft WizardLM 2 8x22B | other | 39.8 |
| 100 | Phi-4 | other | 39.6 |
| 101 | Qwen2.5 32B | alibaba | 39.5 |
| 102 | GLM-4 Flash | other | 39.4 |
| 103 | DeepSeek LLM 67B | deepseek | 39.3 |
| 104 | OLMo 2 1124 7B | other | 39.3 |
| 105 | Mistral Small 3 | mistral | 39.2 |
| 106 | Qwen2.5 14B | alibaba | 38.8 |
| 107 | Llama 3.1 70B | meta | 38.5 |
| 108 | DeepSeek Coder V2 | deepseek | 38.2 |
| 109 | Phi-3.5 MoE | other | 38.2 |
| 110 | OLMo 1.7 7B | other | 38.1 |
| 111 | Mixtral 8x7B | mistral | 38 |
| 112 | StableLM 2 12B | other | 37.9 |
| 113 | Code Llama 13B | meta | 37.8 |
| 114 | Nous Hermes 2 Solar 10.7B | other | 37.7 |
| 115 | Qwen1.5 72B | alibaba | 37.6 |
| 116 | StarCoder2 3B | other | 37.4 |
| 117 | WizardLM Team WizardLM 2 8x22B | other | 37.4 |
| 118 | OLMo 7B | other | 37.1 |
| 119 | DBRX Base | other | 36 |
| 120 | Zephyr ORPO 141B Alpha | other | 36 |
| 121 | Command R+ | cohere | 35.6 |
| 122 | Gemma 2 27B | 35.5 | |
| 123 | DeepSeek V2 | deepseek | 35.2 |
| 124 | Gemini 1.5 Flash-8B | 35 | |
| 125 | OLMo 7B Instruct | other | 34.7 |
| 126 | Code Bison | 34.4 | |
| 127 | Phi-3 Medium | other | 34.3 |
| 128 | Yi 1.5 9B | other | 34.1 |
| 129 | Code Llama 70B | meta | 34 |
| 130 | OpenChat 3.6 8B | other | 33.7 |
| 131 | GLM-4V 9B | other | 33.3 |
| 132 | DeepSeek Coder 7B | deepseek | 32.2 |
| 133 | DeepSeek Coder 33B | deepseek | 32.1 |
| 134 | ChatGLM3 6B | other | 31.8 |
| 135 | Zephyr 7B Beta | other | 31.3 |
| 136 | Llama 3.1 8B | meta | 31.1 |
| 137 | Argilla Notus 7B v1 | other | 31 |
| 138 | StableCode 3B | other | 31 |
| 139 | Microsoft WizardLM 2 7B | other | 30.5 |
| 140 | StarCoder2 7B | other | 30.5 |
| 141 | GPT-3.5 Turbo 16K | openai | 30.4 |
| 142 | Llama 3.2 11B Vision | meta | 29.9 |
| 143 | PaLM 2 | 29.9 | |
| 144 | Yi 6B | other | 29.9 |
| 145 | Codestral | mistral | 29.4 |
| 146 | GPT-3.5 | openai | 29.4 |
| 147 | Gemma 2 9B | 28.9 | |
| 148 | Codestral Mamba | mistral | 28.1 |
| 149 | Flan-T5 XL | other | 27.8 |
| 150 | Gemma 2B | 27.5 | |
| 151 | Jurassic-2 Ultra | other | 27.5 |
| 152 | Phi-1.5 | other | 27.2 |
| 153 | Mistral 7B v0.3 | mistral | 26.7 |
| 154 | GLM-4 9B Chat | other | 26.6 |
| 155 | Mistral 7B v0.1 | mistral | 26.5 |
| 156 | StableLM 2 1.6B | other | 26.4 |
| 157 | Llama Guard 2 8B | meta | 26.3 |
| 158 | Llama 2 13B | meta | 25.9 |
| 159 | Mistral 7B v0.2 | mistral | 25.9 |
| 160 | Gemma 7B | 25.7 | |
| 161 | Capybara 1.5B | other | 25.6 |
| 162 | Hermes 3 Llama 3.1 8B | other | 25.5 |
| 163 | Qwen2 1.5B | alibaba | 25.4 |
| 164 | Phi-1 | other | 25.2 |
| 165 | Phi-3 Vision | other | 24.3 |
| 166 | Command Light | cohere | 24.1 |
| 167 | Grok Vision Beta | xai | 24 |
| 168 | Baichuan2 13B Chat | other | 23.4 |
| 169 | Zephyr 7B Alpha | other | 23.2 |
| 170 | Nous Capybara 34B | other | 23.1 |
| 171 | Claude Instant 1 | anthropic | 22.5 |
| 172 | Llama 2 70B | meta | 22.4 |
| 173 | Flan-UL2 | other | 22 |
| 174 | Claude 2 | anthropic | 21.9 |
| 175 | Flan-T5 XXL | other | 21.4 |
| 176 | Falcon 40B | other | 21.3 |
| 177 | MPT 30B | other | 21.2 |
| 178 | Open-Platypus | other | 21 |
| 179 | Phi-2 | other | 20.5 |
| 180 | Yi 34B | other | 20.5 |
| 181 | Jurassic-2 Mid | other | 20.3 |
| 182 | Chat Bison | 19.8 | |
| 183 | Nous Capybara 7B | other | 19.7 |
| 184 | Llama 2 7B | meta | 19.4 |
| 185 | Grok Beta | xai | 19.1 |
| 186 | Baichuan2 7B Chat | other | 18.8 |
| 187 | MPT 7B | other | 18.4 |
| 188 | Phi-3 Mini | other | 17.9 |
| 189 | GPT-3.5 Turbo | openai | 17.8 |
| 190 | Qwen2.5 3B | alibaba | 17 |
| 191 | Mistral Tiny | mistral | 16 |
| 192 | Claude 2.1 | anthropic | 15.5 |
| 193 | Qwen2.5 0.5B | alibaba | 15.4 |
| 194 | TigerBot 70B Chat | other | 15.4 |
| 195 | OpenChat 3.5 1210 | other | 15.2 |
| 196 | StableLM Zephyr 3B | other | 14.9 |
| 197 | Text Bison | 12.1 | |
| 198 | StableLM 3 4B | other | 11.5 |
| 199 | Llama 3.2 1B | meta | 11.4 |
| 200 | Ministral 3B | mistral | 10.1 |
| 201 | Llama 3.2 3B | meta | 10 |
| 202 | Phi-3.5 Mini | other | 9.6 |
| 203 | Qwen2.5 1.5B | alibaba | 9.5 |
| 204 | Llama Guard 3 8B | meta | 8.4 |
| 205 | Embed English v3 | cohere | 0 |
MATH
Description
12,500 道竞赛数学题(含 7 个学科:代数、几何、数论、微积分、概率等,5 个难度级别),评估模型的复杂数学推理能力。
Core Specifications
| Category | License | Last Updated |
|---|---|---|
| math | MIT | 2021-01-01 |
Benchmark
| Unit |
|---|
| % |