Overview

12,500 道竞赛数学题(含 7 个学科:代数、几何、数论、微积分、概率等,5 个难度级别),评估模型的复杂数学推理能力。

Metrics

Metric Unit Direction
accuracy % ↑ Higher is better

Sources

Model Score Ranking

# Model Vendor Score
1 Qwen2.5 72B alibaba 83.1
2 Grok-2 xai 76.8
3 GPT-4o openai 76.6
4 Claude 3.5 Sonnet (2024-10-22) anthropic 75.1
5 Llama 3.1 405B meta 73.8
6 Llama 3.3 70B meta 73.8
7 Qwen2 72B alibaba 71.9
8 o1 Preview openai 71.5
9 Gemini 2.0 Flash google 71.3
10 Claude 3.5 Sonnet anthropic 71.1
11 Mistral Large 2 mistral 71
12 GPT-4o (2024-08-06) openai 68.2
13 GPT-4o (2024-05-13) openai 67.1
14 Microsoft WizardMath 7B v1 other 65.8
15 Hermes 3 Llama 3.1 405B other 65.7
16 DeepSeek V3 deepseek 61.6
17 Mathstral 7B mistral 61
18 Jamba 1.5 Large other 60
19 Yi Vision other 58.7
20 Gemini 1.5 Pro google 58.5
21 Claude 3 Opus (2024-02-29) anthropic 58.2
22 Gemini 1.0 Pro google 58.2
23 Gemini 1.0 Ultra google 57.1
24 Gemini 2.0 Flash Thinking google 56.6
25 Gemini 1.5 Pro 002 google 55.9
26 Yi Large other 55.9
27 o1 openai 55.7
28 Mistral Medium mistral 54.9
29 Qwen2 57B alibaba 54.8
30 Hermes 3 Llama 3.1 70B other 54.4
31 Qwen1.5 110B alibaba 54.1
32 DeepSeek Math 7B deepseek 53.8
33 GLM-4 Plus other 53.3
34 Gemini 1.5 Flash google 53.2
35 Claude 3.5 Haiku anthropic 53
36 Llama 3.2 90B Vision meta 52.5
37 Claude 3 Haiku anthropic 52.2
38 GPT-4 1106 Preview openai 52.1
39 GPT-4o mini openai 51.8
40 Sonar Large other 51.6
41 GPT-4 Turbo openai 50.9
42 Sonar Reasoning other 50.7
43 Yi 1.5 34B other 50.4
44 DBRX Instruct other 50.1
45 Llama 3 70B meta 50.1
46 GPT-4 Vision Preview openai 49.7
47 DeepSeek V2 Chat deepseek 49.3
48 Mistral Large mistral 49.3
49 Command Nightly cohere 49.2
50 Jamba Instruct other 49
51 GPT-4 Vision openai 48.3
52 Jamba 1.5 other 48.1
53 Llemma 7B other 47.5
54 Code Llama 7B meta 47.1
55 Command R (08-2024) cohere 47.1
56 Sonar Huge other 46.8
57 Claude 3 Sonnet (2024-02-29) anthropic 46.4
58 Qwen1.5 32B alibaba 46.3
59 Qwen1.5 14B alibaba 46.2
60 StarCoder2 15B other 46.2
61 Mixtral 8x22B mistral 46
62 Mistral Small mistral 45.8
63 StarChat2 15B v0.1 other 45.3
64 Grok-2 Mini xai 45.1
65 Nous Hermes 2 Yi 34B other 44.9
66 Falcon 180B other 44.8
67 Claude 3 Haiku (2024-03-07) anthropic 44.7
68 Code Llama 34B meta 44.7
69 Yi 1.5 6B other 44.5
70 o1 mini openai 44.2
71 Command R cohere 43.9
72 GPT-4 openai 43.9
73 Mistral Nemo mistral 43.6
74 Phi-3 Small other 43.5
75 Command R+ (08-2024) cohere 43.3
76 Sonar Small other 43.3
77 OLMo 7B SFT other 43.1
78 Gemini 1.5 Flash-8B 002 google 43
79 Claude 3 Opus anthropic 42.7
80 Nous Hermes 2 Mixtral 8x7B other 42.6
81 Jamba 1.5 Mini other 42.5
82 Claude 3 Sonnet anthropic 42.4
83 GPT-4 0125 Preview openai 42.3
84 Orca 2 13B other 42.3
85 Yi Large Turbo other 42.2
86 Qwen2.5 7B alibaba 41.9
87 Grok-2 Vision xai 41.3
88 Qwen2 7B alibaba 40.7
89 Command R7B cohere 40.6
90 GLM-4 Air other 40.6
91 Microsoft WizardCoder Python 34B other 40.6
92 Gemini 1.5 Flash 002 google 40.5
93 NVIDIA Llama 3.1 Nemotron 70B other 40.3
94 GPT-4 32K openai 40.1
95 Llama 3.1 Nemotron 70B meta 40
96 Llama 3 8B meta 40
97 Gemini 1.0 Flash google 39.9
98 Ministral 8B mistral 39.9
99 Microsoft WizardLM 2 8x22B other 39.8
100 Phi-4 other 39.6
101 Qwen2.5 32B alibaba 39.5
102 GLM-4 Flash other 39.4
103 DeepSeek LLM 67B deepseek 39.3
104 OLMo 2 1124 7B other 39.3
105 Mistral Small 3 mistral 39.2
106 Qwen2.5 14B alibaba 38.8
107 Llama 3.1 70B meta 38.5
108 DeepSeek Coder V2 deepseek 38.2
109 Phi-3.5 MoE other 38.2
110 OLMo 1.7 7B other 38.1
111 Mixtral 8x7B mistral 38
112 StableLM 2 12B other 37.9
113 Code Llama 13B meta 37.8
114 Nous Hermes 2 Solar 10.7B other 37.7
115 Qwen1.5 72B alibaba 37.6
116 StarCoder2 3B other 37.4
117 WizardLM Team WizardLM 2 8x22B other 37.4
118 OLMo 7B other 37.1
119 DBRX Base other 36
120 Zephyr ORPO 141B Alpha other 36
121 Command R+ cohere 35.6
122 Gemma 2 27B google 35.5
123 DeepSeek V2 deepseek 35.2
124 Gemini 1.5 Flash-8B google 35
125 OLMo 7B Instruct other 34.7
126 Code Bison google 34.4
127 Phi-3 Medium other 34.3
128 Yi 1.5 9B other 34.1
129 Code Llama 70B meta 34
130 OpenChat 3.6 8B other 33.7
131 GLM-4V 9B other 33.3
132 DeepSeek Coder 7B deepseek 32.2
133 DeepSeek Coder 33B deepseek 32.1
134 ChatGLM3 6B other 31.8
135 Zephyr 7B Beta other 31.3
136 Llama 3.1 8B meta 31.1
137 Argilla Notus 7B v1 other 31
138 StableCode 3B other 31
139 Microsoft WizardLM 2 7B other 30.5
140 StarCoder2 7B other 30.5
141 GPT-3.5 Turbo 16K openai 30.4
142 Llama 3.2 11B Vision meta 29.9
143 PaLM 2 google 29.9
144 Yi 6B other 29.9
145 Codestral mistral 29.4
146 GPT-3.5 openai 29.4
147 Gemma 2 9B google 28.9
148 Codestral Mamba mistral 28.1
149 Flan-T5 XL other 27.8
150 Gemma 2B google 27.5
151 Jurassic-2 Ultra other 27.5
152 Phi-1.5 other 27.2
153 Mistral 7B v0.3 mistral 26.7
154 GLM-4 9B Chat other 26.6
155 Mistral 7B v0.1 mistral 26.5
156 StableLM 2 1.6B other 26.4
157 Llama Guard 2 8B meta 26.3
158 Llama 2 13B meta 25.9
159 Mistral 7B v0.2 mistral 25.9
160 Gemma 7B google 25.7
161 Capybara 1.5B other 25.6
162 Hermes 3 Llama 3.1 8B other 25.5
163 Qwen2 1.5B alibaba 25.4
164 Phi-1 other 25.2
165 Phi-3 Vision other 24.3
166 Command Light cohere 24.1
167 Grok Vision Beta xai 24
168 Baichuan2 13B Chat other 23.4
169 Zephyr 7B Alpha other 23.2
170 Nous Capybara 34B other 23.1
171 Claude Instant 1 anthropic 22.5
172 Llama 2 70B meta 22.4
173 Flan-UL2 other 22
174 Claude 2 anthropic 21.9
175 Flan-T5 XXL other 21.4
176 Falcon 40B other 21.3
177 MPT 30B other 21.2
178 Open-Platypus other 21
179 Phi-2 other 20.5
180 Yi 34B other 20.5
181 Jurassic-2 Mid other 20.3
182 Chat Bison google 19.8
183 Nous Capybara 7B other 19.7
184 Llama 2 7B meta 19.4
185 Grok Beta xai 19.1
186 Baichuan2 7B Chat other 18.8
187 MPT 7B other 18.4
188 Phi-3 Mini other 17.9
189 GPT-3.5 Turbo openai 17.8
190 Qwen2.5 3B alibaba 17
191 Mistral Tiny mistral 16
192 Claude 2.1 anthropic 15.5
193 Qwen2.5 0.5B alibaba 15.4
194 TigerBot 70B Chat other 15.4
195 OpenChat 3.5 1210 other 15.2
196 StableLM Zephyr 3B other 14.9
197 Text Bison google 12.1
198 StableLM 3 4B other 11.5
199 Llama 3.2 1B meta 11.4
200 Ministral 3B mistral 10.1
201 Llama 3.2 3B meta 10
202 Phi-3.5 Mini other 9.6
203 Qwen2.5 1.5B alibaba 9.5
204 Llama Guard 3 8B meta 8.4
205 Embed English v3 cohere 0

MATH

Description

12,500 道竞赛数学题(含 7 个学科:代数、几何、数论、微积分、概率等,5 个难度级别),评估模型的复杂数学推理能力。

Core Specifications

Category License Last Updated
math MIT 2021-01-01

Benchmark

Unit
%

Sources

Official URL