Overview

评估多步骤软推理能力,包含逻辑、数学和高效推理三个领域各250个问题

Metrics

Metric Unit Direction
accuracy % ↑ Higher is better

Sources

Model Score Ranking

# Model Vendor Score
1 Claude 3.5 Sonnet (2024-10-22) anthropic 71.8
2 o1 openai 71.8
3 Gemini 1.5 Pro 002 google 71.6
4 Gemini 2.0 Flash Thinking google 70.8
5 Sonar Reasoning other 69.4
6 Gemini 2.0 Flash google 69.3
7 Qwen2 72B alibaba 68.9
8 GPT-4 0125 Preview openai 65.8
9 Sonar Large other 65.6
10 GPT-4 Vision Preview openai 64.5
11 Gemini 1.0 Ultra google 64.3
12 Mistral Medium mistral 64.3
13 Claude 3 Sonnet anthropic 64.1
14 o1 Preview openai 64.1
15 Qwen1.5 110B alibaba 63.9
16 Yi Large Turbo other 63.7
17 Claude 3 Sonnet (2024-02-29) anthropic 63.5
18 GPT-4o (2024-08-06) openai 62.9
19 GPT-4o (2024-05-13) openai 62.8
20 Llama 3.3 70B meta 62.3
21 Yi Vision other 61.4
22 GPT-4 Turbo openai 61.1
23 Gemini 1.5 Flash 002 google 59.8
24 Orca 2 13B other 59.8
25 Yi Large other 59.4
26 Falcon 180B other 59.1
27 Qwen1.5 14B alibaba 58.7
28 GPT-4 1106 Preview openai 58.4
29 Claude 3 Haiku anthropic 58.1
30 StableLM 2 12B other 58.1
31 Microsoft WizardLM 2 8x22B other 58
32 Gemini 1.0 Pro google 57.8
33 Grok-2 Vision xai 57.7
34 Hermes 3 Llama 3.1 405B other 57.7
35 Zephyr ORPO 141B Alpha other 57.7
36 Command Nightly cohere 57.5
37 Nous Hermes 2 Mixtral 8x7B other 57.4
38 Claude 3 Opus (2024-02-29) anthropic 57.3
39 Jamba 1.5 other 57.1
40 Command R cohere 57
41 Jamba Instruct other 56.9
42 WizardLM Team WizardLM 2 8x22B other 56.8
43 GLM-4 Flash other 56
44 Gemini 1.0 Flash google 55.9
45 Hermes 3 Llama 3.1 70B other 55.4
46 Sonar Huge other 55.4
47 GPT-4 32K openai 54.9
48 Qwen2.5 32B alibaba 54.9
49 DBRX Instruct other 54.7
50 Gemini 1.5 Flash-8B 002 google 54.7
51 GPT-4 openai 54.6
52 Qwen1.5 32B alibaba 54.5
53 Llama 3.1 Nemotron 70B meta 54.1
54 o1 mini openai 54
55 DBRX Base other 53.8
56 Llama 3.2 90B Vision meta 53.5
57 Mixtral 8x7B mistral 53.5
58 DeepSeek V2 deepseek 53.4
59 Microsoft WizardMath 7B v1 other 53.4
60 Claude 3 Opus anthropic 53.3
61 GPT-4o mini openai 53.1
62 Microsoft WizardCoder Python 34B other 53.1
63 Gemma 2 27B google 53
64 Claude 3.5 Haiku anthropic 52.3
65 Mistral Large mistral 52.1
66 DeepSeek V2 Chat deepseek 52
67 DeepSeek LLM 67B deepseek 51.9
68 StableCode 3B other 51.8
69 StarCoder2 15B other 51.7
70 Grok-2 Mini xai 51.4
71 GLM-4 Plus other 51.3
72 GPT-4 Vision openai 51.2
73 Llama 3 70B meta 51.2
74 Command R+ (08-2024) cohere 51.1
75 Jamba 1.5 Large other 51.1
76 Claude 3 Haiku (2024-03-07) anthropic 51
77 Yi 1.5 34B other 51
78 Qwen1.5 72B alibaba 50.8
79 DeepSeek Coder 7B deepseek 50.5
80 Qwen2.5 14B alibaba 50.5
81 Qwen2 57B alibaba 50.5
82 Phi-3.5 MoE other 50.3
83 Code Bison google 49.8
84 DeepSeek Math 7B deepseek 49.8
85 GLM-4 9B Chat other 49.5
86 Microsoft WizardLM 2 7B other 49.3
87 Mistral Small mistral 49.3
88 Command R (08-2024) cohere 49.1
89 Phi-1 other 49
90 Nous Hermes 2 Yi 34B other 48.8
91 Jamba 1.5 Mini other 48.1
92 Llama 3.1 70B meta 48.1
93 Mistral 7B v0.3 mistral 48.1
94 NVIDIA Llama 3.1 Nemotron 70B other 47.7
95 Sonar Small other 47.6
96 Code Llama 13B meta 47.3
97 Codestral mistral 47.2
98 Codestral Mamba mistral 47
99 Command R7B cohere 47
100 Qwen2.5 7B alibaba 46.9
101 OLMo 7B other 46.8
102 GLM-4 Air other 46.7
103 Yi 1.5 6B other 46.6
104 Yi 1.5 9B other 46.4
105 Gemini 1.5 Flash google 46.1
106 Llama 3 8B meta 46.1
107 Mistral Small 3 mistral 46
108 DeepSeek Coder 33B deepseek 45.8
109 Gemini 1.5 Flash-8B google 45.6
110 Code Llama 34B meta 44.9
111 Grok Vision Beta xai 44.5
112 StarChat2 15B v0.1 other 44.4
113 OLMo 7B SFT other 44.3
114 Gemma 2 9B google 44.1
115 Phi-3 Medium other 43.9
116 Qwen2 7B alibaba 43.9
117 Grok Beta xai 43.7
118 ChatGLM3 6B other 42.9
119 Code Llama 7B meta 42.9
120 Jurassic-2 Ultra other 42.9
121 Llemma 7B other 42.6
122 OLMo 1.7 7B other 42.6
123 Nous Capybara 7B other 42.5
124 OLMo 7B Instruct other 42.5
125 MPT 7B other 42.4
126 Mathstral 7B mistral 42.3
127 Flan-T5 XXL other 42.2
128 Llama 3.1 8B meta 41.7
129 Llama 3.2 11B Vision meta 41.7
130 Mistral Tiny mistral 41.4
131 Text Bison google 40.8
132 Command Light cohere 40.7
133 Zephyr 7B Alpha other 40.5
134 GPT-3.5 openai 40.4
135 DeepSeek Coder V2 deepseek 40.3
136 Flan-T5 XL other 40.2
137 Code Llama 70B meta 40.1
138 Open-Platypus other 40
139 Nous Hermes 2 Solar 10.7B other 39.6
140 TigerBot 70B Chat other 39.3
141 OLMo 2 1124 7B other 39.2
142 Phi-4 other 38.6
143 GLM-4V 9B other 38.5
144 Mistral Nemo mistral 38.5
145 StarCoder2 3B other 38.4
146 StarCoder2 7B other 38.3
147 OpenChat 3.6 8B other 38.2
148 Falcon 40B other 38.1
149 Hermes 3 Llama 3.1 8B other 38
150 Baichuan2 13B Chat other 37.9
151 Qwen2.5 1.5B alibaba 37.8
152 Claude 2.1 anthropic 37.7
153 Mistral 7B v0.1 mistral 37.4
154 Flan-UL2 other 37.2
155 Ministral 8B mistral 37
156 Yi 34B other 36.8
157 Mistral 7B v0.2 mistral 36.7
158 StableLM 3 4B other 36.6
159 Phi-3 Small other 36.5
160 Gemma 7B google 36.2
161 Llama 2 70B meta 36.1
162 Llama 2 7B meta 35.3
163 Claude 2 anthropic 35.2
164 Llama Guard 2 8B meta 34.9
165 Capybara 1.5B other 34
166 Llama 2 13B meta 33.8
167 Claude Instant 1 anthropic 33.5
168 Phi-3 Mini other 33.5
169 Baichuan2 7B Chat other 33.1
170 Chat Bison google 32.8
171 OpenChat 3.5 1210 other 32.6
172 Zephyr 7B Beta other 32.6
173 Jurassic-2 Mid other 32.3
174 Phi-3 Vision other 32.1
175 Nous Capybara 34B other 31.7
176 Phi-3.5 Mini other 31.2
177 Gemma 2B google 30.6
178 Phi-1.5 other 30.3
179 Argilla Notus 7B v1 other 30.1
180 Yi 6B other 29.5
181 Ministral 3B mistral 29.4
182 Llama 3.2 3B meta 29.3
183 PaLM 2 google 29.3
184 GPT-3.5 Turbo openai 29.2
185 GPT-3.5 Turbo 16K openai 29.1
186 MPT 30B other 28.9
187 Qwen2.5 0.5B alibaba 28.9
188 Qwen2 1.5B alibaba 27.8
189 Llama 3.2 1B meta 26.6
190 StableLM 2 1.6B other 26.4
191 Phi-2 other 23.7
192 StableLM Zephyr 3B other 23.5
193 Qwen2.5 3B alibaba 23.3
194 Llama Guard 3 8B meta 22.8
195 Embed English v3 cohere 0

MUSR

Description

评估多步骤软推理能力,包含逻辑、数学和高效推理三个领域各250个问题

Core Specifications

Category License Last Updated
reasoning MIT 2024-02-01

Benchmark

Unit
%

Sources

Official URL