MUSR: Benchmark Detailed Guide
Detailed guide to MUSR: category, metrics, sources and applicable models.
Overview
评估多步骤软推理能力,包含逻辑、数学和高效推理三个领域各250个问题
Metrics
| Metric | Unit | Direction |
|---|---|---|
| accuracy | % | ↑ Higher is better |
Sources
Model Score Ranking
| # | Model | Vendor | Score |
|---|---|---|---|
| 1 | Claude 3.5 Sonnet (2024-10-22) | anthropic | 71.8 |
| 2 | o1 | openai | 71.8 |
| 3 | Gemini 1.5 Pro 002 | 71.6 | |
| 4 | Gemini 2.0 Flash Thinking | 70.8 | |
| 5 | Sonar Reasoning | other | 69.4 |
| 6 | Gemini 2.0 Flash | 69.3 | |
| 7 | Qwen2 72B | alibaba | 68.9 |
| 8 | GPT-4 0125 Preview | openai | 65.8 |
| 9 | Sonar Large | other | 65.6 |
| 10 | GPT-4 Vision Preview | openai | 64.5 |
| 11 | Gemini 1.0 Ultra | 64.3 | |
| 12 | Mistral Medium | mistral | 64.3 |
| 13 | Claude 3 Sonnet | anthropic | 64.1 |
| 14 | o1 Preview | openai | 64.1 |
| 15 | Qwen1.5 110B | alibaba | 63.9 |
| 16 | Yi Large Turbo | other | 63.7 |
| 17 | Claude 3 Sonnet (2024-02-29) | anthropic | 63.5 |
| 18 | GPT-4o (2024-08-06) | openai | 62.9 |
| 19 | GPT-4o (2024-05-13) | openai | 62.8 |
| 20 | Llama 3.3 70B | meta | 62.3 |
| 21 | Yi Vision | other | 61.4 |
| 22 | GPT-4 Turbo | openai | 61.1 |
| 23 | Gemini 1.5 Flash 002 | 59.8 | |
| 24 | Orca 2 13B | other | 59.8 |
| 25 | Yi Large | other | 59.4 |
| 26 | Falcon 180B | other | 59.1 |
| 27 | Qwen1.5 14B | alibaba | 58.7 |
| 28 | GPT-4 1106 Preview | openai | 58.4 |
| 29 | Claude 3 Haiku | anthropic | 58.1 |
| 30 | StableLM 2 12B | other | 58.1 |
| 31 | Microsoft WizardLM 2 8x22B | other | 58 |
| 32 | Gemini 1.0 Pro | 57.8 | |
| 33 | Grok-2 Vision | xai | 57.7 |
| 34 | Hermes 3 Llama 3.1 405B | other | 57.7 |
| 35 | Zephyr ORPO 141B Alpha | other | 57.7 |
| 36 | Command Nightly | cohere | 57.5 |
| 37 | Nous Hermes 2 Mixtral 8x7B | other | 57.4 |
| 38 | Claude 3 Opus (2024-02-29) | anthropic | 57.3 |
| 39 | Jamba 1.5 | other | 57.1 |
| 40 | Command R | cohere | 57 |
| 41 | Jamba Instruct | other | 56.9 |
| 42 | WizardLM Team WizardLM 2 8x22B | other | 56.8 |
| 43 | GLM-4 Flash | other | 56 |
| 44 | Gemini 1.0 Flash | 55.9 | |
| 45 | Hermes 3 Llama 3.1 70B | other | 55.4 |
| 46 | Sonar Huge | other | 55.4 |
| 47 | GPT-4 32K | openai | 54.9 |
| 48 | Qwen2.5 32B | alibaba | 54.9 |
| 49 | DBRX Instruct | other | 54.7 |
| 50 | Gemini 1.5 Flash-8B 002 | 54.7 | |
| 51 | GPT-4 | openai | 54.6 |
| 52 | Qwen1.5 32B | alibaba | 54.5 |
| 53 | Llama 3.1 Nemotron 70B | meta | 54.1 |
| 54 | o1 mini | openai | 54 |
| 55 | DBRX Base | other | 53.8 |
| 56 | Llama 3.2 90B Vision | meta | 53.5 |
| 57 | Mixtral 8x7B | mistral | 53.5 |
| 58 | DeepSeek V2 | deepseek | 53.4 |
| 59 | Microsoft WizardMath 7B v1 | other | 53.4 |
| 60 | Claude 3 Opus | anthropic | 53.3 |
| 61 | GPT-4o mini | openai | 53.1 |
| 62 | Microsoft WizardCoder Python 34B | other | 53.1 |
| 63 | Gemma 2 27B | 53 | |
| 64 | Claude 3.5 Haiku | anthropic | 52.3 |
| 65 | Mistral Large | mistral | 52.1 |
| 66 | DeepSeek V2 Chat | deepseek | 52 |
| 67 | DeepSeek LLM 67B | deepseek | 51.9 |
| 68 | StableCode 3B | other | 51.8 |
| 69 | StarCoder2 15B | other | 51.7 |
| 70 | Grok-2 Mini | xai | 51.4 |
| 71 | GLM-4 Plus | other | 51.3 |
| 72 | GPT-4 Vision | openai | 51.2 |
| 73 | Llama 3 70B | meta | 51.2 |
| 74 | Command R+ (08-2024) | cohere | 51.1 |
| 75 | Jamba 1.5 Large | other | 51.1 |
| 76 | Claude 3 Haiku (2024-03-07) | anthropic | 51 |
| 77 | Yi 1.5 34B | other | 51 |
| 78 | Qwen1.5 72B | alibaba | 50.8 |
| 79 | DeepSeek Coder 7B | deepseek | 50.5 |
| 80 | Qwen2.5 14B | alibaba | 50.5 |
| 81 | Qwen2 57B | alibaba | 50.5 |
| 82 | Phi-3.5 MoE | other | 50.3 |
| 83 | Code Bison | 49.8 | |
| 84 | DeepSeek Math 7B | deepseek | 49.8 |
| 85 | GLM-4 9B Chat | other | 49.5 |
| 86 | Microsoft WizardLM 2 7B | other | 49.3 |
| 87 | Mistral Small | mistral | 49.3 |
| 88 | Command R (08-2024) | cohere | 49.1 |
| 89 | Phi-1 | other | 49 |
| 90 | Nous Hermes 2 Yi 34B | other | 48.8 |
| 91 | Jamba 1.5 Mini | other | 48.1 |
| 92 | Llama 3.1 70B | meta | 48.1 |
| 93 | Mistral 7B v0.3 | mistral | 48.1 |
| 94 | NVIDIA Llama 3.1 Nemotron 70B | other | 47.7 |
| 95 | Sonar Small | other | 47.6 |
| 96 | Code Llama 13B | meta | 47.3 |
| 97 | Codestral | mistral | 47.2 |
| 98 | Codestral Mamba | mistral | 47 |
| 99 | Command R7B | cohere | 47 |
| 100 | Qwen2.5 7B | alibaba | 46.9 |
| 101 | OLMo 7B | other | 46.8 |
| 102 | GLM-4 Air | other | 46.7 |
| 103 | Yi 1.5 6B | other | 46.6 |
| 104 | Yi 1.5 9B | other | 46.4 |
| 105 | Gemini 1.5 Flash | 46.1 | |
| 106 | Llama 3 8B | meta | 46.1 |
| 107 | Mistral Small 3 | mistral | 46 |
| 108 | DeepSeek Coder 33B | deepseek | 45.8 |
| 109 | Gemini 1.5 Flash-8B | 45.6 | |
| 110 | Code Llama 34B | meta | 44.9 |
| 111 | Grok Vision Beta | xai | 44.5 |
| 112 | StarChat2 15B v0.1 | other | 44.4 |
| 113 | OLMo 7B SFT | other | 44.3 |
| 114 | Gemma 2 9B | 44.1 | |
| 115 | Phi-3 Medium | other | 43.9 |
| 116 | Qwen2 7B | alibaba | 43.9 |
| 117 | Grok Beta | xai | 43.7 |
| 118 | ChatGLM3 6B | other | 42.9 |
| 119 | Code Llama 7B | meta | 42.9 |
| 120 | Jurassic-2 Ultra | other | 42.9 |
| 121 | Llemma 7B | other | 42.6 |
| 122 | OLMo 1.7 7B | other | 42.6 |
| 123 | Nous Capybara 7B | other | 42.5 |
| 124 | OLMo 7B Instruct | other | 42.5 |
| 125 | MPT 7B | other | 42.4 |
| 126 | Mathstral 7B | mistral | 42.3 |
| 127 | Flan-T5 XXL | other | 42.2 |
| 128 | Llama 3.1 8B | meta | 41.7 |
| 129 | Llama 3.2 11B Vision | meta | 41.7 |
| 130 | Mistral Tiny | mistral | 41.4 |
| 131 | Text Bison | 40.8 | |
| 132 | Command Light | cohere | 40.7 |
| 133 | Zephyr 7B Alpha | other | 40.5 |
| 134 | GPT-3.5 | openai | 40.4 |
| 135 | DeepSeek Coder V2 | deepseek | 40.3 |
| 136 | Flan-T5 XL | other | 40.2 |
| 137 | Code Llama 70B | meta | 40.1 |
| 138 | Open-Platypus | other | 40 |
| 139 | Nous Hermes 2 Solar 10.7B | other | 39.6 |
| 140 | TigerBot 70B Chat | other | 39.3 |
| 141 | OLMo 2 1124 7B | other | 39.2 |
| 142 | Phi-4 | other | 38.6 |
| 143 | GLM-4V 9B | other | 38.5 |
| 144 | Mistral Nemo | mistral | 38.5 |
| 145 | StarCoder2 3B | other | 38.4 |
| 146 | StarCoder2 7B | other | 38.3 |
| 147 | OpenChat 3.6 8B | other | 38.2 |
| 148 | Falcon 40B | other | 38.1 |
| 149 | Hermes 3 Llama 3.1 8B | other | 38 |
| 150 | Baichuan2 13B Chat | other | 37.9 |
| 151 | Qwen2.5 1.5B | alibaba | 37.8 |
| 152 | Claude 2.1 | anthropic | 37.7 |
| 153 | Mistral 7B v0.1 | mistral | 37.4 |
| 154 | Flan-UL2 | other | 37.2 |
| 155 | Ministral 8B | mistral | 37 |
| 156 | Yi 34B | other | 36.8 |
| 157 | Mistral 7B v0.2 | mistral | 36.7 |
| 158 | StableLM 3 4B | other | 36.6 |
| 159 | Phi-3 Small | other | 36.5 |
| 160 | Gemma 7B | 36.2 | |
| 161 | Llama 2 70B | meta | 36.1 |
| 162 | Llama 2 7B | meta | 35.3 |
| 163 | Claude 2 | anthropic | 35.2 |
| 164 | Llama Guard 2 8B | meta | 34.9 |
| 165 | Capybara 1.5B | other | 34 |
| 166 | Llama 2 13B | meta | 33.8 |
| 167 | Claude Instant 1 | anthropic | 33.5 |
| 168 | Phi-3 Mini | other | 33.5 |
| 169 | Baichuan2 7B Chat | other | 33.1 |
| 170 | Chat Bison | 32.8 | |
| 171 | OpenChat 3.5 1210 | other | 32.6 |
| 172 | Zephyr 7B Beta | other | 32.6 |
| 173 | Jurassic-2 Mid | other | 32.3 |
| 174 | Phi-3 Vision | other | 32.1 |
| 175 | Nous Capybara 34B | other | 31.7 |
| 176 | Phi-3.5 Mini | other | 31.2 |
| 177 | Gemma 2B | 30.6 | |
| 178 | Phi-1.5 | other | 30.3 |
| 179 | Argilla Notus 7B v1 | other | 30.1 |
| 180 | Yi 6B | other | 29.5 |
| 181 | Ministral 3B | mistral | 29.4 |
| 182 | Llama 3.2 3B | meta | 29.3 |
| 183 | PaLM 2 | 29.3 | |
| 184 | GPT-3.5 Turbo | openai | 29.2 |
| 185 | GPT-3.5 Turbo 16K | openai | 29.1 |
| 186 | MPT 30B | other | 28.9 |
| 187 | Qwen2.5 0.5B | alibaba | 28.9 |
| 188 | Qwen2 1.5B | alibaba | 27.8 |
| 189 | Llama 3.2 1B | meta | 26.6 |
| 190 | StableLM 2 1.6B | other | 26.4 |
| 191 | Phi-2 | other | 23.7 |
| 192 | StableLM Zephyr 3B | other | 23.5 |
| 193 | Qwen2.5 3B | alibaba | 23.3 |
| 194 | Llama Guard 3 8B | meta | 22.8 |
| 195 | Embed English v3 | cohere | 0 |
MUSR
Description
评估多步骤软推理能力,包含逻辑、数学和高效推理三个领域各250个问题
Core Specifications
| Category | License | Last Updated |
|---|---|---|
| reasoning | MIT | 2024-02-01 |
Benchmark
| Unit |
|---|
| % |