MMLU (Massive Multitask Language Understanding): Benchmark Detailed Guide
Detailed guide to MMLU (Massive Multitask Language Understanding): category, metrics, sources and applicable models.
Overview
57 学科多选题知识评测,覆盖人文、社科、STEM、医学等领域,共 14042 道题,评估模型的广泛世界知识与问题解决能力。
Metrics
| Metric | Unit | Direction |
|---|---|---|
| accuracy | % | ↑ Higher is better |
Sources
Model Score Ranking
| # | Model | Vendor | Score |
|---|---|---|---|
| 1 | GPT-4o (2024-05-13) | openai | 89.8 |
| 2 | Claude 3.5 Sonnet | anthropic | 88.7 |
| 3 | GPT-4o | openai | 88.7 |
| 4 | Llama 3.1 405B | meta | 88.6 |
| 5 | DeepSeek V3 | deepseek | 88.5 |
| 6 | Sonar Reasoning | other | 87.7 |
| 7 | GPT-4o (2024-08-06) | openai | 87.6 |
| 8 | Grok-2 | xai | 87.5 |
| 9 | Claude 3.5 Sonnet (2024-10-22) | anthropic | 87.2 |
| 10 | Gemini 1.5 Pro 002 | 87 | |
| 11 | Gemini 2.0 Flash Thinking | 86.5 | |
| 12 | Gemini 2.0 Flash | 86.3 | |
| 13 | o1 | openai | 86.3 |
| 14 | Qwen1.5 110B | alibaba | 86.1 |
| 15 | Qwen2.5 72B | alibaba | 86.1 |
| 16 | Gemini 1.5 Pro | 85.9 | |
| 17 | GPT-4 | openai | 85.8 |
| 18 | Claude 3 Opus (2024-02-29) | anthropic | 85.7 |
| 19 | Yi Vision | other | 85.5 |
| 20 | o1 Preview | openai | 85.3 |
| 21 | Command Nightly | cohere | 85.2 |
| 22 | GPT-4 Vision Preview | openai | 85 |
| 23 | Gemini 1.0 Pro | 84.5 | |
| 24 | GPT-4 Turbo | openai | 84.4 |
| 25 | Yi Large Turbo | other | 84.4 |
| 26 | GLM-4 Plus | other | 84.3 |
| 27 | Jamba 1.5 Large | other | 84.3 |
| 28 | Claude 3 Sonnet (2024-02-29) | anthropic | 84 |
| 29 | Mistral Large 2 | mistral | 84 |
| 30 | GPT-4 Vision | openai | 83.9 |
| 31 | Hermes 3 Llama 3.1 405B | other | 83.9 |
| 32 | Grok-2 Vision | xai | 83.4 |
| 33 | Llama 3.3 70B | meta | 83.4 |
| 34 | Mistral Medium | mistral | 82.9 |
| 35 | Yi Large | other | 82.8 |
| 36 | GPT-4 0125 Preview | openai | 82.6 |
| 37 | Sonar Large | other | 82.5 |
| 38 | Command R+ (08-2024) | cohere | 82.3 |
| 39 | Mistral Large | mistral | 82 |
| 40 | Qwen1.5 32B | alibaba | 81.9 |
| 41 | GPT-4 1106 Preview | openai | 81.8 |
| 42 | Claude 3 Sonnet | anthropic | 81.5 |
| 43 | Llama 3.1 Nemotron 70B | meta | 81.5 |
| 44 | Nous Hermes 2 Mixtral 8x7B | other | 81.4 |
| 45 | DBRX Instruct | other | 81.3 |
| 46 | Qwen1.5 72B | alibaba | 81.3 |
| 47 | Claude 3 Opus | anthropic | 81.2 |
| 48 | Mistral Small | mistral | 81.2 |
| 49 | Gemini 1.5 Flash 002 | 81.1 | |
| 50 | Sonar Huge | other | 81 |
| 51 | GPT-4 32K | openai | 80.8 |
| 52 | Grok-2 Mini | xai | 80.6 |
| 53 | NVIDIA Llama 3.1 Nemotron 70B | other | 80.4 |
| 54 | Gemini 1.0 Ultra | 80.3 | |
| 55 | Qwen2 72B | alibaba | 80.2 |
| 56 | GLM-4 Flash | other | 80 |
| 57 | Sonar Small | other | 80 |
| 58 | DBRX Base | other | 79.9 |
| 59 | GPT-4o mini | openai | 79.7 |
| 60 | Llama 3 70B | meta | 79.5 |
| 61 | Zephyr ORPO 141B Alpha | other | 79.5 |
| 62 | Command R | cohere | 79.3 |
| 63 | WizardLM Team WizardLM 2 8x22B | other | 79.2 |
| 64 | GLM-4 Air | other | 78.9 |
| 65 | Microsoft WizardLM 2 8x22B | other | 78.9 |
| 66 | Gemini 1.0 Flash | 78.7 | |
| 67 | Jamba 1.5 | other | 78.5 |
| 68 | Qwen2.5 32B | alibaba | 78.5 |
| 69 | Orca 2 13B | other | 78.4 |
| 70 | Qwen1.5 14B | alibaba | 78.4 |
| 71 | DeepSeek V2 | deepseek | 78.2 |
| 72 | o1 mini | openai | 78.2 |
| 73 | Claude 3 Haiku | anthropic | 77.9 |
| 74 | Jamba 1.5 Mini | other | 77.9 |
| 75 | Command R (08-2024) | cohere | 77.8 |
| 76 | Mixtral 8x22B | mistral | 77.8 |
| 77 | Qwen2 57B | alibaba | 77.7 |
| 78 | Claude 3.5 Haiku | anthropic | 77.6 |
| 79 | Gemini 1.5 Flash-8B 002 | 77.5 | |
| 80 | Gemma 2 27B | 77.4 | |
| 81 | StableCode 3B | other | 77.4 |
| 82 | Mixtral 8x7B | mistral | 77.2 |
| 83 | Claude 3 Haiku (2024-03-07) | anthropic | 77.1 |
| 84 | Mistral Small 3 | mistral | 76.9 |
| 85 | Nous Hermes 2 Yi 34B | other | 76.9 |
| 86 | Falcon 180B | other | 76.7 |
| 87 | Hermes 3 Llama 3.1 70B | other | 76.7 |
| 88 | DeepSeek V2 Chat | deepseek | 76.6 |
| 89 | Gemini 1.5 Flash-8B | 76.6 | |
| 90 | Llama 3.2 90B Vision | meta | 76.5 |
| 91 | Jamba Instruct | other | 76.2 |
| 92 | Yi 1.5 34B | other | 76.1 |
| 93 | Gemini 1.5 Flash | 76 | |
| 94 | Phi-3.5 MoE | other | 75.7 |
| 95 | StableLM 2 12B | other | 75.7 |
| 96 | Llama 3.1 70B | meta | 75.6 |
| 97 | Qwen2.5 14B | alibaba | 75.4 |
| 98 | DeepSeek LLM 67B | deepseek | 75.2 |
| 99 | Codestral | mistral | 75.1 |
| 100 | Command R+ | cohere | 75 |
| 101 | StarChat2 15B v0.1 | other | 74.9 |
| 102 | Code Llama 34B | meta | 74.5 |
| 103 | Qwen2.5 7B | alibaba | 73.6 |
| 104 | GLM-4V 9B | other | 73.4 |
| 105 | OLMo 7B SFT | other | 73.4 |
| 106 | Hermes 3 Llama 3.1 8B | other | 73.1 |
| 107 | Llama 3.1 8B | meta | 73 |
| 108 | Mistral 7B v0.3 | mistral | 72.6 |
| 109 | DeepSeek Coder V2 | deepseek | 72.3 |
| 110 | Phi-3 Small | other | 71.8 |
| 111 | TigerBot 70B Chat | other | 71.7 |
| 112 | Argilla Notus 7B v1 | other | 71.4 |
| 113 | Llama 2 7B | meta | 70.8 |
| 114 | GLM-4 9B Chat | other | 70.7 |
| 115 | OpenChat 3.5 1210 | other | 70.5 |
| 116 | Gemma 7B | 69.6 | |
| 117 | DeepSeek Math 7B | deepseek | 69.3 |
| 118 | Mistral Nemo | mistral | 69.2 |
| 119 | OpenChat 3.6 8B | other | 69.2 |
| 120 | Phi-4 | other | 69.2 |
| 121 | Mathstral 7B | mistral | 69.1 |
| 122 | Microsoft WizardMath 7B v1 | other | 69.1 |
| 123 | Open-Platypus | other | 68.7 |
| 124 | Mistral 7B v0.1 | mistral | 68.4 |
| 125 | PaLM 2 | 68.4 | |
| 126 | OLMo 2 1124 7B | other | 68.3 |
| 127 | Phi-1 | other | 68.1 |
| 128 | Yi 6B | other | 68 |
| 129 | Llama 3 8B | meta | 67.8 |
| 130 | Yi 1.5 9B | other | 67.8 |
| 131 | StarCoder2 7B | other | 67.4 |
| 132 | Claude 2 | anthropic | 66.3 |
| 133 | Llemma 7B | other | 66 |
| 134 | Code Llama 13B | meta | 65.9 |
| 135 | DeepSeek Coder 33B | deepseek | 65.9 |
| 136 | Microsoft WizardLM 2 7B | other | 65.7 |
| 137 | OLMo 7B Instruct | other | 65.6 |
| 138 | OLMo 1.7 7B | other | 65.5 |
| 139 | Zephyr 7B Beta | other | 65.1 |
| 140 | Code Bison | 65 | |
| 141 | Command R7B | cohere | 64.9 |
| 142 | Ministral 8B | mistral | 64.9 |
| 143 | Mistral 7B v0.2 | mistral | 64.4 |
| 144 | Nous Capybara 7B | other | 64.4 |
| 145 | Gemma 2 9B | 64.3 | |
| 146 | GPT-3.5 Turbo 16K | openai | 64.2 |
| 147 | DeepSeek Coder 7B | deepseek | 63.6 |
| 148 | Nous Hermes 2 Solar 10.7B | other | 63.4 |
| 149 | Qwen2 7B | alibaba | 63.4 |
| 150 | Grok Vision Beta | xai | 63.3 |
| 151 | Chat Bison | 63.1 | |
| 152 | OLMo 7B | other | 62.6 |
| 153 | Jurassic-2 Ultra | other | 62.5 |
| 154 | Code Llama 70B | meta | 62.1 |
| 155 | Phi-3 Medium | other | 62.1 |
| 156 | Llama 3.2 11B Vision | meta | 61.9 |
| 157 | GPT-3.5 Turbo | openai | 61.1 |
| 158 | Yi 1.5 6B | other | 61.1 |
| 159 | Claude 2.1 | anthropic | 60.9 |
| 160 | StarCoder2 15B | other | 60.6 |
| 161 | StarCoder2 3B | other | 60.6 |
| 162 | Flan-T5 XL | other | 60.4 |
| 163 | Microsoft WizardCoder Python 34B | other | 60 |
| 164 | Codestral Mamba | mistral | 59.3 |
| 165 | Flan-UL2 | other | 58.7 |
| 166 | Baichuan2 13B Chat | other | 58.6 |
| 167 | Code Llama 7B | meta | 58.5 |
| 168 | Phi-1.5 | other | 58.4 |
| 169 | ChatGLM3 6B | other | 58 |
| 170 | Baichuan2 7B Chat | other | 57.9 |
| 171 | Phi-2 | other | 57.5 |
| 172 | Claude Instant 1 | anthropic | 57.1 |
| 173 | Qwen2.5 3B | alibaba | 56.8 |
| 174 | Falcon 40B | other | 56.7 |
| 175 | Llama 2 13B | meta | 56.4 |
| 176 | Phi-3 Vision | other | 56.4 |
| 177 | Jurassic-2 Mid | other | 55.9 |
| 178 | Phi-3.5 Mini | other | 55.7 |
| 179 | Yi 34B | other | 55.5 |
| 180 | MPT 30B | other | 55.2 |
| 181 | Qwen2.5 1.5B | alibaba | 54.5 |
| 182 | Mistral Tiny | mistral | 54.2 |
| 183 | Text Bison | 53.9 | |
| 184 | Nous Capybara 34B | other | 53.2 |
| 185 | Llama 2 70B | meta | 53 |
| 186 | Command Light | cohere | 52.8 |
| 187 | StableLM 2 1.6B | other | 52.7 |
| 188 | Llama 3.2 3B | meta | 52.5 |
| 189 | Zephyr 7B Alpha | other | 52.2 |
| 190 | Flan-T5 XXL | other | 52 |
| 191 | Capybara 1.5B | other | 51.7 |
| 192 | MPT 7B | other | 51.7 |
| 193 | Grok Beta | xai | 51 |
| 194 | Llama Guard 2 8B | meta | 51 |
| 195 | StableLM 3 4B | other | 50.6 |
| 196 | GPT-3.5 | openai | 50.3 |
| 197 | Llama 3.2 1B | meta | 48.6 |
| 198 | Gemma 2B | 48 | |
| 199 | StableLM Zephyr 3B | other | 44.3 |
| 200 | Phi-3 Mini | other | 43.6 |
| 201 | Ministral 3B | mistral | 42.4 |
| 202 | Qwen2.5 0.5B | alibaba | 41.5 |
| 203 | Qwen2 1.5B | alibaba | 40.5 |
| 204 | Llama Guard 3 8B | meta | 40.1 |
| 205 | Embed English v3 | cohere | 0 |
MMLU (Massive Multitask Language Understanding)
Description
57 学科多选题知识评测,覆盖人文、社科、STEM、医学等领域,共 14042 道题,评估模型的广泛世界知识与问题解决能力。
Core Specifications
| Category | License | Last Updated |
|---|---|---|
| knowledge | MIT | 2024-01-01 |
Benchmark
| Unit |
|---|
| % |