Overview

57 学科多选题知识评测,覆盖人文、社科、STEM、医学等领域,共 14042 道题,评估模型的广泛世界知识与问题解决能力。

Metrics

Metric Unit Direction
accuracy % ↑ Higher is better

Sources

Model Score Ranking

# Model Vendor Score
1 GPT-4o (2024-05-13) openai 89.8
2 Claude 3.5 Sonnet anthropic 88.7
3 GPT-4o openai 88.7
4 Llama 3.1 405B meta 88.6
5 DeepSeek V3 deepseek 88.5
6 Sonar Reasoning other 87.7
7 GPT-4o (2024-08-06) openai 87.6
8 Grok-2 xai 87.5
9 Claude 3.5 Sonnet (2024-10-22) anthropic 87.2
10 Gemini 1.5 Pro 002 google 87
11 Gemini 2.0 Flash Thinking google 86.5
12 Gemini 2.0 Flash google 86.3
13 o1 openai 86.3
14 Qwen1.5 110B alibaba 86.1
15 Qwen2.5 72B alibaba 86.1
16 Gemini 1.5 Pro google 85.9
17 GPT-4 openai 85.8
18 Claude 3 Opus (2024-02-29) anthropic 85.7
19 Yi Vision other 85.5
20 o1 Preview openai 85.3
21 Command Nightly cohere 85.2
22 GPT-4 Vision Preview openai 85
23 Gemini 1.0 Pro google 84.5
24 GPT-4 Turbo openai 84.4
25 Yi Large Turbo other 84.4
26 GLM-4 Plus other 84.3
27 Jamba 1.5 Large other 84.3
28 Claude 3 Sonnet (2024-02-29) anthropic 84
29 Mistral Large 2 mistral 84
30 GPT-4 Vision openai 83.9
31 Hermes 3 Llama 3.1 405B other 83.9
32 Grok-2 Vision xai 83.4
33 Llama 3.3 70B meta 83.4
34 Mistral Medium mistral 82.9
35 Yi Large other 82.8
36 GPT-4 0125 Preview openai 82.6
37 Sonar Large other 82.5
38 Command R+ (08-2024) cohere 82.3
39 Mistral Large mistral 82
40 Qwen1.5 32B alibaba 81.9
41 GPT-4 1106 Preview openai 81.8
42 Claude 3 Sonnet anthropic 81.5
43 Llama 3.1 Nemotron 70B meta 81.5
44 Nous Hermes 2 Mixtral 8x7B other 81.4
45 DBRX Instruct other 81.3
46 Qwen1.5 72B alibaba 81.3
47 Claude 3 Opus anthropic 81.2
48 Mistral Small mistral 81.2
49 Gemini 1.5 Flash 002 google 81.1
50 Sonar Huge other 81
51 GPT-4 32K openai 80.8
52 Grok-2 Mini xai 80.6
53 NVIDIA Llama 3.1 Nemotron 70B other 80.4
54 Gemini 1.0 Ultra google 80.3
55 Qwen2 72B alibaba 80.2
56 GLM-4 Flash other 80
57 Sonar Small other 80
58 DBRX Base other 79.9
59 GPT-4o mini openai 79.7
60 Llama 3 70B meta 79.5
61 Zephyr ORPO 141B Alpha other 79.5
62 Command R cohere 79.3
63 WizardLM Team WizardLM 2 8x22B other 79.2
64 GLM-4 Air other 78.9
65 Microsoft WizardLM 2 8x22B other 78.9
66 Gemini 1.0 Flash google 78.7
67 Jamba 1.5 other 78.5
68 Qwen2.5 32B alibaba 78.5
69 Orca 2 13B other 78.4
70 Qwen1.5 14B alibaba 78.4
71 DeepSeek V2 deepseek 78.2
72 o1 mini openai 78.2
73 Claude 3 Haiku anthropic 77.9
74 Jamba 1.5 Mini other 77.9
75 Command R (08-2024) cohere 77.8
76 Mixtral 8x22B mistral 77.8
77 Qwen2 57B alibaba 77.7
78 Claude 3.5 Haiku anthropic 77.6
79 Gemini 1.5 Flash-8B 002 google 77.5
80 Gemma 2 27B google 77.4
81 StableCode 3B other 77.4
82 Mixtral 8x7B mistral 77.2
83 Claude 3 Haiku (2024-03-07) anthropic 77.1
84 Mistral Small 3 mistral 76.9
85 Nous Hermes 2 Yi 34B other 76.9
86 Falcon 180B other 76.7
87 Hermes 3 Llama 3.1 70B other 76.7
88 DeepSeek V2 Chat deepseek 76.6
89 Gemini 1.5 Flash-8B google 76.6
90 Llama 3.2 90B Vision meta 76.5
91 Jamba Instruct other 76.2
92 Yi 1.5 34B other 76.1
93 Gemini 1.5 Flash google 76
94 Phi-3.5 MoE other 75.7
95 StableLM 2 12B other 75.7
96 Llama 3.1 70B meta 75.6
97 Qwen2.5 14B alibaba 75.4
98 DeepSeek LLM 67B deepseek 75.2
99 Codestral mistral 75.1
100 Command R+ cohere 75
101 StarChat2 15B v0.1 other 74.9
102 Code Llama 34B meta 74.5
103 Qwen2.5 7B alibaba 73.6
104 GLM-4V 9B other 73.4
105 OLMo 7B SFT other 73.4
106 Hermes 3 Llama 3.1 8B other 73.1
107 Llama 3.1 8B meta 73
108 Mistral 7B v0.3 mistral 72.6
109 DeepSeek Coder V2 deepseek 72.3
110 Phi-3 Small other 71.8
111 TigerBot 70B Chat other 71.7
112 Argilla Notus 7B v1 other 71.4
113 Llama 2 7B meta 70.8
114 GLM-4 9B Chat other 70.7
115 OpenChat 3.5 1210 other 70.5
116 Gemma 7B google 69.6
117 DeepSeek Math 7B deepseek 69.3
118 Mistral Nemo mistral 69.2
119 OpenChat 3.6 8B other 69.2
120 Phi-4 other 69.2
121 Mathstral 7B mistral 69.1
122 Microsoft WizardMath 7B v1 other 69.1
123 Open-Platypus other 68.7
124 Mistral 7B v0.1 mistral 68.4
125 PaLM 2 google 68.4
126 OLMo 2 1124 7B other 68.3
127 Phi-1 other 68.1
128 Yi 6B other 68
129 Llama 3 8B meta 67.8
130 Yi 1.5 9B other 67.8
131 StarCoder2 7B other 67.4
132 Claude 2 anthropic 66.3
133 Llemma 7B other 66
134 Code Llama 13B meta 65.9
135 DeepSeek Coder 33B deepseek 65.9
136 Microsoft WizardLM 2 7B other 65.7
137 OLMo 7B Instruct other 65.6
138 OLMo 1.7 7B other 65.5
139 Zephyr 7B Beta other 65.1
140 Code Bison google 65
141 Command R7B cohere 64.9
142 Ministral 8B mistral 64.9
143 Mistral 7B v0.2 mistral 64.4
144 Nous Capybara 7B other 64.4
145 Gemma 2 9B google 64.3
146 GPT-3.5 Turbo 16K openai 64.2
147 DeepSeek Coder 7B deepseek 63.6
148 Nous Hermes 2 Solar 10.7B other 63.4
149 Qwen2 7B alibaba 63.4
150 Grok Vision Beta xai 63.3
151 Chat Bison google 63.1
152 OLMo 7B other 62.6
153 Jurassic-2 Ultra other 62.5
154 Code Llama 70B meta 62.1
155 Phi-3 Medium other 62.1
156 Llama 3.2 11B Vision meta 61.9
157 GPT-3.5 Turbo openai 61.1
158 Yi 1.5 6B other 61.1
159 Claude 2.1 anthropic 60.9
160 StarCoder2 15B other 60.6
161 StarCoder2 3B other 60.6
162 Flan-T5 XL other 60.4
163 Microsoft WizardCoder Python 34B other 60
164 Codestral Mamba mistral 59.3
165 Flan-UL2 other 58.7
166 Baichuan2 13B Chat other 58.6
167 Code Llama 7B meta 58.5
168 Phi-1.5 other 58.4
169 ChatGLM3 6B other 58
170 Baichuan2 7B Chat other 57.9
171 Phi-2 other 57.5
172 Claude Instant 1 anthropic 57.1
173 Qwen2.5 3B alibaba 56.8
174 Falcon 40B other 56.7
175 Llama 2 13B meta 56.4
176 Phi-3 Vision other 56.4
177 Jurassic-2 Mid other 55.9
178 Phi-3.5 Mini other 55.7
179 Yi 34B other 55.5
180 MPT 30B other 55.2
181 Qwen2.5 1.5B alibaba 54.5
182 Mistral Tiny mistral 54.2
183 Text Bison google 53.9
184 Nous Capybara 34B other 53.2
185 Llama 2 70B meta 53
186 Command Light cohere 52.8
187 StableLM 2 1.6B other 52.7
188 Llama 3.2 3B meta 52.5
189 Zephyr 7B Alpha other 52.2
190 Flan-T5 XXL other 52
191 Capybara 1.5B other 51.7
192 MPT 7B other 51.7
193 Grok Beta xai 51
194 Llama Guard 2 8B meta 51
195 StableLM 3 4B other 50.6
196 GPT-3.5 openai 50.3
197 Llama 3.2 1B meta 48.6
198 Gemma 2B google 48
199 StableLM Zephyr 3B other 44.3
200 Phi-3 Mini other 43.6
201 Ministral 3B mistral 42.4
202 Qwen2.5 0.5B alibaba 41.5
203 Qwen2 1.5B alibaba 40.5
204 Llama Guard 3 8B meta 40.1
205 Embed English v3 cohere 0

MMLU (Massive Multitask Language Understanding)

Description

57 学科多选题知识评测,覆盖人文、社科、STEM、医学等领域,共 14042 道题,评估模型的广泛世界知识与问题解决能力。

Core Specifications

Category License Last Updated
knowledge MIT 2024-01-01

Benchmark

Unit
%

Sources

Official URL