Overview

OpenAI 发布的 164 道 Python 编程任务,每题包含函数签名、文档字符串、函数体与单元测试,评估模型的代码生成能力(pass@1)。

Metrics

Metric Unit Direction
pass@1 pass@1 ↑ Higher is better

Sources

Model Score Ranking

# Model Vendor Score
1 Claude 3.5 Sonnet anthropic 92
2 Mistral Large 2 mistral 92
3 Gemini 2.0 Flash google 90.7
4 Gemini 1.5 Pro 002 google 90.3
5 GPT-4o openai 90.2
6 o1 Preview openai 90.1
7 Hermes 3 Llama 3.1 405B other 89
8 Llama 3.1 405B meta 89
9 DeepSeek Coder V2 deepseek 88.9
10 Grok-2 xai 88.4
11 Phi-1 other 88.1
12 Gemini 2.0 Flash Thinking google 87.2
13 Llama 3.3 70B meta 87
14 Qwen2.5 72B alibaba 86.6
15 Codestral Mamba mistral 86.2
16 Claude 3.5 Sonnet (2024-10-22) anthropic 86.1
17 o1 openai 85.3
18 GPT-4o (2024-08-06) openai 85.1
19 Grok-2 Vision xai 84.2
20 GPT-4 1106 Preview openai 83.7
21 Claude 3 Opus anthropic 83.3
22 GPT-4o (2024-05-13) openai 83.2
23 DeepSeek V3 deepseek 82.6
24 StarCoder2 3B other 82.1
25 Gemini 1.0 Ultra google 81.7
26 Sonar Reasoning other 81.1
27 GPT-4 Turbo openai 80.5
28 Sonar Large other 80.4
29 GPT-4 0125 Preview openai 80.3
30 Claude 3 Sonnet (2024-02-29) anthropic 80.1
31 Llama 3.1 Nemotron 70B meta 79.8
32 Microsoft WizardCoder Python 34B other 79.8
33 Claude 3 Opus (2024-02-29) anthropic 79.7
34 Gemini 1.5 Flash-8B 002 google 79.7
35 Llama 3.1 70B meta 79.7
36 Mixtral 8x7B mistral 79
37 Nous Hermes 2 Mixtral 8x7B other 78.7
38 GPT-4o mini openai 78
39 Mistral Medium mistral 78
40 o1 mini openai 78
41 Grok-2 Mini xai 77.9
42 Microsoft WizardLM 2 8x22B other 77.9
43 Qwen1.5 32B alibaba 77.9
44 GPT-4 openai 77.6
45 StableCode 3B other 77
46 GPT-4 32K openai 76.8
47 Qwen1.5 110B alibaba 76.5
48 Qwen2 72B alibaba 76.5
49 Code Llama 7B meta 76.4
50 Gemini 1.0 Pro google 76.4
51 StableLM 2 12B other 76.4
52 Code Bison google 76.2
53 Yi Vision other 76.1
54 DeepSeek Coder 7B deepseek 75.9
55 Jamba 1.5 other 75.9
56 Command R cohere 75.8
57 GPT-4 Vision openai 75.8
58 Hermes 3 Llama 3.1 70B other 75.8
59 Sonar Huge other 75.8
60 DeepSeek V2 deepseek 75.7
61 Command R+ (08-2024) cohere 75
62 Claude 3 Haiku (2024-03-07) anthropic 74.8
63 Command Nightly cohere 74.7
64 Zephyr ORPO 141B Alpha other 74.7
65 StarCoder2 7B other 74.6
66 Claude 3 Sonnet anthropic 74.5
67 Yi 1.5 34B other 74.5
68 Mistral Small 3 mistral 74.4
69 Orca 2 13B other 74.3
70 Mistral Small mistral 74
71 Claude 3.5 Haiku anthropic 73.8
72 Jamba 1.5 Large other 73.7
73 Command R (08-2024) cohere 73.6
74 StarCoder2 15B other 73.5
75 Mistral Large mistral 73.4
76 Nous Hermes 2 Yi 34B other 73.2
77 Llama 3.2 90B Vision meta 73.1
78 Llama 3 70B meta 73
79 DBRX Instruct other 72.8
80 Sonar Small other 72.8
81 GLM-4 Plus other 72.7
82 Gemma 2 27B google 72.4
83 Yi Large Turbo other 72.4
84 Yi Large other 72.2
85 DBRX Base other 72.1
86 Qwen2.5 14B alibaba 72
87 Gemini 1.5 Pro google 71.9
88 DeepSeek Coder 33B deepseek 71.8
89 Code Llama 34B meta 71.7
90 Command R+ cohere 70.7
91 Gemini 1.5 Flash-8B google 70.7
92 GPT-4 Vision Preview openai 70.6
93 WizardLM Team WizardLM 2 8x22B other 70.2
94 Code Llama 13B meta 69.5
95 Phi-3.5 MoE other 69.5
96 Claude 3 Haiku anthropic 69.4
97 Mistral Nemo mistral 69.4
98 Codestral mistral 69.2
99 StarChat2 15B v0.1 other 69
100 Falcon 180B other 68.9
101 Mistral 7B v0.3 mistral 68.8
102 Microsoft WizardLM 2 7B other 68.6
103 DeepSeek LLM 67B deepseek 68.3
104 GLM-4 Air other 68.3
105 Phi-3 Small other 68.3
106 Qwen2.5 32B alibaba 68
107 GLM-4 Flash other 67.6
108 Qwen2 57B alibaba 67.6
109 Gemini 1.5 Flash google 67.5
110 Hermes 3 Llama 3.1 8B other 67.1
111 DeepSeek V2 Chat deepseek 67
112 Jamba 1.5 Mini other 67
113 OLMo 1.7 7B other 67
114 Qwen2 7B alibaba 67
115 Qwen2.5 7B alibaba 66.5
116 NVIDIA Llama 3.1 Nemotron 70B other 66.4
117 Gemini 1.5 Flash 002 google 66.2
118 GLM-4V 9B other 66
119 Jamba Instruct other 66
120 Qwen1.5 14B alibaba 65.8
121 Gemini 1.0 Flash google 65.7
122 Code Llama 70B meta 65.6
123 Qwen1.5 72B alibaba 65.2
124 Llama 3 8B meta 65
125 Llemma 7B other 64.2
126 Llama 3.1 8B meta 63.4
127 Ministral 8B mistral 63.3
128 Yi 1.5 6B other 63.1
129 GLM-4 9B Chat other 62.9
130 Phi-4 other 62.8
131 OLMo 7B other 62.2
132 OpenChat 3.6 8B other 61.5
133 Microsoft WizardMath 7B v1 other 61.1
134 Gemma 7B google 61
135 Gemma 2 9B google 60.9
136 Mistral 7B v0.2 mistral 59.7
137 Nous Hermes 2 Solar 10.7B other 59.2
138 Yi 1.5 9B other 57.3
139 Mistral 7B v0.1 mistral 56.4
140 DeepSeek Math 7B deepseek 54
141 OLMo 7B SFT other 53.9
142 Phi-3 Medium other 53.9
143 OLMo 2 1124 7B other 53.2
144 GPT-3.5 Turbo openai 51.3
145 Command Light cohere 51.2
146 ChatGLM3 6B other 51
147 Jurassic-2 Ultra other 50.9
148 Llama 2 70B meta 50.6
149 OLMo 7B Instruct other 50.5
150 Command R7B cohere 50.1
151 Llama 3.2 11B Vision meta 50.1
152 Nous Capybara 34B other 50.1
153 Mathstral 7B mistral 49.9
154 MPT 30B other 49.7
155 Baichuan2 13B Chat other 49.6
156 Baichuan2 7B Chat other 49.2
157 Qwen2.5 3B alibaba 49
158 Open-Platypus other 48.2
159 Phi-3 Vision other 48.2
160 Qwen2 1.5B alibaba 47.2
161 StableLM Zephyr 3B other 47.1
162 Flan-UL2 other 46.8
163 Mistral Tiny mistral 45.8
164 Mixtral 8x22B mistral 45.2
165 Text Bison google 45.1
166 PaLM 2 google 44.8
167 Grok Vision Beta xai 44
168 Phi-3.5 Mini other 44
169 Gemma 2B google 43.2
170 Llama 2 7B meta 42.6
171 Llama 3.2 3B meta 42.5
172 Yi 6B other 42.3
173 OpenChat 3.5 1210 other 41.9
174 Llama 2 13B meta 40.7
175 Falcon 40B other 39.7
176 Jurassic-2 Mid other 39.7
177 Phi-2 other 39.7
178 Phi-1.5 other 39.2
179 Nous Capybara 7B other 38.8
180 Zephyr 7B Beta other 38.4
181 MPT 7B other 38.1
182 StableLM 2 1.6B other 37.3
183 Ministral 3B mistral 37.2
184 Argilla Notus 7B v1 other 36.8
185 GPT-3.5 Turbo 16K openai 36.8
186 Chat Bison google 36
187 TigerBot 70B Chat other 36
188 Zephyr 7B Alpha other 35.9
189 Yi 34B other 35.8
190 Flan-T5 XXL other 35.4
191 Qwen2.5 1.5B alibaba 35.2
192 Flan-T5 XL other 34
193 StableLM 3 4B other 32.6
194 Claude 2 anthropic 31.7
195 Grok Beta xai 31
196 Qwen2.5 0.5B alibaba 31
197 Capybara 1.5B other 30.1
198 Llama 3.2 1B meta 29.9
199 Claude Instant 1 anthropic 28.5
200 GPT-3.5 openai 28.2
201 Claude 2.1 anthropic 28.1
202 Llama Guard 3 8B meta 28
203 Phi-3 Mini other 26.9
204 Llama Guard 2 8B meta 26.6
205 Embed English v3 cohere 0

HumanEval

Description

OpenAI 发布的 164 道 Python 编程任务,每题包含函数签名、文档字符串、函数体与单元测试,评估模型的代码生成能力(pass@1)。

Core Specifications

Category License Last Updated
coding MIT 2022-01-01

Benchmark

Unit
pass@1

Sources

Official URL