HumanEval:基准测试详解
详细介绍 HumanEval 基准:分类、评估指标、数据来源与适用模型。
基准介绍
OpenAI 发布的 164 道 Python 编程任务,每题包含函数签名、文档字符串、函数体与单元测试,评估模型的代码生成能力(pass@1)。
评测指标
| 指标名 | 单位 | 方向 |
|---|---|---|
| pass@1 | pass@1 | ↑ 越高越好 |
数据来源
模型得分排名
| # | 模型 | 厂商 | 得分 |
|---|---|---|---|
| 1 | Claude 3.5 Sonnet | anthropic | 92 |
| 2 | Mistral Large 2 | mistral | 92 |
| 3 | Gemini 2.0 Flash | 90.7 | |
| 4 | Gemini 1.5 Pro 002 | 90.3 | |
| 5 | GPT-4o | openai | 90.2 |
| 6 | o1 Preview | openai | 90.1 |
| 7 | Hermes 3 Llama 3.1 405B | other | 89 |
| 8 | Llama 3.1 405B | meta | 89 |
| 9 | DeepSeek Coder V2 | deepseek | 88.9 |
| 10 | Grok-2 | xai | 88.4 |
| 11 | Phi-1 | other | 88.1 |
| 12 | Gemini 2.0 Flash Thinking | 87.2 | |
| 13 | Llama 3.3 70B | meta | 87 |
| 14 | Qwen2.5 72B | alibaba | 86.6 |
| 15 | Codestral Mamba | mistral | 86.2 |
| 16 | Claude 3.5 Sonnet (2024-10-22) | anthropic | 86.1 |
| 17 | o1 | openai | 85.3 |
| 18 | GPT-4o (2024-08-06) | openai | 85.1 |
| 19 | Grok-2 Vision | xai | 84.2 |
| 20 | GPT-4 1106 Preview | openai | 83.7 |
| 21 | Claude 3 Opus | anthropic | 83.3 |
| 22 | GPT-4o (2024-05-13) | openai | 83.2 |
| 23 | DeepSeek V3 | deepseek | 82.6 |
| 24 | StarCoder2 3B | other | 82.1 |
| 25 | Gemini 1.0 Ultra | 81.7 | |
| 26 | Sonar Reasoning | other | 81.1 |
| 27 | GPT-4 Turbo | openai | 80.5 |
| 28 | Sonar Large | other | 80.4 |
| 29 | GPT-4 0125 Preview | openai | 80.3 |
| 30 | Claude 3 Sonnet (2024-02-29) | anthropic | 80.1 |
| 31 | Llama 3.1 Nemotron 70B | meta | 79.8 |
| 32 | Microsoft WizardCoder Python 34B | other | 79.8 |
| 33 | Claude 3 Opus (2024-02-29) | anthropic | 79.7 |
| 34 | Gemini 1.5 Flash-8B 002 | 79.7 | |
| 35 | Llama 3.1 70B | meta | 79.7 |
| 36 | Mixtral 8x7B | mistral | 79 |
| 37 | Nous Hermes 2 Mixtral 8x7B | other | 78.7 |
| 38 | GPT-4o mini | openai | 78 |
| 39 | Mistral Medium | mistral | 78 |
| 40 | o1 mini | openai | 78 |
| 41 | Grok-2 Mini | xai | 77.9 |
| 42 | Microsoft WizardLM 2 8x22B | other | 77.9 |
| 43 | Qwen1.5 32B | alibaba | 77.9 |
| 44 | GPT-4 | openai | 77.6 |
| 45 | StableCode 3B | other | 77 |
| 46 | GPT-4 32K | openai | 76.8 |
| 47 | Qwen1.5 110B | alibaba | 76.5 |
| 48 | Qwen2 72B | alibaba | 76.5 |
| 49 | Code Llama 7B | meta | 76.4 |
| 50 | Gemini 1.0 Pro | 76.4 | |
| 51 | StableLM 2 12B | other | 76.4 |
| 52 | Code Bison | 76.2 | |
| 53 | Yi Vision | other | 76.1 |
| 54 | DeepSeek Coder 7B | deepseek | 75.9 |
| 55 | Jamba 1.5 | other | 75.9 |
| 56 | Command R | cohere | 75.8 |
| 57 | GPT-4 Vision | openai | 75.8 |
| 58 | Hermes 3 Llama 3.1 70B | other | 75.8 |
| 59 | Sonar Huge | other | 75.8 |
| 60 | DeepSeek V2 | deepseek | 75.7 |
| 61 | Command R+ (08-2024) | cohere | 75 |
| 62 | Claude 3 Haiku (2024-03-07) | anthropic | 74.8 |
| 63 | Command Nightly | cohere | 74.7 |
| 64 | Zephyr ORPO 141B Alpha | other | 74.7 |
| 65 | StarCoder2 7B | other | 74.6 |
| 66 | Claude 3 Sonnet | anthropic | 74.5 |
| 67 | Yi 1.5 34B | other | 74.5 |
| 68 | Mistral Small 3 | mistral | 74.4 |
| 69 | Orca 2 13B | other | 74.3 |
| 70 | Mistral Small | mistral | 74 |
| 71 | Claude 3.5 Haiku | anthropic | 73.8 |
| 72 | Jamba 1.5 Large | other | 73.7 |
| 73 | Command R (08-2024) | cohere | 73.6 |
| 74 | StarCoder2 15B | other | 73.5 |
| 75 | Mistral Large | mistral | 73.4 |
| 76 | Nous Hermes 2 Yi 34B | other | 73.2 |
| 77 | Llama 3.2 90B Vision | meta | 73.1 |
| 78 | Llama 3 70B | meta | 73 |
| 79 | DBRX Instruct | other | 72.8 |
| 80 | Sonar Small | other | 72.8 |
| 81 | GLM-4 Plus | other | 72.7 |
| 82 | Gemma 2 27B | 72.4 | |
| 83 | Yi Large Turbo | other | 72.4 |
| 84 | Yi Large | other | 72.2 |
| 85 | DBRX Base | other | 72.1 |
| 86 | Qwen2.5 14B | alibaba | 72 |
| 87 | Gemini 1.5 Pro | 71.9 | |
| 88 | DeepSeek Coder 33B | deepseek | 71.8 |
| 89 | Code Llama 34B | meta | 71.7 |
| 90 | Command R+ | cohere | 70.7 |
| 91 | Gemini 1.5 Flash-8B | 70.7 | |
| 92 | GPT-4 Vision Preview | openai | 70.6 |
| 93 | WizardLM Team WizardLM 2 8x22B | other | 70.2 |
| 94 | Code Llama 13B | meta | 69.5 |
| 95 | Phi-3.5 MoE | other | 69.5 |
| 96 | Claude 3 Haiku | anthropic | 69.4 |
| 97 | Mistral Nemo | mistral | 69.4 |
| 98 | Codestral | mistral | 69.2 |
| 99 | StarChat2 15B v0.1 | other | 69 |
| 100 | Falcon 180B | other | 68.9 |
| 101 | Mistral 7B v0.3 | mistral | 68.8 |
| 102 | Microsoft WizardLM 2 7B | other | 68.6 |
| 103 | DeepSeek LLM 67B | deepseek | 68.3 |
| 104 | GLM-4 Air | other | 68.3 |
| 105 | Phi-3 Small | other | 68.3 |
| 106 | Qwen2.5 32B | alibaba | 68 |
| 107 | GLM-4 Flash | other | 67.6 |
| 108 | Qwen2 57B | alibaba | 67.6 |
| 109 | Gemini 1.5 Flash | 67.5 | |
| 110 | Hermes 3 Llama 3.1 8B | other | 67.1 |
| 111 | DeepSeek V2 Chat | deepseek | 67 |
| 112 | Jamba 1.5 Mini | other | 67 |
| 113 | OLMo 1.7 7B | other | 67 |
| 114 | Qwen2 7B | alibaba | 67 |
| 115 | Qwen2.5 7B | alibaba | 66.5 |
| 116 | NVIDIA Llama 3.1 Nemotron 70B | other | 66.4 |
| 117 | Gemini 1.5 Flash 002 | 66.2 | |
| 118 | GLM-4V 9B | other | 66 |
| 119 | Jamba Instruct | other | 66 |
| 120 | Qwen1.5 14B | alibaba | 65.8 |
| 121 | Gemini 1.0 Flash | 65.7 | |
| 122 | Code Llama 70B | meta | 65.6 |
| 123 | Qwen1.5 72B | alibaba | 65.2 |
| 124 | Llama 3 8B | meta | 65 |
| 125 | Llemma 7B | other | 64.2 |
| 126 | Llama 3.1 8B | meta | 63.4 |
| 127 | Ministral 8B | mistral | 63.3 |
| 128 | Yi 1.5 6B | other | 63.1 |
| 129 | GLM-4 9B Chat | other | 62.9 |
| 130 | Phi-4 | other | 62.8 |
| 131 | OLMo 7B | other | 62.2 |
| 132 | OpenChat 3.6 8B | other | 61.5 |
| 133 | Microsoft WizardMath 7B v1 | other | 61.1 |
| 134 | Gemma 7B | 61 | |
| 135 | Gemma 2 9B | 60.9 | |
| 136 | Mistral 7B v0.2 | mistral | 59.7 |
| 137 | Nous Hermes 2 Solar 10.7B | other | 59.2 |
| 138 | Yi 1.5 9B | other | 57.3 |
| 139 | Mistral 7B v0.1 | mistral | 56.4 |
| 140 | DeepSeek Math 7B | deepseek | 54 |
| 141 | OLMo 7B SFT | other | 53.9 |
| 142 | Phi-3 Medium | other | 53.9 |
| 143 | OLMo 2 1124 7B | other | 53.2 |
| 144 | GPT-3.5 Turbo | openai | 51.3 |
| 145 | Command Light | cohere | 51.2 |
| 146 | ChatGLM3 6B | other | 51 |
| 147 | Jurassic-2 Ultra | other | 50.9 |
| 148 | Llama 2 70B | meta | 50.6 |
| 149 | OLMo 7B Instruct | other | 50.5 |
| 150 | Command R7B | cohere | 50.1 |
| 151 | Llama 3.2 11B Vision | meta | 50.1 |
| 152 | Nous Capybara 34B | other | 50.1 |
| 153 | Mathstral 7B | mistral | 49.9 |
| 154 | MPT 30B | other | 49.7 |
| 155 | Baichuan2 13B Chat | other | 49.6 |
| 156 | Baichuan2 7B Chat | other | 49.2 |
| 157 | Qwen2.5 3B | alibaba | 49 |
| 158 | Open-Platypus | other | 48.2 |
| 159 | Phi-3 Vision | other | 48.2 |
| 160 | Qwen2 1.5B | alibaba | 47.2 |
| 161 | StableLM Zephyr 3B | other | 47.1 |
| 162 | Flan-UL2 | other | 46.8 |
| 163 | Mistral Tiny | mistral | 45.8 |
| 164 | Mixtral 8x22B | mistral | 45.2 |
| 165 | Text Bison | 45.1 | |
| 166 | PaLM 2 | 44.8 | |
| 167 | Grok Vision Beta | xai | 44 |
| 168 | Phi-3.5 Mini | other | 44 |
| 169 | Gemma 2B | 43.2 | |
| 170 | Llama 2 7B | meta | 42.6 |
| 171 | Llama 3.2 3B | meta | 42.5 |
| 172 | Yi 6B | other | 42.3 |
| 173 | OpenChat 3.5 1210 | other | 41.9 |
| 174 | Llama 2 13B | meta | 40.7 |
| 175 | Falcon 40B | other | 39.7 |
| 176 | Jurassic-2 Mid | other | 39.7 |
| 177 | Phi-2 | other | 39.7 |
| 178 | Phi-1.5 | other | 39.2 |
| 179 | Nous Capybara 7B | other | 38.8 |
| 180 | Zephyr 7B Beta | other | 38.4 |
| 181 | MPT 7B | other | 38.1 |
| 182 | StableLM 2 1.6B | other | 37.3 |
| 183 | Ministral 3B | mistral | 37.2 |
| 184 | Argilla Notus 7B v1 | other | 36.8 |
| 185 | GPT-3.5 Turbo 16K | openai | 36.8 |
| 186 | Chat Bison | 36 | |
| 187 | TigerBot 70B Chat | other | 36 |
| 188 | Zephyr 7B Alpha | other | 35.9 |
| 189 | Yi 34B | other | 35.8 |
| 190 | Flan-T5 XXL | other | 35.4 |
| 191 | Qwen2.5 1.5B | alibaba | 35.2 |
| 192 | Flan-T5 XL | other | 34 |
| 193 | StableLM 3 4B | other | 32.6 |
| 194 | Claude 2 | anthropic | 31.7 |
| 195 | Grok Beta | xai | 31 |
| 196 | Qwen2.5 0.5B | alibaba | 31 |
| 197 | Capybara 1.5B | other | 30.1 |
| 198 | Llama 3.2 1B | meta | 29.9 |
| 199 | Claude Instant 1 | anthropic | 28.5 |
| 200 | GPT-3.5 | openai | 28.2 |
| 201 | Claude 2.1 | anthropic | 28.1 |
| 202 | Llama Guard 3 8B | meta | 28 |
| 203 | Phi-3 Mini | other | 26.9 |
| 204 | Llama Guard 2 8B | meta | 26.6 |
| 205 | Embed English v3 | cohere | 0 |
HumanEval
描述
OpenAI 发布的 164 道 Python 编程任务,每题包含函数签名、文档字符串、函数体与单元测试,评估模型的代码生成能力(pass@1)。
核心规格
| 分类 | 许可证 | 最后更新 |
|---|---|---|
| coding | MIT | 2022-01-01 |
基准
| 单位 |
|---|
| pass@1 |