IFEval:基准测试详解
详细介绍 IFEval 基准:分类、评估指标、数据来源与适用模型。
基准介绍
评估大型语言模型严格遵循指令的能力,包含500个可验证的指令
评测指标
| 指标名 | 单位 | 方向 |
|---|---|---|
| prompt_strict_acc | % | ↑ 越高越好 |
| prompt_loose_acc | % | ↑ 越高越好 |
数据来源
模型得分排名
| # | 模型 | 厂商 | 得分 |
|---|---|---|---|
| 1 | Gemini 1.5 Pro 002 | 87.9 | |
| 2 | GPT-4o (2024-05-13) | openai | 86.7 |
| 3 | Gemini 2.0 Flash | 85.9 | |
| 4 | GPT-4o (2024-08-06) | openai | 85.5 |
| 5 | o1 Preview | openai | 84.4 |
| 6 | Qwen1.5 110B | alibaba | 83.4 |
| 7 | Sonar Reasoning | other | 83.1 |
| 8 | Gemini 2.0 Flash Thinking | 82.8 | |
| 9 | o1 | openai | 82.2 |
| 10 | Claude 3.5 Sonnet (2024-10-22) | anthropic | 81.6 |
| 11 | Gemini 1.0 Ultra | 81.6 | |
| 12 | Yi Vision | other | 81.4 |
| 13 | Command Nightly | cohere | 80.8 |
| 14 | StableLM 2 12B | other | 80 |
| 15 | Qwen1.5 14B | alibaba | 79.9 |
| 16 | DBRX Base | other | 79.3 |
| 17 | Llama 3.3 70B | meta | 79.3 |
| 18 | Yi 1.5 34B | other | 79.3 |
| 19 | Claude 3 Opus | anthropic | 79.2 |
| 20 | GLM-4 Plus | other | 79.2 |
| 21 | Jamba Instruct | other | 79.2 |
| 22 | Sonar Huge | other | 78.7 |
| 23 | DeepSeek V2 | deepseek | 78.6 |
| 24 | Hermes 3 Llama 3.1 405B | other | 78.2 |
| 25 | GPT-4 Turbo | openai | 77.9 |
| 26 | Mistral Medium | mistral | 77.4 |
| 27 | Gemini 1.5 Flash 002 | 77.2 | |
| 28 | Grok-2 Vision | xai | 76.9 |
| 29 | Mistral Small 3 | mistral | 76.9 |
| 30 | Qwen2 72B | alibaba | 76.9 |
| 31 | DeepSeek LLM 67B | deepseek | 76.8 |
| 32 | Qwen1.5 72B | alibaba | 76.4 |
| 33 | Mistral Large | mistral | 75.8 |
| 34 | GLM-4 Air | other | 75.7 |
| 35 | Claude 3 Opus (2024-02-29) | anthropic | 75.5 |
| 36 | Claude 3 Sonnet (2024-02-29) | anthropic | 75.5 |
| 37 | Claude 3 Haiku (2024-03-07) | anthropic | 75.3 |
| 38 | Microsoft WizardLM 2 8x22B | other | 75.3 |
| 39 | Nous Hermes 2 Mixtral 8x7B | other | 75.3 |
| 40 | Gemini 1.0 Pro | 75.1 | |
| 41 | Sonar Large | other | 74.9 |
| 42 | GPT-4 32K | openai | 74.7 |
| 43 | GPT-4 | openai | 74.6 |
| 44 | Llama 3.1 Nemotron 70B | meta | 74.5 |
| 45 | Llama 3.2 90B Vision | meta | 74.5 |
| 46 | Claude 3 Sonnet | anthropic | 74.3 |
| 47 | Gemini 1.5 Flash-8B 002 | 74.3 | |
| 48 | Phi-3.5 MoE | other | 74.3 |
| 49 | Orca 2 13B | other | 74.2 |
| 50 | Command R | cohere | 74.1 |
| 51 | Qwen2.5 14B | alibaba | 74.1 |
| 52 | Mistral Small | mistral | 74 |
| 53 | GPT-4 1106 Preview | openai | 73.9 |
| 54 | Hermes 3 Llama 3.1 70B | other | 73.8 |
| 55 | Llama 3.1 70B | meta | 73.7 |
| 56 | Claude 3.5 Haiku | anthropic | 73.6 |
| 57 | GPT-4 Vision | openai | 73.5 |
| 58 | Grok-2 Mini | xai | 73.5 |
| 59 | Qwen2.5 32B | alibaba | 73.5 |
| 60 | Yi Large | other | 73.4 |
| 61 | DBRX Instruct | other | 73.1 |
| 62 | o1 mini | openai | 73.1 |
| 63 | GPT-4 Vision Preview | openai | 72.7 |
| 64 | Gemma 2 27B | 72.5 | |
| 65 | Llama 3 70B | meta | 72.2 |
| 66 | Command R+ (08-2024) | cohere | 72.1 |
| 67 | Mixtral 8x7B | mistral | 72.1 |
| 68 | Sonar Small | other | 72.1 |
| 69 | Jamba 1.5 Mini | other | 71.9 |
| 70 | Gemini 1.0 Flash | 71.8 | |
| 71 | DeepSeek V2 Chat | deepseek | 71.5 |
| 72 | GPT-4 0125 Preview | openai | 71.4 |
| 73 | Qwen2 57B | alibaba | 71.2 |
| 74 | Claude 3 Haiku | anthropic | 70.8 |
| 75 | Jamba 1.5 | other | 70.8 |
| 76 | Jamba 1.5 Large | other | 70.8 |
| 77 | WizardLM Team WizardLM 2 8x22B | other | 70.7 |
| 78 | Yi Large Turbo | other | 70.7 |
| 79 | Falcon 180B | other | 70.5 |
| 80 | Command R (08-2024) | cohere | 70.3 |
| 81 | Zephyr ORPO 141B Alpha | other | 69.7 |
| 82 | NVIDIA Llama 3.1 Nemotron 70B | other | 69.6 |
| 83 | OLMo 7B | other | 68.4 |
| 84 | Nous Hermes 2 Yi 34B | other | 68.3 |
| 85 | Llama 3 8B | meta | 68.2 |
| 86 | Qwen1.5 32B | alibaba | 68.2 |
| 87 | Gemini 1.5 Flash-8B | 68.1 | |
| 88 | Gemini 1.5 Flash | 68 | |
| 89 | GPT-4o mini | openai | 67.9 |
| 90 | GLM-4V 9B | other | 67.4 |
| 91 | Phi-3 Medium | other | 67.4 |
| 92 | Command R7B | cohere | 67.2 |
| 93 | Gemma 7B | 66.9 | |
| 94 | GLM-4 Flash | other | 65.9 |
| 95 | Codestral Mamba | mistral | 65.5 |
| 96 | Yi 1.5 9B | other | 65.4 |
| 97 | Gemma 2 9B | 64.3 | |
| 98 | Code Llama 70B | meta | 63.8 |
| 99 | Mathstral 7B | mistral | 63.6 |
| 100 | OLMo 7B Instruct | other | 63.6 |
| 101 | StarChat2 15B v0.1 | other | 63.3 |
| 102 | Llama 3.1 8B | meta | 62.1 |
| 103 | Qwen2 7B | alibaba | 61.8 |
| 104 | Code Llama 7B | meta | 61.2 |
| 105 | OLMo 1.7 7B | other | 61.2 |
| 106 | DeepSeek Math 7B | deepseek | 61.1 |
| 107 | Phi-3 Small | other | 61 |
| 108 | Mistral 7B v0.3 | mistral | 60.9 |
| 109 | Hermes 3 Llama 3.1 8B | other | 60.7 |
| 110 | GLM-4 9B Chat | other | 59.9 |
| 111 | MPT 7B | other | 59.9 |
| 112 | Phi-4 | other | 59.6 |
| 113 | StarCoder2 15B | other | 59.6 |
| 114 | Text Bison | 59.5 | |
| 115 | Codestral | mistral | 59.4 |
| 116 | Code Bison | 59.2 | |
| 117 | Mistral 7B v0.2 | mistral | 59.2 |
| 118 | Mistral 7B v0.1 | mistral | 58.9 |
| 119 | DeepSeek Coder 33B | deepseek | 58.8 |
| 120 | Yi 1.5 6B | other | 58.8 |
| 121 | Code Llama 34B | meta | 58.7 |
| 122 | Llama 2 7B | meta | 58.5 |
| 123 | Llama 3.2 11B Vision | meta | 58.4 |
| 124 | Microsoft WizardMath 7B v1 | other | 57.7 |
| 125 | GPT-3.5 Turbo | openai | 57.4 |
| 126 | Qwen2 1.5B | alibaba | 57.4 |
| 127 | Falcon 40B | other | 57.2 |
| 128 | Llama 2 70B | meta | 57.1 |
| 129 | Microsoft WizardLM 2 7B | other | 56.9 |
| 130 | DeepSeek Coder V2 | deepseek | 56.7 |
| 131 | OpenChat 3.6 8B | other | 56.7 |
| 132 | Phi-3 Vision | other | 56.4 |
| 133 | Grok Beta | xai | 56.2 |
| 134 | OLMo 7B SFT | other | 56.2 |
| 135 | Ministral 8B | mistral | 56.1 |
| 136 | Mistral Nemo | mistral | 56 |
| 137 | OLMo 2 1124 7B | other | 56 |
| 138 | Qwen2.5 3B | alibaba | 56 |
| 139 | Flan-T5 XL | other | 55.7 |
| 140 | TigerBot 70B Chat | other | 55.7 |
| 141 | Flan-UL2 | other | 55.6 |
| 142 | Nous Hermes 2 Solar 10.7B | other | 55.6 |
| 143 | DeepSeek Coder 7B | deepseek | 55.4 |
| 144 | Qwen2.5 7B | alibaba | 55.4 |
| 145 | StableLM Zephyr 3B | other | 55.3 |
| 146 | StarCoder2 3B | other | 55.2 |
| 147 | Claude 2 | anthropic | 55 |
| 148 | Claude 2.1 | anthropic | 55 |
| 149 | Phi-3 Mini | other | 54.9 |
| 150 | Yi 6B | other | 54.9 |
| 151 | StarCoder2 7B | other | 54 |
| 152 | Command Light | cohere | 53.9 |
| 153 | Microsoft WizardCoder Python 34B | other | 53.8 |
| 154 | Phi-1 | other | 53.8 |
| 155 | Claude Instant 1 | anthropic | 53.2 |
| 156 | Llemma 7B | other | 52.9 |
| 157 | Nous Capybara 34B | other | 52.9 |
| 158 | GPT-3.5 | openai | 52.8 |
| 159 | Capybara 1.5B | other | 52.4 |
| 160 | StableCode 3B | other | 51.8 |
| 161 | Mistral Tiny | mistral | 51.7 |
| 162 | Llama Guard 2 8B | meta | 51.6 |
| 163 | StableLM 3 4B | other | 51.6 |
| 164 | Flan-T5 XXL | other | 51.3 |
| 165 | MPT 30B | other | 51.2 |
| 166 | Zephyr 7B Beta | other | 51.2 |
| 167 | Zephyr 7B Alpha | other | 50.9 |
| 168 | OpenChat 3.5 1210 | other | 50.4 |
| 169 | Baichuan2 7B Chat | other | 50.1 |
| 170 | Code Llama 13B | meta | 50.1 |
| 171 | Llama Guard 3 8B | meta | 50 |
| 172 | Llama 2 13B | meta | 49.9 |
| 173 | Jurassic-2 Mid | other | 49.8 |
| 174 | Phi-2 | other | 48.8 |
| 175 | Phi-3.5 Mini | other | 48.7 |
| 176 | Llama 3.2 3B | meta | 48.6 |
| 177 | Baichuan2 13B Chat | other | 48.5 |
| 178 | GPT-3.5 Turbo 16K | openai | 48.5 |
| 179 | Qwen2.5 0.5B | alibaba | 48.5 |
| 180 | Chat Bison | 48.2 | |
| 181 | Yi 34B | other | 48.2 |
| 182 | Phi-1.5 | other | 48 |
| 183 | Jurassic-2 Ultra | other | 47.9 |
| 184 | Ministral 3B | mistral | 47.8 |
| 185 | Grok Vision Beta | xai | 46.8 |
| 186 | Gemma 2B | 46.4 | |
| 187 | Open-Platypus | other | 46 |
| 188 | PaLM 2 | 44.7 | |
| 189 | Nous Capybara 7B | other | 44.2 |
| 190 | Llama 3.2 1B | meta | 44.1 |
| 191 | StableLM 2 1.6B | other | 44.1 |
| 192 | Argilla Notus 7B v1 | other | 42.8 |
| 193 | ChatGLM3 6B | other | 42.5 |
| 194 | Qwen2.5 1.5B | alibaba | 41.9 |
| 195 | Embed English v3 | cohere | 0 |
IFEval
描述
评估大型语言模型严格遵循指令的能力,包含500个可验证的指令
核心规格
| 分类 | 许可证 | 最后更新 |
|---|---|---|
| reasoning | Apache-2.0 | 2024-02-01 |
基准
| 单位 |
|---|
| % |
| % |