基准介绍

评估大型语言模型严格遵循指令的能力,包含500个可验证的指令

评测指标

指标名 单位 方向
prompt_strict_acc % ↑ 越高越好
prompt_loose_acc % ↑ 越高越好

数据来源

模型得分排名

# 模型 厂商 得分
1 Gemini 1.5 Pro 002 google 87.9
2 GPT-4o (2024-05-13) openai 86.7
3 Gemini 2.0 Flash google 85.9
4 GPT-4o (2024-08-06) openai 85.5
5 o1 Preview openai 84.4
6 Qwen1.5 110B alibaba 83.4
7 Sonar Reasoning other 83.1
8 Gemini 2.0 Flash Thinking google 82.8
9 o1 openai 82.2
10 Claude 3.5 Sonnet (2024-10-22) anthropic 81.6
11 Gemini 1.0 Ultra google 81.6
12 Yi Vision other 81.4
13 Command Nightly cohere 80.8
14 StableLM 2 12B other 80
15 Qwen1.5 14B alibaba 79.9
16 DBRX Base other 79.3
17 Llama 3.3 70B meta 79.3
18 Yi 1.5 34B other 79.3
19 Claude 3 Opus anthropic 79.2
20 GLM-4 Plus other 79.2
21 Jamba Instruct other 79.2
22 Sonar Huge other 78.7
23 DeepSeek V2 deepseek 78.6
24 Hermes 3 Llama 3.1 405B other 78.2
25 GPT-4 Turbo openai 77.9
26 Mistral Medium mistral 77.4
27 Gemini 1.5 Flash 002 google 77.2
28 Grok-2 Vision xai 76.9
29 Mistral Small 3 mistral 76.9
30 Qwen2 72B alibaba 76.9
31 DeepSeek LLM 67B deepseek 76.8
32 Qwen1.5 72B alibaba 76.4
33 Mistral Large mistral 75.8
34 GLM-4 Air other 75.7
35 Claude 3 Opus (2024-02-29) anthropic 75.5
36 Claude 3 Sonnet (2024-02-29) anthropic 75.5
37 Claude 3 Haiku (2024-03-07) anthropic 75.3
38 Microsoft WizardLM 2 8x22B other 75.3
39 Nous Hermes 2 Mixtral 8x7B other 75.3
40 Gemini 1.0 Pro google 75.1
41 Sonar Large other 74.9
42 GPT-4 32K openai 74.7
43 GPT-4 openai 74.6
44 Llama 3.1 Nemotron 70B meta 74.5
45 Llama 3.2 90B Vision meta 74.5
46 Claude 3 Sonnet anthropic 74.3
47 Gemini 1.5 Flash-8B 002 google 74.3
48 Phi-3.5 MoE other 74.3
49 Orca 2 13B other 74.2
50 Command R cohere 74.1
51 Qwen2.5 14B alibaba 74.1
52 Mistral Small mistral 74
53 GPT-4 1106 Preview openai 73.9
54 Hermes 3 Llama 3.1 70B other 73.8
55 Llama 3.1 70B meta 73.7
56 Claude 3.5 Haiku anthropic 73.6
57 GPT-4 Vision openai 73.5
58 Grok-2 Mini xai 73.5
59 Qwen2.5 32B alibaba 73.5
60 Yi Large other 73.4
61 DBRX Instruct other 73.1
62 o1 mini openai 73.1
63 GPT-4 Vision Preview openai 72.7
64 Gemma 2 27B google 72.5
65 Llama 3 70B meta 72.2
66 Command R+ (08-2024) cohere 72.1
67 Mixtral 8x7B mistral 72.1
68 Sonar Small other 72.1
69 Jamba 1.5 Mini other 71.9
70 Gemini 1.0 Flash google 71.8
71 DeepSeek V2 Chat deepseek 71.5
72 GPT-4 0125 Preview openai 71.4
73 Qwen2 57B alibaba 71.2
74 Claude 3 Haiku anthropic 70.8
75 Jamba 1.5 other 70.8
76 Jamba 1.5 Large other 70.8
77 WizardLM Team WizardLM 2 8x22B other 70.7
78 Yi Large Turbo other 70.7
79 Falcon 180B other 70.5
80 Command R (08-2024) cohere 70.3
81 Zephyr ORPO 141B Alpha other 69.7
82 NVIDIA Llama 3.1 Nemotron 70B other 69.6
83 OLMo 7B other 68.4
84 Nous Hermes 2 Yi 34B other 68.3
85 Llama 3 8B meta 68.2
86 Qwen1.5 32B alibaba 68.2
87 Gemini 1.5 Flash-8B google 68.1
88 Gemini 1.5 Flash google 68
89 GPT-4o mini openai 67.9
90 GLM-4V 9B other 67.4
91 Phi-3 Medium other 67.4
92 Command R7B cohere 67.2
93 Gemma 7B google 66.9
94 GLM-4 Flash other 65.9
95 Codestral Mamba mistral 65.5
96 Yi 1.5 9B other 65.4
97 Gemma 2 9B google 64.3
98 Code Llama 70B meta 63.8
99 Mathstral 7B mistral 63.6
100 OLMo 7B Instruct other 63.6
101 StarChat2 15B v0.1 other 63.3
102 Llama 3.1 8B meta 62.1
103 Qwen2 7B alibaba 61.8
104 Code Llama 7B meta 61.2
105 OLMo 1.7 7B other 61.2
106 DeepSeek Math 7B deepseek 61.1
107 Phi-3 Small other 61
108 Mistral 7B v0.3 mistral 60.9
109 Hermes 3 Llama 3.1 8B other 60.7
110 GLM-4 9B Chat other 59.9
111 MPT 7B other 59.9
112 Phi-4 other 59.6
113 StarCoder2 15B other 59.6
114 Text Bison google 59.5
115 Codestral mistral 59.4
116 Code Bison google 59.2
117 Mistral 7B v0.2 mistral 59.2
118 Mistral 7B v0.1 mistral 58.9
119 DeepSeek Coder 33B deepseek 58.8
120 Yi 1.5 6B other 58.8
121 Code Llama 34B meta 58.7
122 Llama 2 7B meta 58.5
123 Llama 3.2 11B Vision meta 58.4
124 Microsoft WizardMath 7B v1 other 57.7
125 GPT-3.5 Turbo openai 57.4
126 Qwen2 1.5B alibaba 57.4
127 Falcon 40B other 57.2
128 Llama 2 70B meta 57.1
129 Microsoft WizardLM 2 7B other 56.9
130 DeepSeek Coder V2 deepseek 56.7
131 OpenChat 3.6 8B other 56.7
132 Phi-3 Vision other 56.4
133 Grok Beta xai 56.2
134 OLMo 7B SFT other 56.2
135 Ministral 8B mistral 56.1
136 Mistral Nemo mistral 56
137 OLMo 2 1124 7B other 56
138 Qwen2.5 3B alibaba 56
139 Flan-T5 XL other 55.7
140 TigerBot 70B Chat other 55.7
141 Flan-UL2 other 55.6
142 Nous Hermes 2 Solar 10.7B other 55.6
143 DeepSeek Coder 7B deepseek 55.4
144 Qwen2.5 7B alibaba 55.4
145 StableLM Zephyr 3B other 55.3
146 StarCoder2 3B other 55.2
147 Claude 2 anthropic 55
148 Claude 2.1 anthropic 55
149 Phi-3 Mini other 54.9
150 Yi 6B other 54.9
151 StarCoder2 7B other 54
152 Command Light cohere 53.9
153 Microsoft WizardCoder Python 34B other 53.8
154 Phi-1 other 53.8
155 Claude Instant 1 anthropic 53.2
156 Llemma 7B other 52.9
157 Nous Capybara 34B other 52.9
158 GPT-3.5 openai 52.8
159 Capybara 1.5B other 52.4
160 StableCode 3B other 51.8
161 Mistral Tiny mistral 51.7
162 Llama Guard 2 8B meta 51.6
163 StableLM 3 4B other 51.6
164 Flan-T5 XXL other 51.3
165 MPT 30B other 51.2
166 Zephyr 7B Beta other 51.2
167 Zephyr 7B Alpha other 50.9
168 OpenChat 3.5 1210 other 50.4
169 Baichuan2 7B Chat other 50.1
170 Code Llama 13B meta 50.1
171 Llama Guard 3 8B meta 50
172 Llama 2 13B meta 49.9
173 Jurassic-2 Mid other 49.8
174 Phi-2 other 48.8
175 Phi-3.5 Mini other 48.7
176 Llama 3.2 3B meta 48.6
177 Baichuan2 13B Chat other 48.5
178 GPT-3.5 Turbo 16K openai 48.5
179 Qwen2.5 0.5B alibaba 48.5
180 Chat Bison google 48.2
181 Yi 34B other 48.2
182 Phi-1.5 other 48
183 Jurassic-2 Ultra other 47.9
184 Ministral 3B mistral 47.8
185 Grok Vision Beta xai 46.8
186 Gemma 2B google 46.4
187 Open-Platypus other 46
188 PaLM 2 google 44.7
189 Nous Capybara 7B other 44.2
190 Llama 3.2 1B meta 44.1
191 StableLM 2 1.6B other 44.1
192 Argilla Notus 7B v1 other 42.8
193 ChatGLM3 6B other 42.5
194 Qwen2.5 1.5B alibaba 41.9
195 Embed English v3 cohere 0

IFEval

描述

评估大型语言模型严格遵循指令的能力,包含500个可验证的指令

核心规格

分类 许可证 最后更新
reasoning Apache-2.0 2024-02-01

基准

单位
%
%

数据来源

官方地址