Overview

BIG-Bench 的 23 个最难任务子集,覆盖逻辑推理、符号操作、常识推理等领域,传统模型表现不佳,用于评估模型的复杂推理能力。

Metrics

Metric Unit Direction
accuracy % ↑ Higher is better

Sources

Model Score Ranking

# Model Vendor Score
1 Gemini 2.0 Flash Thinking google 87.9
2 Gemini 1.5 Pro 002 google 87.2
3 GPT-4o (2024-05-13) openai 86.7
4 DeepSeek V3 deepseek 84.9
5 GPT-4o (2024-08-06) openai 84.9
6 Claude 3.5 Sonnet anthropic 84.5
7 o1 Preview openai 84.5
8 Claude 3.5 Sonnet (2024-10-22) anthropic 84.2
9 Gemini 1.5 Pro google 84
10 Grok-2 xai 84
11 Llama 3.3 70B meta 83.9
12 Gemini 1.0 Ultra google 83.8
13 GLM-4 Plus other 83.6
14 Gemini 2.0 Flash google 83.4
15 Mistral Medium mistral 83.4
16 Qwen2 72B alibaba 83.4
17 Sonar Reasoning other 83.4
18 GPT-4o openai 83.1
19 o1 openai 83.1
20 Llama 3.1 405B meta 82.9
21 Jamba 1.5 Large other 82.8
22 Command Nightly cohere 82.7
23 Qwen2.5 72B alibaba 82.4
24 Qwen1.5 110B alibaba 82.3
25 Command R+ (08-2024) cohere 82.1
26 Hermes 3 Llama 3.1 405B other 82.1
27 Claude 3 Opus anthropic 81.6
28 Sonar Large other 81.6
29 Mistral Large 2 mistral 81
30 GPT-4 openai 80.9
31 GPT-4 Vision openai 80.9
32 GPT-4 1106 Preview openai 80
33 WizardLM Team WizardLM 2 8x22B other 80
34 Gemini 1.5 Flash 002 google 79.8
35 Claude 3 Opus (2024-02-29) anthropic 79.5
36 Phi-3.5 MoE other 79.4
37 GPT-4 32K openai 79.2
38 Gemini 1.0 Pro google 79.1
39 GLM-4 Air other 79
40 Command R (08-2024) cohere 78.8
41 Sonar Huge other 78.6
42 GPT-4 Turbo openai 78.5
43 Command R cohere 78.4
44 Mixtral 8x7B mistral 78.4
45 Yi Vision other 78.3
46 GPT-4 0125 Preview openai 78.2
47 Llama 3 70B meta 77.6
48 Jamba 1.5 other 77.3
49 o1 mini openai 77.2
50 Yi Large other 77.1
51 Qwen2.5 32B alibaba 76.8
52 Mistral Large mistral 76.7
53 NVIDIA Llama 3.1 Nemotron 70B other 76.6
54 Qwen1.5 72B alibaba 76.5
55 Falcon 180B other 76.4
56 Yi Large Turbo other 76.4
57 Claude 3 Sonnet (2024-02-29) anthropic 76.2
58 Claude 3 Sonnet anthropic 76.1
59 Llama 3.2 90B Vision meta 76.1
60 Hermes 3 Llama 3.1 70B other 75.9
61 Claude 3 Haiku anthropic 75.7
62 Jamba 1.5 Mini other 75.6
63 Nous Hermes 2 Yi 34B other 75.6
64 DeepSeek V2 Chat deepseek 75.4
65 Grok-2 Vision xai 75.4
66 GPT-4 Vision Preview openai 75
67 Gemini 1.0 Flash google 74.8
68 Gemini 1.5 Flash-8B google 74.8
69 Grok-2 Mini xai 74.6
70 Mixtral 8x22B mistral 74.5
71 Microsoft WizardMath 7B v1 other 74.4
72 Yi 1.5 34B other 74.1
73 Gemini 1.5 Flash-8B 002 google 73.6
74 Microsoft WizardLM 2 8x22B other 73.6
75 Llemma 7B other 73.1
76 Qwen1.5 32B alibaba 73
77 DeepSeek LLM 67B deepseek 72.6
78 Nous Hermes 2 Mixtral 8x7B other 72.4
79 Qwen2 57B alibaba 72.4
80 DBRX Base other 72.3
81 DeepSeek Coder V2 deepseek 72.3
82 Claude 3 Haiku (2024-03-07) anthropic 72.2
83 Qwen2.5 14B alibaba 72.2
84 Code Llama 70B meta 72.1
85 Jamba Instruct other 71.8
86 Nous Hermes 2 Solar 10.7B other 71.8
87 Phi-3 Small other 71.8
88 DBRX Instruct other 71.7
89 Gemini 1.5 Flash google 71.7
90 Sonar Small other 71.7
91 Mistral Small mistral 71.6
92 Qwen1.5 14B alibaba 71.6
93 OLMo 7B Instruct other 71.5
94 Code Bison google 71.4
95 Phi-3 Medium other 71.4
96 Mistral 7B v0.2 mistral 71.3
97 Mistral Small 3 mistral 71.3
98 Gemma 2 27B google 71.2
99 Mistral Nemo mistral 71.2
100 Code Llama 13B meta 71
101 OpenChat 3.6 8B other 71
102 Orca 2 13B other 70.8
103 StableLM 2 12B other 70.6
104 StarChat2 15B v0.1 other 70.6
105 DeepSeek V2 deepseek 70.5
106 GLM-4 Flash other 70.5
107 Yi 1.5 6B other 70.5
108 GPT-4o mini openai 70.3
109 Llama 3.1 Nemotron 70B meta 70.3
110 Zephyr ORPO 141B Alpha other 70.3
111 Claude 3.5 Haiku anthropic 70.2
112 Llama 3.1 70B meta 70.2
113 Phi-1 other 69.9
114 Mathstral 7B mistral 69.8
115 Microsoft WizardLM 2 7B other 69.4
116 Llama 3.2 11B Vision meta 69.1
117 OLMo 1.7 7B other 68.9
118 OLMo 7B other 67.7
119 GLM-4 9B Chat other 67.2
120 Microsoft WizardCoder Python 34B other 67.2
121 Llama 3.1 8B meta 67.1
122 Code Llama 7B meta 66.9
123 Phi-4 other 66.9
124 Gemma 2 9B google 66.7
125 Command R+ cohere 66.1
126 Command R7B cohere 66.1
127 Codestral Mamba mistral 66
128 DeepSeek Math 7B deepseek 65.9
129 Gemma 7B google 65.7
130 Mistral 7B v0.1 mistral 65.4
131 Yi 1.5 9B other 65.3
132 GLM-4V 9B other 65.1
133 Llama 3 8B meta 65
134 Baichuan2 13B Chat other 64.2
135 OLMo 7B SFT other 64
136 GPT-3.5 openai 63.9
137 Claude 2.1 anthropic 63.8
138 Open-Platypus other 63.7
139 Qwen2 7B alibaba 63.3
140 Claude Instant 1 anthropic 63.1
141 OLMo 2 1124 7B other 63.1
142 StarCoder2 15B other 62.8
143 Hermes 3 Llama 3.1 8B other 62.6
144 Llama 2 70B meta 62.5
145 Argilla Notus 7B v1 other 62.4
146 Llama 2 13B meta 62.1
147 StarCoder2 3B other 62.1
148 Qwen2.5 7B alibaba 61.9
149 DeepSeek Coder 33B deepseek 61.6
150 Flan-T5 XXL other 60.5
151 Mistral 7B v0.3 mistral 60.5
152 OpenChat 3.5 1210 other 60.4
153 Claude 2 anthropic 60.2
154 Ministral 8B mistral 60
155 StarCoder2 7B other 60
156 MPT 7B other 59.9
157 DeepSeek Coder 7B deepseek 59.7
158 Code Llama 34B meta 59.1
159 StableCode 3B other 58.4
160 Codestral mistral 58.3
161 Grok Vision Beta xai 56.8
162 Nous Capybara 34B other 56.1
163 Phi-3 Mini other 55.9
164 ChatGLM3 6B other 55.7
165 Command Light cohere 55.5
166 Yi 6B other 55.3
167 StableLM 3 4B other 54.9
168 Llama 2 7B meta 54.7
169 Chat Bison google 54.6
170 Baichuan2 7B Chat other 53
171 Gemma 2B google 53
172 Jurassic-2 Ultra other 52.9
173 PaLM 2 google 52.9
174 Flan-T5 XL other 52.5
175 TigerBot 70B Chat other 51.9
176 Jurassic-2 Mid other 51.8
177 Qwen2 1.5B alibaba 51.5
178 Qwen2.5 1.5B alibaba 51.5
179 Mistral Tiny mistral 51.3
180 Qwen2.5 0.5B alibaba 51
181 Zephyr 7B Alpha other 50.7
182 Falcon 40B other 50.6
183 GPT-3.5 Turbo 16K openai 50.2
184 Grok Beta xai 50.2
185 Ministral 3B mistral 50.1
186 Zephyr 7B Beta other 50
187 Nous Capybara 7B other 49.6
188 Phi-1.5 other 49.5
189 Flan-UL2 other 49.2
190 MPT 30B other 49.2
191 Qwen2.5 3B alibaba 48.8
192 Yi 34B other 48.8
193 Capybara 1.5B other 48.3
194 GPT-3.5 Turbo openai 48.2
195 Text Bison google 48.2
196 Phi-2 other 47.8
197 Llama Guard 2 8B meta 47.6
198 Phi-3.5 Mini other 47.5
199 Llama 3.2 3B meta 46
200 Llama Guard 3 8B meta 44.1
201 StableLM Zephyr 3B other 43.9
202 StableLM 2 1.6B other 40.7
203 Phi-3 Vision other 40.1
204 Llama 3.2 1B meta 38.8
205 Embed English v3 cohere 0

BBH (BIG-Bench Hard)

Description

BIG-Bench 的 23 个最难任务子集,覆盖逻辑推理、符号操作、常识推理等领域,传统模型表现不佳,用于评估模型的复杂推理能力。

Core Specifications

Category License Last Updated
reasoning MIT 2022-10-01

Benchmark

Unit
%

Sources

Official URL