BBH (BIG-Bench Hard):基准测试详解
详细介绍 BBH (BIG-Bench Hard) 基准:分类、评估指标、数据来源与适用模型。
基准介绍
BIG-Bench 的 23 个最难任务子集,覆盖逻辑推理、符号操作、常识推理等领域,传统模型表现不佳,用于评估模型的复杂推理能力。
评测指标
| 指标名 | 单位 | 方向 |
|---|---|---|
| accuracy | % | ↑ 越高越好 |
数据来源
模型得分排名
| # | 模型 | 厂商 | 得分 |
|---|---|---|---|
| 1 | Gemini 2.0 Flash Thinking | 87.9 | |
| 2 | Gemini 1.5 Pro 002 | 87.2 | |
| 3 | GPT-4o (2024-05-13) | openai | 86.7 |
| 4 | DeepSeek V3 | deepseek | 84.9 |
| 5 | GPT-4o (2024-08-06) | openai | 84.9 |
| 6 | Claude 3.5 Sonnet | anthropic | 84.5 |
| 7 | o1 Preview | openai | 84.5 |
| 8 | Claude 3.5 Sonnet (2024-10-22) | anthropic | 84.2 |
| 9 | Gemini 1.5 Pro | 84 | |
| 10 | Grok-2 | xai | 84 |
| 11 | Llama 3.3 70B | meta | 83.9 |
| 12 | Gemini 1.0 Ultra | 83.8 | |
| 13 | GLM-4 Plus | other | 83.6 |
| 14 | Gemini 2.0 Flash | 83.4 | |
| 15 | Mistral Medium | mistral | 83.4 |
| 16 | Qwen2 72B | alibaba | 83.4 |
| 17 | Sonar Reasoning | other | 83.4 |
| 18 | GPT-4o | openai | 83.1 |
| 19 | o1 | openai | 83.1 |
| 20 | Llama 3.1 405B | meta | 82.9 |
| 21 | Jamba 1.5 Large | other | 82.8 |
| 22 | Command Nightly | cohere | 82.7 |
| 23 | Qwen2.5 72B | alibaba | 82.4 |
| 24 | Qwen1.5 110B | alibaba | 82.3 |
| 25 | Command R+ (08-2024) | cohere | 82.1 |
| 26 | Hermes 3 Llama 3.1 405B | other | 82.1 |
| 27 | Claude 3 Opus | anthropic | 81.6 |
| 28 | Sonar Large | other | 81.6 |
| 29 | Mistral Large 2 | mistral | 81 |
| 30 | GPT-4 | openai | 80.9 |
| 31 | GPT-4 Vision | openai | 80.9 |
| 32 | GPT-4 1106 Preview | openai | 80 |
| 33 | WizardLM Team WizardLM 2 8x22B | other | 80 |
| 34 | Gemini 1.5 Flash 002 | 79.8 | |
| 35 | Claude 3 Opus (2024-02-29) | anthropic | 79.5 |
| 36 | Phi-3.5 MoE | other | 79.4 |
| 37 | GPT-4 32K | openai | 79.2 |
| 38 | Gemini 1.0 Pro | 79.1 | |
| 39 | GLM-4 Air | other | 79 |
| 40 | Command R (08-2024) | cohere | 78.8 |
| 41 | Sonar Huge | other | 78.6 |
| 42 | GPT-4 Turbo | openai | 78.5 |
| 43 | Command R | cohere | 78.4 |
| 44 | Mixtral 8x7B | mistral | 78.4 |
| 45 | Yi Vision | other | 78.3 |
| 46 | GPT-4 0125 Preview | openai | 78.2 |
| 47 | Llama 3 70B | meta | 77.6 |
| 48 | Jamba 1.5 | other | 77.3 |
| 49 | o1 mini | openai | 77.2 |
| 50 | Yi Large | other | 77.1 |
| 51 | Qwen2.5 32B | alibaba | 76.8 |
| 52 | Mistral Large | mistral | 76.7 |
| 53 | NVIDIA Llama 3.1 Nemotron 70B | other | 76.6 |
| 54 | Qwen1.5 72B | alibaba | 76.5 |
| 55 | Falcon 180B | other | 76.4 |
| 56 | Yi Large Turbo | other | 76.4 |
| 57 | Claude 3 Sonnet (2024-02-29) | anthropic | 76.2 |
| 58 | Claude 3 Sonnet | anthropic | 76.1 |
| 59 | Llama 3.2 90B Vision | meta | 76.1 |
| 60 | Hermes 3 Llama 3.1 70B | other | 75.9 |
| 61 | Claude 3 Haiku | anthropic | 75.7 |
| 62 | Jamba 1.5 Mini | other | 75.6 |
| 63 | Nous Hermes 2 Yi 34B | other | 75.6 |
| 64 | DeepSeek V2 Chat | deepseek | 75.4 |
| 65 | Grok-2 Vision | xai | 75.4 |
| 66 | GPT-4 Vision Preview | openai | 75 |
| 67 | Gemini 1.0 Flash | 74.8 | |
| 68 | Gemini 1.5 Flash-8B | 74.8 | |
| 69 | Grok-2 Mini | xai | 74.6 |
| 70 | Mixtral 8x22B | mistral | 74.5 |
| 71 | Microsoft WizardMath 7B v1 | other | 74.4 |
| 72 | Yi 1.5 34B | other | 74.1 |
| 73 | Gemini 1.5 Flash-8B 002 | 73.6 | |
| 74 | Microsoft WizardLM 2 8x22B | other | 73.6 |
| 75 | Llemma 7B | other | 73.1 |
| 76 | Qwen1.5 32B | alibaba | 73 |
| 77 | DeepSeek LLM 67B | deepseek | 72.6 |
| 78 | Nous Hermes 2 Mixtral 8x7B | other | 72.4 |
| 79 | Qwen2 57B | alibaba | 72.4 |
| 80 | DBRX Base | other | 72.3 |
| 81 | DeepSeek Coder V2 | deepseek | 72.3 |
| 82 | Claude 3 Haiku (2024-03-07) | anthropic | 72.2 |
| 83 | Qwen2.5 14B | alibaba | 72.2 |
| 84 | Code Llama 70B | meta | 72.1 |
| 85 | Jamba Instruct | other | 71.8 |
| 86 | Nous Hermes 2 Solar 10.7B | other | 71.8 |
| 87 | Phi-3 Small | other | 71.8 |
| 88 | DBRX Instruct | other | 71.7 |
| 89 | Gemini 1.5 Flash | 71.7 | |
| 90 | Sonar Small | other | 71.7 |
| 91 | Mistral Small | mistral | 71.6 |
| 92 | Qwen1.5 14B | alibaba | 71.6 |
| 93 | OLMo 7B Instruct | other | 71.5 |
| 94 | Code Bison | 71.4 | |
| 95 | Phi-3 Medium | other | 71.4 |
| 96 | Mistral 7B v0.2 | mistral | 71.3 |
| 97 | Mistral Small 3 | mistral | 71.3 |
| 98 | Gemma 2 27B | 71.2 | |
| 99 | Mistral Nemo | mistral | 71.2 |
| 100 | Code Llama 13B | meta | 71 |
| 101 | OpenChat 3.6 8B | other | 71 |
| 102 | Orca 2 13B | other | 70.8 |
| 103 | StableLM 2 12B | other | 70.6 |
| 104 | StarChat2 15B v0.1 | other | 70.6 |
| 105 | DeepSeek V2 | deepseek | 70.5 |
| 106 | GLM-4 Flash | other | 70.5 |
| 107 | Yi 1.5 6B | other | 70.5 |
| 108 | GPT-4o mini | openai | 70.3 |
| 109 | Llama 3.1 Nemotron 70B | meta | 70.3 |
| 110 | Zephyr ORPO 141B Alpha | other | 70.3 |
| 111 | Claude 3.5 Haiku | anthropic | 70.2 |
| 112 | Llama 3.1 70B | meta | 70.2 |
| 113 | Phi-1 | other | 69.9 |
| 114 | Mathstral 7B | mistral | 69.8 |
| 115 | Microsoft WizardLM 2 7B | other | 69.4 |
| 116 | Llama 3.2 11B Vision | meta | 69.1 |
| 117 | OLMo 1.7 7B | other | 68.9 |
| 118 | OLMo 7B | other | 67.7 |
| 119 | GLM-4 9B Chat | other | 67.2 |
| 120 | Microsoft WizardCoder Python 34B | other | 67.2 |
| 121 | Llama 3.1 8B | meta | 67.1 |
| 122 | Code Llama 7B | meta | 66.9 |
| 123 | Phi-4 | other | 66.9 |
| 124 | Gemma 2 9B | 66.7 | |
| 125 | Command R+ | cohere | 66.1 |
| 126 | Command R7B | cohere | 66.1 |
| 127 | Codestral Mamba | mistral | 66 |
| 128 | DeepSeek Math 7B | deepseek | 65.9 |
| 129 | Gemma 7B | 65.7 | |
| 130 | Mistral 7B v0.1 | mistral | 65.4 |
| 131 | Yi 1.5 9B | other | 65.3 |
| 132 | GLM-4V 9B | other | 65.1 |
| 133 | Llama 3 8B | meta | 65 |
| 134 | Baichuan2 13B Chat | other | 64.2 |
| 135 | OLMo 7B SFT | other | 64 |
| 136 | GPT-3.5 | openai | 63.9 |
| 137 | Claude 2.1 | anthropic | 63.8 |
| 138 | Open-Platypus | other | 63.7 |
| 139 | Qwen2 7B | alibaba | 63.3 |
| 140 | Claude Instant 1 | anthropic | 63.1 |
| 141 | OLMo 2 1124 7B | other | 63.1 |
| 142 | StarCoder2 15B | other | 62.8 |
| 143 | Hermes 3 Llama 3.1 8B | other | 62.6 |
| 144 | Llama 2 70B | meta | 62.5 |
| 145 | Argilla Notus 7B v1 | other | 62.4 |
| 146 | Llama 2 13B | meta | 62.1 |
| 147 | StarCoder2 3B | other | 62.1 |
| 148 | Qwen2.5 7B | alibaba | 61.9 |
| 149 | DeepSeek Coder 33B | deepseek | 61.6 |
| 150 | Flan-T5 XXL | other | 60.5 |
| 151 | Mistral 7B v0.3 | mistral | 60.5 |
| 152 | OpenChat 3.5 1210 | other | 60.4 |
| 153 | Claude 2 | anthropic | 60.2 |
| 154 | Ministral 8B | mistral | 60 |
| 155 | StarCoder2 7B | other | 60 |
| 156 | MPT 7B | other | 59.9 |
| 157 | DeepSeek Coder 7B | deepseek | 59.7 |
| 158 | Code Llama 34B | meta | 59.1 |
| 159 | StableCode 3B | other | 58.4 |
| 160 | Codestral | mistral | 58.3 |
| 161 | Grok Vision Beta | xai | 56.8 |
| 162 | Nous Capybara 34B | other | 56.1 |
| 163 | Phi-3 Mini | other | 55.9 |
| 164 | ChatGLM3 6B | other | 55.7 |
| 165 | Command Light | cohere | 55.5 |
| 166 | Yi 6B | other | 55.3 |
| 167 | StableLM 3 4B | other | 54.9 |
| 168 | Llama 2 7B | meta | 54.7 |
| 169 | Chat Bison | 54.6 | |
| 170 | Baichuan2 7B Chat | other | 53 |
| 171 | Gemma 2B | 53 | |
| 172 | Jurassic-2 Ultra | other | 52.9 |
| 173 | PaLM 2 | 52.9 | |
| 174 | Flan-T5 XL | other | 52.5 |
| 175 | TigerBot 70B Chat | other | 51.9 |
| 176 | Jurassic-2 Mid | other | 51.8 |
| 177 | Qwen2 1.5B | alibaba | 51.5 |
| 178 | Qwen2.5 1.5B | alibaba | 51.5 |
| 179 | Mistral Tiny | mistral | 51.3 |
| 180 | Qwen2.5 0.5B | alibaba | 51 |
| 181 | Zephyr 7B Alpha | other | 50.7 |
| 182 | Falcon 40B | other | 50.6 |
| 183 | GPT-3.5 Turbo 16K | openai | 50.2 |
| 184 | Grok Beta | xai | 50.2 |
| 185 | Ministral 3B | mistral | 50.1 |
| 186 | Zephyr 7B Beta | other | 50 |
| 187 | Nous Capybara 7B | other | 49.6 |
| 188 | Phi-1.5 | other | 49.5 |
| 189 | Flan-UL2 | other | 49.2 |
| 190 | MPT 30B | other | 49.2 |
| 191 | Qwen2.5 3B | alibaba | 48.8 |
| 192 | Yi 34B | other | 48.8 |
| 193 | Capybara 1.5B | other | 48.3 |
| 194 | GPT-3.5 Turbo | openai | 48.2 |
| 195 | Text Bison | 48.2 | |
| 196 | Phi-2 | other | 47.8 |
| 197 | Llama Guard 2 8B | meta | 47.6 |
| 198 | Phi-3.5 Mini | other | 47.5 |
| 199 | Llama 3.2 3B | meta | 46 |
| 200 | Llama Guard 3 8B | meta | 44.1 |
| 201 | StableLM Zephyr 3B | other | 43.9 |
| 202 | StableLM 2 1.6B | other | 40.7 |
| 203 | Phi-3 Vision | other | 40.1 |
| 204 | Llama 3.2 1B | meta | 38.8 |
| 205 | Embed English v3 | cohere | 0 |
BBH (BIG-Bench Hard)
描述
BIG-Bench 的 23 个最难任务子集,覆盖逻辑推理、符号操作、常识推理等领域,传统模型表现不佳,用于评估模型的复杂推理能力。
核心规格
| 分类 | 许可证 | 最后更新 |
|---|---|---|
| reasoning | MIT | 2022-10-01 |
基准
| 单位 |
|---|
| % |