Benchmark Performance
| Benchmark |
Score |
Unit |
Evaluated At |
Notes |
Source |
| MMLU |
66.3 |
% |
2023-07-11 |
5-shot |
view |
| HUMANEVAL |
31.7 |
pass@1 |
2023-07-11 |
— |
view |
| GSM8K |
32.5 |
% |
2023-07-11 |
0-shot CoT |
view |
| MATH |
21.9 |
% |
2023-07-11 |
0-shot CoT |
view |
| BBH |
60.2 |
% |
2023-07-11 |
3-shot CoT |
view |
| GPQA |
31.4 |
% |
2023-07-11 |
0-shot |
view |
| IFEVAL |
55 |
% |
2023-07-11 |
prompt_strict |
view |
| ARC |
80.4 |
% |
2023-07-11 |
challenge |
view |
| MUSR |
35.2 |
% |
2023-07-11 |
0-shot |
view |
| WINOGRANDE |
66.4 |
% |
2023-07-11 |
0-shot |
view |
Pricing
| Tier |
Price |
Currency |
| Input | $11 / Mtok | USD |
| Output | $32 / Mtok | USD |
| Cache Read | $0 / Mtok | USD |
| Cache Write | $0 / Mtok | USD |
Source:
https://www.anthropic.com/pricing
· as of 2023-07-11
Compliance
- Data Residency: US
- SOC2: ✓
- HIPAA: ✗
- GDPR: ✓
- ISO 27001: ✗
Claude 2
Model Overview
Anthropic Claude 2, 100K 上下文, 在编码/数学/推理上超越 Claude 1.3, 首次开放 API。
Core Specifications
| Vendor |
Version |
Release Date |
Context Window |
Input Modalities |
Output Modalities |
License |
| Anthropic |
2 |
2023-07-11 |
100K |
text |
text |
Proprietary |
| Benchmark |
Score |
Unit |
Notes |
| MMLU (Massive Multitask Language Understanding) |
66.3 |
% |
5-shot |
| HumanEval |
31.7 |
pass@1 |
— |
| GSM8K (Grade School Math 8K) |
32.5 |
% |
0-shot CoT |
| MATH |
21.9 |
% |
0-shot CoT |
| BBH (BIG-Bench Hard) |
60.2 |
% |
3-shot CoT |
| GPQA |
31.4 |
% |
0-shot |
| IFEval |
55.0 |
% |
prompt_strict |
| ARC |
80.4 |
% |
challenge |
| MUSR |
35.2 |
% |
0-shot |
| WinoGrande |
66.4 |
% |
0-shot |
Pricing
| Input |
Output |
Cache Read |
Cache Write |
| — |
— |
— |
— |
per million tokens
Strengths
Weaknesses
- HumanEval 31.7, coding weak.
- Proprietary, not self-hostable.
Use Cases
- Long document summarization
References