Key Specifications

Vendorother
Versionnous-capybara-34b
Release Date2023-11-15
Context Window4096 tokens
Input Modalitiestext
Output Modalitiestext
LicenseApache 2.0
Documentationhttps://huggingface.co/models

Benchmark Performance

Benchmark Score Unit Evaluated At Notes Source
MMLU 53.2 % 2023-11-15 5-shot view
HUMANEVAL 50.1 pass@1 2023-11-15 view
GSM8K 31.9 % 2023-11-15 0-shot CoT view
MATH 23.1 % 2023-11-15 0-shot CoT view
BBH 56.1 % 2023-11-15 3-shot CoT view
GPQA 29.7 % 2023-11-15 0-shot view
IFEVAL 52.9 % 2023-11-15 prompt_strict view
ARC 80.7 % 2023-11-15 challenge view
MUSR 31.7 % 2023-11-15 0-shot view
WINOGRANDE 67.3 % 2023-11-15 0-shot view

Pricing

Tier Price Currency
Input$0.4 / MtokUSD
Output$0.4 / MtokUSD
Cache Read$0 / MtokUSD
Cache Write$0 / MtokUSD

Source: https://huggingface.co/models · as of 2023-11-15

Compliance

  • Data Residency: self-host
  • SOC2: ✗
  • HIPAA: ✗
  • GDPR: ✗
  • ISO 27001: ✗

Nous Capybara 34B

Model Overview

NousResearch Capybara 34B 对话模型, 4K 上下文, 基于 Yi 34B 微调, 改进长对话与推理能力。

Core Specifications

Vendor Version Release Date Context Window Input Modalities Output Modalities License
Other nous-capybara-34b 2023-11-15 4K text text Apache 2.0

Benchmark Performance

Benchmark Score Unit Notes
MMLU (Massive Multitask Language Understanding) 53.2 % 5-shot
HumanEval 50.1 pass@1
GSM8K (Grade School Math 8K) 31.9 % 0-shot CoT
MATH 23.1 % 0-shot CoT
BBH (BIG-Bench Hard) 56.1 % 3-shot CoT
GPQA 29.7 % 0-shot
IFEval 52.9 % prompt_strict
ARC 80.7 % challenge
MUSR 31.7 % 0-shot
WinoGrande 67.3 % 0-shot

Pricing

Input Output Cache Read Cache Write

per million tokens

Strengths

  • Reliable general-purpose model.

Weaknesses

  • MMLU 53.2, weak knowledge reasoning.
  • Proprietary, not self-hostable.
  • Context window 4K is limited.

Use Cases

  • General chat and Q&A

References