Key Specifications

Vendorother
Versionsonar-reasoning
Release Date2024-12-18
Context Window127072 tokens
Input Modalitiestext
Output Modalitiestext
LicenseProprietary
Documentationhttps://huggingface.co/models

Benchmark Performance

Benchmark Score Unit Evaluated At Notes Source
MMLU 87.7 % 2024-12-18 5-shot view
HUMANEVAL 81.1 pass@1 2024-12-18 view
GSM8K 89.8 % 2024-12-18 0-shot CoT view
MATH 50.7 % 2024-12-18 0-shot CoT view
BBH 83.4 % 2024-12-18 3-shot CoT view
GPQA 50.1 % 2024-12-18 0-shot view
IFEVAL 83.1 % 2024-12-18 prompt_strict view
ARC 96.2 % 2024-12-18 challenge view
MUSR 69.4 % 2024-12-18 0-shot view
WINOGRANDE 87.2 % 2024-12-18 0-shot view

Pricing

Tier Price Currency
Input$2 / MtokUSD
Output$8 / MtokUSD
Cache Read$0 / MtokUSD
Cache Write$0 / MtokUSD

Source: https://huggingface.co/models · as of 2024-12-18

Compliance

  • Data Residency: self-host
  • SOC2: ✗
  • HIPAA: ✗
  • GDPR: ✗
  • ISO 27001: ✗

Sonar Reasoning

Model Overview

Perplexity Sonar Reasoning 推理型在线 RAG 模型, 127K 上下文, 基于 DeepSeek R1 微调, 链式思维推理。

Core Specifications

Vendor Version Release Date Context Window Input Modalities Output Modalities License
Other sonar-reasoning 2024-12-18 127K text text Proprietary

Benchmark Performance

Benchmark Score Unit Notes
MMLU (Massive Multitask Language Understanding) 87.7 % 5-shot
HumanEval 81.1 pass@1
GSM8K (Grade School Math 8K) 89.8 % 0-shot CoT
MATH 50.7 % 0-shot CoT
BBH (BIG-Bench Hard) 83.4 % 3-shot CoT
GPQA 50.1 % 0-shot
IFEval 83.1 % prompt_strict
ARC 96.2 % challenge
MUSR 69.4 % 0-shot
WinoGrande 87.2 % 0-shot

Pricing

Input Output Cache Read Cache Write

per million tokens

Strengths

  • MMLU score 87.7, strong knowledge reasoning.
  • HumanEval 81.1, excellent code generation.
  • GSM8K 89.8, robust math reasoning.

Weaknesses

  • Proprietary, not self-hostable.

Use Cases

  • Code generation and debugging

References