Key Specifications

Vendorxai
Version2
Release Date2024-08-13
Context Window131072 tokens
Input Modalitiestext, image
Output Modalitiestext
LicenseProprietary
Documentationhttps://docs.x.ai/

Benchmark Performance

Benchmark Score Unit Evaluated At Notes Source
MMLU 87.5 % 2024-08-13 5-shot view
HUMANEVAL 88.4 pass@1 2024-08-13 view
GSM8K 93.2 % 2024-08-13 0-shot CoT view
MATH 76.8 % 2024-08-13 0-shot CoT view
BBH 84 % 2024-08-13 3-shot CoT view

Pricing

Tier Price Currency
Input$2 / MtokUSD
Output$10 / MtokUSD
Cache Read$0 / MtokUSD
Cache Write$0 / MtokUSD

Source: https://x.ai/api · as of 2024-08-13

Compliance

  • Data Residency: US
  • SOC2: ✗
  • HIPAA: ✗
  • GDPR: ✗
  • ISO 27001: ✗

Grok-2

Model Overview

Grok-2 is xAI’s flagship model, released on August 13, 2024. The model is notable for its tight integration with the X (formerly Twitter) platform, providing real-time awareness of current events and trending topics — a capability no other flagship model matches. Grok-2 accepts both text and image inputs, supports a 131K context window, and is positioned as a competitive alternative to GPT-4o and Claude 3.5 Sonnet on academic benchmarks. On MMLU, Grok-2 scores 87.5, within striking distance of GPT-4o (88.7) and Claude 3.5 Sonnet (88.7). The model is available via the X platform’s Grok chatbot (for Premium subscribers) and the xAI API. Grok-2 is also notable for its relatively permissive content policy compared to other flagship models, which xAI positions as a feature for users seeking fewer refusals.

Key Specifications

Attribute Value
Vendor xAI
Version 2
Release Date 2024-08-13
Context Window 131,072 tokens
Input Modalities text, image
Output Modalities text
License Proprietary
Documentation https://docs.x.ai/

Benchmark Performance

Benchmark Score Unit Notes
MMLU 87.5 % 5-shot
HumanEval 88.4 pass@1
GSM8K 93.2 % 0-shot CoT
MATH 76.8 % 0-shot CoT
BBH 84.0 % 3-shot CoT

Pricing

Tier Price (per 1M tokens)
Input $2.00
Output $10.00
Cache Read $0.00
Cache Write $0.00

Pricing source: https://x.ai/api (as of 2024-08-13).

Strengths

  • Real-time awareness via X platform integration, unique among flagships.
  • Strong benchmark performance, competitive with GPT-4o on most metrics.
  • Multimodal input (text + image) support.
  • More permissive content policy for users seeking fewer refusals.

Weaknesses

  • Proprietary model with no self-host option and limited enterprise compliance certifications.
  • No audio modality, unlike GPT-4o.
  • Tightly coupled to the X ecosystem, which may be a concern for some enterprise buyers.
  • Limited track record vs established vendors like OpenAI and Anthropic.

Use Cases

  • Real-time news and event-aware chatbots.
  • Social media content generation and trend analysis.
  • Customer-facing applications where current events matter.
  • Image-based question answering in conjunction with text.

References