Key Specifications

Vendormeta
Version3.2-11b-vision
Release Date2024-09-25
Context Window128000 tokens
Input Modalitiestext, image
Output Modalitiestext
LicenseLlama 3.2 Community License
Documentationhttps://llama.meta.com/docs/

Benchmark Performance

Benchmark Score Unit Evaluated At Notes Source
MMLU 61.9 % 2024-09-25 5-shot view
HUMANEVAL 50.1 pass@1 2024-09-25 view
GSM8K 68.2 % 2024-09-25 0-shot CoT view
MATH 29.9 % 2024-09-25 0-shot CoT view
BBH 69.1 % 2024-09-25 3-shot CoT view
GPQA 35.5 % 2024-09-25 0-shot view
IFEVAL 58.4 % 2024-09-25 prompt_strict view
ARC 86 % 2024-09-25 challenge view
MUSR 41.7 % 2024-09-25 0-shot view
WINOGRANDE 72.7 % 2024-09-25 0-shot view

Pricing

Tier Price Currency
Input$0.55 / MtokUSD
Output$0.55 / MtokUSD
Cache Read$0 / MtokUSD
Cache Write$0 / MtokUSD

Source: https://ai.meta.com/blog/ · as of 2024-09-25

Compliance

  • Data Residency: self-host
  • SOC2: ✗
  • HIPAA: ✗
  • GDPR: ✗
  • ISO 27001: ✗

Llama 3.2 11B Vision

Model Overview

Meta Llama 3.2 11B Vision 经济型多模态开源模型, 128K 上下文, 11B 参数, 适合边缘视觉任务。

Core Specifications

Vendor Version Release Date Context Window Input Modalities Output Modalities License
Meta 3.2-11b-vision 2024-09-25 128K text, image text Llama 3.2 Community License

Benchmark Performance

Benchmark Score Unit Notes
MMLU (Massive Multitask Language Understanding) 61.9 % 5-shot
HumanEval 50.1 pass@1
GSM8K (Grade School Math 8K) 68.2 % 0-shot CoT
MATH 29.9 % 0-shot CoT
BBH (BIG-Bench Hard) 69.1 % 3-shot CoT
GPQA 35.5 % 0-shot
IFEval 58.4 % prompt_strict
ARC 86.0 % challenge
MUSR 41.7 % 0-shot
WinoGrande 72.7 % 0-shot

Pricing

Input Output Cache Read Cache Write

per million tokens

Strengths

  • Input Modalities: text, image, audio.

Weaknesses

  • Proprietary, not self-hostable.

Use Cases

  • Vision and image understanding

References