All Comparisons
Browse all model comparisons on AI Benchmark Hub.
All model comparisons on AI Benchmark Hub.
-
GPT-4o vs Claude 3.5 Sonnet: Benchmark Comparison
See Editor's Take section for recommendation.
-
GPT-4o vs Gemini 1.5 Pro: Benchmark Comparison
See Editor's Take section for recommendation.
-
Claude 3.5 Sonnet vs Llama 3.1 405B: Benchmark Comparison
See Editor's Take section for recommendation.
-
DeepSeek V3 vs Mistral Large 2: Benchmark Comparison
See Editor's Take section.
-
DeepSeek V3 vs Mixtral 8x22B: Benchmark Comparison
See Editor's Take section.
-
o1 vs DeepSeek V3: Benchmark Comparison
See Editor's Take section.
-
o1 vs Llama 3.1 405B: Benchmark Comparison
See Editor's Take section.
-
o1 vs o1 mini: Benchmark Comparison
See Editor's Take section.
-
Phi-4 vs Qwen2.5 14B: Benchmark Comparison
See Editor's Take section.
-
Gemini 2.0 Flash vs Gemini 2.0 Flash Thinking: Benchmark Comparison
See Editor's Take section.
-
Llama 3.3 70B vs Llama 3.1 70B: Benchmark Comparison
See Editor's Take section.
-
Claude 3.5 Haiku vs Gemini 1.5 Flash: Benchmark Comparison
See Editor's Take section.
-
Claude 3.5 Haiku vs Llama 3.1 8B: Benchmark Comparison
See Editor's Take section.
-
NVIDIA Llama 3.1 Nemotron 70B vs Llama 3.1 70B: Benchmark Comparison
See Editor's Take section.
-
Llama 3.2 11B Vision vs Gemma 2 9B: Benchmark Comparison
See Editor's Take section.
-
Llama 3.2 90B Vision vs GPT-4o: Benchmark Comparison
See Editor's Take section.
-
Qwen2.5 32B vs Mistral Small 3: Benchmark Comparison
See Editor's Take section.
-
Qwen2.5 32B vs Qwen2.5 14B: Benchmark Comparison
See Editor's Take section.
-
Qwen2.5 72B vs DeepSeek V3: Benchmark Comparison
See Editor's Take section.
-
Qwen2.5 72B vs Mistral Large 2: Benchmark Comparison
See Editor's Take section.
-
Qwen2.5 72B vs Mixtral 8x22B: Benchmark Comparison
See Editor's Take section.
-
o1 Preview vs DeepSeek V3: Benchmark Comparison
See Editor's Take section.
-
Grok-2 vs Claude 3 Opus: Benchmark Comparison
See Editor's Take section.
-
Grok-2 vs DeepSeek V3: Benchmark Comparison
See Editor's Take section.
-
Grok-2 vs Llama 3.1 405B: Benchmark Comparison
See Editor's Take section.
-
GLM-4 Plus vs DeepSeek V3: Benchmark Comparison
See Editor's Take section.
-
GLM-4 Plus vs Qwen2.5 72B: Benchmark Comparison
See Editor's Take section.
-
Hermes 3 Llama 3.1 405B vs Llama 3.1 405B: Benchmark Comparison
See Editor's Take section.
-
Jamba 1.5 Large vs Mistral Large 2: Benchmark Comparison
See Editor's Take section.
-
Mistral Large 2 vs Mistral Large: Benchmark Comparison
See Editor's Take section.
-
Llama 3.1 405B vs Command R+: Benchmark Comparison
See Editor's Take section.
-
Llama 3.1 405B vs DeepSeek V3: Benchmark Comparison
See Editor's Take section.
-
Llama 3.1 405B vs Mistral Large 2: Benchmark Comparison
See Editor's Take section.
-
Llama 3.1 405B vs Mixtral 8x22B: Benchmark Comparison
See Editor's Take section.
-
Llama 3.1 405B vs Qwen2.5 72B: Benchmark Comparison
See Editor's Take section.
-
Llama 3.1 70B vs DeepSeek V2: Benchmark Comparison
See Editor's Take section.
-
Llama 3.1 70B vs Mixtral 8x7B: Benchmark Comparison
See Editor's Take section.
-
Llama 3.1 70B vs Qwen2.5 72B: Benchmark Comparison
See Editor's Take section.
-
Llama 3.1 8B vs Mistral 7B v0.3: Benchmark Comparison
See Editor's Take section.
-
Llama 3.1 8B vs Qwen2.5 7B: Benchmark Comparison
See Editor's Take section.
-
GPT-4o mini vs Claude 3.5 Haiku: Benchmark Comparison
See Editor's Take section.
-
GPT-4o mini vs Gemini 1.5 Flash: Benchmark Comparison
See Editor's Take section.
-
GPT-4o mini vs Llama 3.1 8B: Benchmark Comparison
See Editor's Take section.
-
GPT-4o mini vs Mistral 7B v0.3: Benchmark Comparison
See Editor's Take section.
-
GPT-4o mini vs Qwen2.5 7B: Benchmark Comparison
See Editor's Take section.
-
Gemma 2 27B vs Gemma 2 9B: Benchmark Comparison
See Editor's Take section.
-
Gemma 2 9B vs Gemma 7B: Benchmark Comparison
See Editor's Take section.
-
DeepSeek Coder V2 vs Codestral: Benchmark Comparison
See Editor's Take section.
-
DeepSeek Coder V2 vs DeepSeek V3: Benchmark Comparison
See Editor's Take section.
-
Claude 3.5 Sonnet vs Claude 3 Opus: Benchmark Comparison
See Editor's Take section.
-
Claude 3.5 Sonnet vs DeepSeek V3: Benchmark Comparison
See Editor's Take section.
-
Claude 3.5 Sonnet vs Gemini 1.5 Pro: Benchmark Comparison
See Editor's Take section.
-
Claude 3.5 Sonnet vs Grok-2: Benchmark Comparison
See Editor's Take section.
-
Claude 3.5 Sonnet vs Qwen2.5 72B: Benchmark Comparison
See Editor's Take section.
-
Qwen2 72B vs Qwen2.5 72B: Benchmark Comparison
See Editor's Take section.
-
GLM-4 9B Chat vs Qwen2.5 7B: Benchmark Comparison
See Editor's Take section.
-
Phi-3 Medium vs Phi-3 Small: Benchmark Comparison
See Editor's Take section.
-
Phi-3 Vision vs GPT-4o: Benchmark Comparison
See Editor's Take section.
-
Gemini 1.5 Flash vs Llama 3.1 8B: Benchmark Comparison
See Editor's Take section.
-
GPT-4o vs Claude 3 Opus: Benchmark Comparison
See Editor's Take section.
-
GPT-4o vs Claude 3.5 Haiku: Benchmark Comparison
See Editor's Take section.
-
GPT-4o vs DeepSeek V3: Benchmark Comparison
See Editor's Take section.
-
GPT-4o vs Gemini 2.0 Flash: Benchmark Comparison
See Editor's Take section.
-
GPT-4o vs GPT-4 Turbo: Benchmark Comparison
See Editor's Take section.
-
GPT-4o vs GPT-4o (2024-08-06): Benchmark Comparison
See Editor's Take section.
-
GPT-4o vs GPT-4o mini: Benchmark Comparison
See Editor's Take section.
-
GPT-4o vs Grok-2: Benchmark Comparison
See Editor's Take section.
-
GPT-4o vs Llama 3.1 405B: Benchmark Comparison
See Editor's Take section.
-
GPT-4o vs Mistral Large 2: Benchmark Comparison
See Editor's Take section.
-
GPT-4o vs o1 Preview: Benchmark Comparison
See Editor's Take section.
-
GPT-4o vs o1: Benchmark Comparison
See Editor's Take section.
-
GPT-4o vs Qwen2.5 72B: Benchmark Comparison
See Editor's Take section.
-
Yi Large vs Qwen2.5 72B: Benchmark Comparison
See Editor's Take section.
-
Yi Large vs Yi 34B: Benchmark Comparison
See Editor's Take section.
-
DeepSeek V2 vs DeepSeek V3: Benchmark Comparison
See Editor's Take section.
-
Llama 3 70B vs Llama 3 8B: Benchmark Comparison
See Editor's Take section.
-
Mixtral 8x22B vs Mistral Large 2: Benchmark Comparison
See Editor's Take section.
-
DBRX Instruct vs Mixtral 8x22B: Benchmark Comparison
See Editor's Take section.
-
Command R vs Command R+: Benchmark Comparison
See Editor's Take section.
-
Claude 3 Haiku vs Claude 3 Sonnet: Benchmark Comparison
See Editor's Take section.
-
Claude 3 Opus vs Claude 3 Sonnet: Benchmark Comparison
See Editor's Take section.
-
StarCoder2 15B vs Code Llama 34B: Benchmark Comparison
See Editor's Take section.
-
Mistral Large vs Mistral Medium: Benchmark Comparison
See Editor's Take section.
-
Qwen1.5 72B vs Qwen2 72B: Benchmark Comparison
See Editor's Take section.
-
Gemini 1.5 Pro vs Claude 3 Opus: Benchmark Comparison
See Editor's Take section.
-
Gemini 1.5 Pro vs DeepSeek V3: Benchmark Comparison
See Editor's Take section.
-
Gemini 1.5 Pro vs Gemini 1.5 Flash: Benchmark Comparison
See Editor's Take section.
-
Gemini 1.5 Pro vs Gemini 2.0 Flash: Benchmark Comparison
See Editor's Take section.
-
Gemini 1.5 Pro vs Llama 3.1 405B: Benchmark Comparison
See Editor's Take section.
-
Gemini 1.5 Pro vs Qwen2.5 72B: Benchmark Comparison
See Editor's Take section.
-
Code Llama 70B vs Codestral: Benchmark Comparison
See Editor's Take section.
-
DeepSeek Coder 33B vs Code Llama 34B: Benchmark Comparison
See Editor's Take section.
-
Mixtral 8x7B vs Mixtral 8x22B: Benchmark Comparison
See Editor's Take section.
-
Gemini 1.0 Pro vs Gemini 1.0 Ultra: Benchmark Comparison
See Editor's Take section.
-
Claude 2.1 vs Claude 2: Benchmark Comparison
See Editor's Take section.
-
Yi 34B vs Yi 6B: Benchmark Comparison
See Editor's Take section.
-
Falcon 180B vs Llama 2 70B: Benchmark Comparison
See Editor's Take section.
-
Llama 2 70B vs Llama 3 70B: Benchmark Comparison
See Editor's Take section.
-
GPT-4 vs GPT-4 Turbo: Benchmark Comparison
See Editor's Take section.
-
GPT-3.5 Turbo vs GPT-4: Benchmark Comparison
See Editor's Take section.