Benchmarks
Which models perform best
Our ranked picks per capability, resolved to the newest live model in each family, alongside the public leaderboards you can check yourself.
Live from OpenRouter · updated

Methodology
We do not publish our own benchmark numbers. Each category lists AI HQ's editorial ranking, informed by the public benchmarks linked below and our own hands-on use, with live pricing attached so you can weigh capability against cost. Rankings are reviewed when major models ship. Treat them as a starting point, not a verdict.
Coding
Real-world software tasks: multi-file edits, bug fixing and agentic tool use.
- 1Anthropic: Claude Opus 5
Anthropic · 1.0M context
$5.00 in
$25.00 out
- 2OpenAI: GPT-5.6 Luna Pro
OpenAI · 1.1M context
$0.200 in
$1.20 out
- 3Google: Gemini 3.1 Pro Preview
Google · 1.0M context
$2.00 in
$12.00 out
- 4DeepSeek: DeepSeek V4.1 Flash
DeepSeek · 1.0M context
$0.150 in
$0.600 out
Benchmarks we watch
Reasoning
Maths, logic and multi-step problem solving.
- 1OpenAI: GPT-5.6 Luna Pro
OpenAI · 1.1M context
$0.200 in
$1.20 out
- 2Anthropic: Claude Opus 5
Anthropic · 1.0M context
$5.00 in
$25.00 out
- 3Google: Gemini 3.1 Pro Preview
Google · 1.0M context
$2.00 in
$12.00 out
- 4SpaceXAI: Grok 4.6
xAI · 500K context
$2.00 in
$6.00 out
Benchmarks we watch
Writing
Tone, structure and instruction following for long-form text.
- 1Anthropic: Claude Sonnet 5
Anthropic · 1.0M context
$2.00 in
$10.00 out
- 2OpenAI: GPT-5.6 Luna Pro
OpenAI · 1.1M context
$0.200 in
$1.20 out
- 3Google: Gemini 3.1 Pro Preview
Google · 1.0M context
$2.00 in
$12.00 out
Benchmarks we watch
Vision
Understanding images, charts, documents and video.
- 1Google: Gemini 3.1 Pro Preview
Google · 1.0M context
$2.00 in
$12.00 out
- 2OpenAI: GPT-5.6 Luna Pro
OpenAI · 1.1M context
$0.200 in
$1.20 out
- 3Anthropic: Claude Sonnet 5
Anthropic · 1.0M context
$2.00 in
$10.00 out
Benchmarks we watch
Agent Workflows
Long-horizon tool use, planning and recovering from errors.
- 1Anthropic: Claude Opus 5
Anthropic · 1.0M context
$5.00 in
$25.00 out
- 2OpenAI: GPT-5.6 Luna Pro
OpenAI · 1.1M context
$0.200 in
$1.20 out
- 3Google: Gemini 3.1 Pro Preview
Google · 1.0M context
$2.00 in
$12.00 out
Benchmarks we watch
Research
Long-context comprehension and grounded, cited answers.
- 1Google: Gemini 3.1 Pro Preview
Google · 1.0M context
$2.00 in
$12.00 out
- 2Anthropic: Claude Sonnet 5
Anthropic · 1.0M context
$2.00 in
$10.00 out
- 3Perplexity: Sonar Pro Search
Perplexity · 200K context
$3.00 in
$15.00 out
- 4OpenAI: GPT-5.6 Luna Pro
OpenAI · 1.1M context
$0.200 in
$1.20 out
Speed & Value
Fast, inexpensive models for high-volume workloads. Our shortlist of capable small models, ordered by live blended price (cheapest first).
- 1DeepSeek: DeepSeek V4.1 Flash
DeepSeek · 1.0M context
$0.150 in
$0.600 out
- 2Mistral: Mistral Small 4
Mistral · 262K context
$0.150 in
$0.600 out
- 3OpenAI: GPT-5 Mini
OpenAI · 400K context
$0.250 in
$2.00 out
- 4Google: Gemini 3.8 Flash
Google · 1.0M context
$0.750 in
$3.75 out
Benchmarks we watch
Newsletter
Stay Ahead of AI
Get the most important developments in artificial intelligence delivered directly to your inbox.
No hype. No spam.
Just trusted insights, major model releases, pricing updates, benchmark changes, product reviews, and practical guidance from across the AI ecosystem.
- Weekly AI Briefing
- Major Model Releases
- Pricing & Benchmark Updates
- Unsubscribe Anytime
The AI HQ Briefing
One email a week. Read in five minutes.