Skip to main content
AIHQ

Benchmarks

Which models perform best

Our ranked picks per capability, resolved to the newest live model in each family, alongside the public leaderboards you can check yourself.

Live from OpenRouter · updated

A laptop displays several coloured measurement curves in a workspace.
Measurement in progress. The curves on screen are not AI HQ benchmark data.Illustrative photographPhoto: ThisIsEngineering / Pexels

Methodology

We do not publish our own benchmark numbers. Each category lists AI HQ's editorial ranking, informed by the public benchmarks linked below and our own hands-on use, with live pricing attached so you can weigh capability against cost. Rankings are reviewed when major models ship. Treat them as a starting point, not a verdict.

Coding

Real-world software tasks: multi-file edits, bug fixing and agentic tool use.

  1. 1
    Anthropic: Claude Opus 5

    Anthropic · 1.0M context

    $5.00 in

    $25.00 out

  2. 2
    OpenAI: GPT-5.6 Luna Pro

    OpenAI · 1.1M context

    $0.200 in

    $1.20 out

  3. 3
    Google: Gemini 3.1 Pro Preview

    Google · 1.0M context

    $2.00 in

    $12.00 out

  4. 4
    DeepSeek: DeepSeek V4.1 Flash

    DeepSeek · 1.0M context

    $0.150 in

    $0.600 out

Reasoning

Maths, logic and multi-step problem solving.

  1. 1
    OpenAI: GPT-5.6 Luna Pro

    OpenAI · 1.1M context

    $0.200 in

    $1.20 out

  2. 2
    Anthropic: Claude Opus 5

    Anthropic · 1.0M context

    $5.00 in

    $25.00 out

  3. 3
    Google: Gemini 3.1 Pro Preview

    Google · 1.0M context

    $2.00 in

    $12.00 out

  4. 4
    SpaceXAI: Grok 4.6

    xAI · 500K context

    $2.00 in

    $6.00 out

Writing

Tone, structure and instruction following for long-form text.

  1. 1
    Anthropic: Claude Sonnet 5

    Anthropic · 1.0M context

    $2.00 in

    $10.00 out

  2. 2
    OpenAI: GPT-5.6 Luna Pro

    OpenAI · 1.1M context

    $0.200 in

    $1.20 out

  3. 3
    Google: Gemini 3.1 Pro Preview

    Google · 1.0M context

    $2.00 in

    $12.00 out

Benchmarks we watch

Vision

Understanding images, charts, documents and video.

  1. 1
    Google: Gemini 3.1 Pro Preview

    Google · 1.0M context

    $2.00 in

    $12.00 out

  2. 2
    OpenAI: GPT-5.6 Luna Pro

    OpenAI · 1.1M context

    $0.200 in

    $1.20 out

  3. 3
    Anthropic: Claude Sonnet 5

    Anthropic · 1.0M context

    $2.00 in

    $10.00 out

Benchmarks we watch

Agent Workflows

Long-horizon tool use, planning and recovering from errors.

  1. 1
    Anthropic: Claude Opus 5

    Anthropic · 1.0M context

    $5.00 in

    $25.00 out

  2. 2
    OpenAI: GPT-5.6 Luna Pro

    OpenAI · 1.1M context

    $0.200 in

    $1.20 out

  3. 3
    Google: Gemini 3.1 Pro Preview

    Google · 1.0M context

    $2.00 in

    $12.00 out

Benchmarks we watch

Research

Long-context comprehension and grounded, cited answers.

  1. 1
    Google: Gemini 3.1 Pro Preview

    Google · 1.0M context

    $2.00 in

    $12.00 out

  2. 2
    Anthropic: Claude Sonnet 5

    Anthropic · 1.0M context

    $2.00 in

    $10.00 out

  3. 3
    Perplexity: Sonar Pro Search

    Perplexity · 200K context

    $3.00 in

    $15.00 out

  4. 4
    OpenAI: GPT-5.6 Luna Pro

    OpenAI · 1.1M context

    $0.200 in

    $1.20 out

Benchmarks we watch

Speed & Value

Fast, inexpensive models for high-volume workloads. Our shortlist of capable small models, ordered by live blended price (cheapest first).

  1. 1
    DeepSeek: DeepSeek V4.1 Flash

    DeepSeek · 1.0M context

    $0.150 in

    $0.600 out

  2. 2
    Mistral: Mistral Small 4

    Mistral · 262K context

    $0.150 in

    $0.600 out

  3. 3
    OpenAI: GPT-5 Mini

    OpenAI · 400K context

    $0.250 in

    $2.00 out

  4. 4
    Google: Gemini 3.8 Flash

    Google · 1.0M context

    $0.750 in

    $3.75 out

Benchmarks we watch

Newsletter

Stay Ahead of AI

Get the most important developments in artificial intelligence delivered directly to your inbox.

No hype. No spam.

Just trusted insights, major model releases, pricing updates, benchmark changes, product reviews, and practical guidance from across the AI ecosystem.

  • Weekly AI Briefing
  • Major Model Releases
  • Pricing & Benchmark Updates
  • Unsubscribe Anytime

The AI HQ Briefing

One email a week. Read in five minutes.

By subscribing you agree to receive the AI HQ newsletter. Your address is processed by our email delivery provider and never sold. Unsubscribe anytime.