Skip to main content
AIHQ

Beginner

What is a Context Window?

Why a model 'forgets', what 128K or 1M means, and how it affects cost.

AI HQ Editorial3 min read
Two blue equipment racks filled with electronic modules and wiring in a laboratory.
Two blue equipment racks filled with electronic modules and wiring in a laboratory.Illustrative photographPhoto: Eric Stoynov / Unsplash

The model's working memory

The context window is the maximum amount of text (in tokens) a model can consider at once: your instructions, everything pasted in, and the conversation so far, plus its reply. A 128K window is roughly 100,000 words; a 1M window is about eight novels.

When a conversation exceeds the window, older content is dropped or summarised. That is the 'forgetting' people notice in long chats.

Bigger is not free

Everything in the window is billed as input tokens on every turn. Long conversations and large pasted documents are the fastest way to run up a bill. Context length is listed for every model on our Models page.

Learn

Keep reading

More →

Beginner

AI Pricing Explained

Tokens, input vs output, blended cost and subscriptions vs API. How an AI bill is actually calculated.

6 min read

Newsletter

Stay Ahead of AI

Get the most important developments in artificial intelligence delivered directly to your inbox.

No hype. No spam.

Just trusted insights, major model releases, pricing updates, benchmark changes, product reviews, and practical guidance from across the AI ecosystem.

  • Weekly AI Briefing
  • Major Model Releases
  • Pricing & Benchmark Updates
  • Unsubscribe Anytime

The AI HQ Briefing

One email a week. Read in five minutes.

By subscribing you agree to receive the AI HQ newsletter. Your address is processed by our email delivery provider and never sold. Unsubscribe anytime.