Back
LLM Basics

A Guide to Tokens and Context Windows in LLMs

Tokens and context windows decide how much text an AI model can process at once. This guide explains both ideas in plain language, shows why models sometimes forget earlier parts of a conversation, and offers practical ways to work within these limits.

On this page
What a token isWhat the context window isWhy models seem to forgetWorking within the limits

Every large language model reads text as tokens and can only consider a fixed number of them at once. Those two facts explain most of the practical behaviour people notice: why long chats drift, why pasting a whole document sometimes fails, and why pricing is quoted per million tokens.

01

What a token is

A token is a chunk of text produced by the model’s tokenizer. Most modern tokenizers use byte-pair encoding, which learns common character sequences from training data. Frequent words become a single token; rarer words are split into several pieces.

In English, one token averages roughly four characters, or about three-quarters of a word. Code, numbers and non-English languages usually need more tokens for the same amount of meaning, which is worth remembering when estimating cost.

02

What the context window is

The context window is the maximum number of tokens the model can process in a single request. It covers everything: the system prompt, conversation history, any documents or retrieved passages, and the response the model writes.

Windows have grown from a few thousand tokens in early models to hundreds of thousands in current ones. A larger window lets you include more material, but it does not guarantee the model uses all of it equally well.

03

Why models seem to forget

When a conversation exceeds the window, applications must drop or compress older turns. Anything outside the window is simply not seen by the model on that request.

Even inside the window, research such as the “Lost in the Middle” study found that models tend to use information at the start and end of long inputs more reliably than information buried in the middle. Placement matters.

04

Working within the limits

Retrieve rather than paste: send only the passages relevant to the question instead of whole documents. Summarise older conversation turns so key facts survive in fewer tokens. Put the most important instructions and evidence near the start or end of the prompt.

Budget explicitly. Decide how many tokens go to instructions, history, retrieved context and the answer, and enforce those limits in code rather than hoping requests fit.

Key takeaways
Tokens are sub-word units; English averages about four characters per token.
The context window is a shared budget for instructions, history, context and output.
Long inputs are not used uniformly, so position the most important content deliberately.
Retrieval, summarisation and explicit token budgets keep applications reliable as conversations grow.
Next article · AI Agents
Difference between AI Agents vs Agentic AI →