LLM Basics
Understand how large language models work, from tokens and context to embeddings and the limits worth designing around.
This series explains what a large language model does when you send it a prompt: tokens, attention, context windows, embeddings, training and the limits worth designing around.
It assumes you know what training and inference mean. Each topic cites the research papers and documentation it draws on.
What You’ll Learn
A large language model is a neural network, almost always a transformer, trained on very large amounts of text to predict the next token. GPT-3 showed that scaling this simple objective to 175 billion parameters produced models that could perform new tasks from a few examples in the prompt.
An LLM has read an enormous amount of text and learned what tends to come next. Useful abilities, such as summarising or answering questions, emerge from doing that at scale.
Assistants such as ChatGPT and Claude start as pre-trained language models and are then instruction-tuned so they follow requests helpfully.