What Is LLM Fine-Tuning? A Complete Guide
Take a model that already knows a great deal in general, and give it focused additional training on examples specific to your business so its default behaviour matches what you need.
Fine-tuning takes a model that already understands language in general and trains it further on examples specific to your task, so its default behaviour matches what you need without long instructions in every prompt.
How it works
A pre-trained model is trained for additional steps on a curated dataset of input–output pairs. The model’s weights shift toward producing outputs like those examples. Instruction-tuned assistants were created this way, using human-written demonstrations and feedback, as described in the InstructGPT work.
Full versus parameter-efficient
Full fine-tuning updates every weight and needs significant compute. Parameter-efficient methods such as LoRA train small low-rank adapter matrices, and QLoRA combines this with quantisation so large models can be tuned on a single GPU.
Preparing data
Quality matters more than quantity. A few hundred to a few thousand clean, consistent examples that reflect real inputs usually outperform large noisy datasets. Hold back a test set so improvements can be measured honestly.
When to use it
Fine-tune for consistent structure, tone, domain terminology or a narrow classification or extraction task. Do not fine-tune to add frequently changing facts; retrieval is the better tool for that.