Back
LLM Basics

What Is LLM Fine-Tuning? A Complete Guide

Take a model that already knows a great deal in general, and give it focused additional training on examples specific to your business so its default behaviour matches what you need.

On this page
How it worksFull versus parameter-efficientPreparing dataWhen to use it

Fine-tuning takes a model that already understands language in general and trains it further on examples specific to your task, so its default behaviour matches what you need without long instructions in every prompt.

01

How it works

A pre-trained model is trained for additional steps on a curated dataset of input–output pairs. The model’s weights shift toward producing outputs like those examples. Instruction-tuned assistants were created this way, using human-written demonstrations and feedback, as described in the InstructGPT work.

02

Full versus parameter-efficient

Full fine-tuning updates every weight and needs significant compute. Parameter-efficient methods such as LoRA train small low-rank adapter matrices, and QLoRA combines this with quantisation so large models can be tuned on a single GPU.

03

Preparing data

Quality matters more than quantity. A few hundred to a few thousand clean, consistent examples that reflect real inputs usually outperform large noisy datasets. Hold back a test set so improvements can be measured honestly.

04

When to use it

Fine-tune for consistent structure, tone, domain terminology or a narrow classification or extraction task. Do not fine-tune to add frequently changing facts; retrieval is the better tool for that.

Key takeaways
Fine-tuning adapts behaviour by continuing training on your examples.
LoRA and QLoRA make tuning affordable by training small adapters.
Clean, representative data beats large volumes of noisy data.
Use fine-tuning for behaviour, retrieval for changing knowledge.
Next article · Product
AI SaaS v AI Augmented App. Lets see the difference →