Skip to content

What is Fine-Tuning?

Generative AI & LLMs, explained by the engineers who build it. Definition, how it works, use cases and common questions.

Fine-Tuning definition

Fine-tuning is the process of taking a pretrained machine learning model, such as a large language model, and training it further on a smaller, task-specific dataset so it performs better on that task. It adjusts the model's weights to teach a consistent format, tone, vocabulary or skill that prompting alone does not reliably achieve.

How does fine-tuning work?

You prepare a dataset of examples showing the input and the ideal output, often a few hundred to a few thousand for a focused task. Training then runs for a small number of passes with a low learning rate, nudging the pretrained weights toward your examples without erasing what the model already knows. A held-out set of examples is used to check that quality actually improved over the base model with a good prompt.

Full fine-tuning updates every weight, which needs significant GPU memory for large models. Parameter-efficient fine-tuning instead trains a small set of added weights. LoRA, the most common method, adds low-rank matrices to selected layers, and QLoRA does the same on a quantized model to fit on smaller GPUs. Hosted services from OpenAI, Google's Gemini Enterprise Agent Platform and Amazon Bedrock offer fine-tuning without managing hardware.

Types of fine-tuning

  • Supervised fine-tuning (SFT): training on input and ideal output pairs.
  • Instruction tuning: SFT on many varied tasks so a model follows instructions in general.
  • Preference tuning: RLHF or DPO, using ranked answers to steer style and safety.
  • Parameter-efficient methods: LoRA, QLoRA and adapters.
  • Continued pretraining: more training on raw domain text, such as legal or medical documents.
  • Distillation: training a smaller model to copy a larger model's outputs.
  • Domain adaptation: tuning speech or vision models on specialized recordings or images.

When to fine-tune, and when not to

Fine-tune when you need consistent behavior that prompts cannot hold: a strict output format, a specific brand voice, a specialized classification or extraction task, or the same quality from a smaller, cheaper and faster model. Fine-tuning also shortens prompts, since the instructions and examples no longer need to be sent with every request.

Do not fine-tune to add knowledge that changes, such as prices, policies or product documentation. The model may blend facts unpredictably, and every update would need retraining. Retrieval-augmented generation handles changing knowledge better and can cite sources. Always test strong prompting and retrieval first, because they are cheaper to build and maintain.

Example: a document extraction model

An insurer extracts a dozen fields from claim forms. A large model with a detailed prompt works but is slow and costly at volume. The team collects a few thousand human-verified extractions and fine-tunes a smaller open model with LoRA. The tuned model matches the large model on the evaluation set for this narrow task, runs on the insurer's own servers and answers faster. Low-confidence extractions still go to a human reviewer, and the team retrains when claim forms change.

Risks and costs

Fine-tuning quality depends on data quality: inconsistent or wrong examples teach inconsistent or wrong behavior. Over-training can cause the model to lose general abilities, known as catastrophic forgetting, or to memorize sensitive training data. Tuned models also need maintenance, since moving to a newer base model means retraining. Nexzem treats fine-tuning datasets as versioned assets and compares every tuned model against a prompted baseline before deployment.

Fine-Tuning: common questions

Something else on your mind? Ask a consultant and get a reply within one business day.

How much data do you need to fine-tune an LLM?

For a narrow task, a few hundred high-quality examples can produce clear improvement, and a few thousand is common. Quality and consistency matter more than volume. Broader behavior changes or new domains need far more. Start small, measure against a held-out set and add data where the model still fails.

What is LoRA fine-tuning?

LoRA, or low-rank adaptation, freezes the original model weights and trains small additional matrices inserted into selected layers. It needs far less memory and storage than full fine-tuning, and the resulting adapter is a small file that can be swapped in and out, so one base model can serve several tuned tasks.

Is fine-tuning better than RAG?

They solve different problems. RAG gives a model access to current, specific knowledge and citations. Fine-tuning changes how the model behaves: format, tone and task skill. For a question-answering assistant over company documents, RAG is usually the right starting point, sometimes combined with a fine-tuned model for style.

Keep exploring the generative ai & llms glossary

Need Fine-Tuning in your product?

A solutions consultant replies within one business day with next steps, a rough estimate and a suggested team.