Low-Rank Adaptation definition
LoRA (Low-Rank Adaptation) is a parameter-efficient fine-tuning technique that adapts a large pretrained model by training small low-rank matrices added to some of its layers, while the original weights stay frozen. Because only a tiny fraction of parameters is trained, LoRA cuts memory and compute needs dramatically and produces small adapter files that can be swapped per task.
How LoRA works
Fine-tuning normally updates every weight in a model, billions of numbers for a modern LLM, which needs large GPUs and produces a full copy of the model per task. LoRA, introduced by Microsoft researchers in 2021, builds on the observation that the change needed to adapt a model tends to have low intrinsic rank. Instead of updating a large weight matrix directly, it learns two small matrices whose product approximates the update and adds that product at inference time.
With a rank of, say, 8 or 16, the trainable parameters can drop to a small fraction of a percent of the original model. The base model stays frozen and shared, and each task gets its own adapter file of a few megabytes to a few hundred megabytes, which can be loaded, swapped or merged into the base weights for deployment. Serving frameworks such as vLLM can even switch between many adapters per request on a single base model.
LoRA vs full fine-tuning
LoRA is the default starting point for most custom model work today because of the practical advantages below, although full fine-tuning can still edge ahead on very large or very different datasets, and it is worth testing both when the budget allows:
- Much lower GPU memory, often enough to train mid-sized models on a single GPU
- Faster training and cheaper experiments
- Small adapter files instead of full model copies, easy to version and store
- Many adapters can share one base model in production, one per customer or task
- Less risk of catastrophic forgetting, because the original weights are unchanged
QLoRA and key settings
QLoRA combines LoRA with 4-bit quantization of the frozen base model, so even large open models can be fine-tuned on a single high-memory GPU with little loss in quality. Libraries such as Hugging Face PEFT, Unsloth and Axolotl make LoRA and QLoRA training largely a matter of configuration rather than custom code.
The main settings are the rank (higher means more capacity and more parameters), alpha (a scaling factor), dropout and which layers receive adapters, commonly the attention projections and sometimes the feed-forward layers. Data quality matters far more than settings: a few thousand clean, representative examples usually beat a large, noisy set.
When to use LoRA
LoRA suits teams that need a model to follow a specific format, style or vocabulary consistently, such as producing structured outputs from medical notes, writing in a brand voice or classifying documents into a company's own taxonomy, where prompting alone has plateaued. It also lets smaller open models match larger hosted ones on narrow tasks, cutting inference cost and keeping data in your own environment.
It is not a good way to teach a model large amounts of changing knowledge; retrieval-augmented generation usually handles that better, as our RAG vs fine-tuning comparison explains. Nexzem trains LoRA adapters on open models when evaluation shows a clear gain over prompting alone.