Transfer Learning definition
Transfer learning is a machine learning technique in which a model trained on one large task is reused as the starting point for a different but related task. Instead of training from scratch, teams adapt a pretrained model with a smaller dataset, which cuts training time, compute cost and the amount of labeled data needed.
How does transfer learning work?
A model pretrained on a huge, general dataset, such as ImageNet for images or large web text collections for language, has already learned broadly useful features. In a vision model, early layers detect edges and textures that matter for almost any image. Transfer learning keeps those learned layers, replaces the final output layer with one suited to the new task, and trains on a much smaller dataset of task-specific examples.
There are two basic approaches. Feature extraction freezes the pretrained layers and trains only the new head, which is fast and works well when the new data resembles the original. Fine-tuning unfreezes some or all layers and trains them with a low learning rate, which adapts better to a different domain but needs more data and care to avoid overfitting.
Transfer learning strategies
- Feature extraction: freeze the base model and train a small classifier on its outputs.
- Partial fine-tuning: retrain the top layers, keep lower layers fixed.
- Full fine-tuning: update every weight, usually with a small learning rate.
- Parameter-efficient fine-tuning: methods such as LoRA and adapters train a small set of added weights.
- Domain-adaptive pretraining: continue pretraining on unlabeled text or images from your domain, then fine-tune.
Examples of transfer learning
Worked example: a dermatology clinic wants to sort skin images into a few lesion categories but has only a few thousand labeled photos. Training a deep network from scratch on that data would overfit badly. Starting from a model pretrained on ImageNet and fine-tuning it on the clinic's images produces a far stronger classifier, which specialists then validate against their own diagnoses before any clinical use.
- Vision: ResNet, EfficientNet and vision transformers adapted for defect detection or medical imaging.
- Language: BERT-style models fine-tuned for ticket routing, sentiment or entity extraction.
- Speech: Whisper adapted to industry vocabulary or regional accents.
- Generative AI: open language models fine-tuned on company documents or support transcripts.
Transfer learning vs training from scratch
Fine-tuning is the most common form of transfer learning, so the terms are often used together. Training from scratch only makes sense when the domain is very different from anything available pretrained, such as unusual sensor signals, and when you have a large dataset and compute budget. For most business projects, a pretrained starting point is cheaper, faster and more accurate.
Transfer can also hurt. When the source and target domains differ too much, pretrained features may mislead the model, a problem called negative transfer. Comparing the fine-tuned model against a simple baseline trained only on the new data catches this early and shows whether the pretrained starting point is actually helping.
Risks to check before you reuse a model
Pretrained models carry the biases and blind spots of their training data, and their licenses vary: some weights forbid commercial use or impose conditions. Teams should check the license, review the model card and test on their own edge cases. Nexzem reviews model licenses and data provenance at the start of every transfer learning project, before any client data touches the model.