Foundation Model definition
A foundation model is a large AI model trained on broad, diverse data at scale, usually with self-supervised learning, that can be adapted to a wide range of downstream tasks. Large language models such as GPT, Claude, Gemini and Llama are foundation models, as are image models and multimodal models that handle text, images, audio and video together.
What makes a model a foundation model
Stanford researchers popularized the term in 2021 to describe models that serve as a common base, or foundation, for many applications. Two traits define them: they are trained on enormous, broad datasets, such as large portions of the public web, books and code, and the capabilities they learn transfer to tasks they were never explicitly trained for, from summarizing contracts to writing software.
Training is self-supervised: the model predicts missing or next pieces of data, such as the next word, so no human labeling is needed for the bulk of training. Post-training with instruction data and human feedback, such as RLHF, then turns the raw model into a helpful assistant that follows instructions and declines harmful requests.
Examples of foundation models
Foundation models now exist for most kinds of data, and new versions arrive frequently, so teams should treat any specific model choice as replaceable rather than permanent. Well-known families, grouped by the kind of data they handle, include:
- Language and multimodal: OpenAI GPT, Anthropic Claude and Google Gemini, plus open-weight families such as DeepSeek, Qwen, Gemma, gpt-oss, Mistral and Llama
- Image generation: OpenAI's GPT Image models, Google Imagen, FLUX, Stable Diffusion and Midjourney
- Speech: transcription models such as Whisper and NVIDIA's Parakeet and Canary, plus a range of text-to-speech models
- Embeddings: models that turn text or images into vectors for search, from OpenAI, Cohere, Google and open-source projects
- Domain models: foundation models trained for code, biology, weather or time series forecasting
How businesses adapt foundation models
Few companies train foundation models; most build on them. The usual ladder starts with prompt engineering, adds retrieval of company documents with RAG so answers are grounded in current data, and moves to fine-tuning or LoRA adapters when a task needs consistent behavior that prompting cannot reach. Tools and agents then let the model take actions in business systems.
Each step adds cost and complexity, so the right level depends on the task. A support assistant might need only prompting and retrieval, while a model that extracts structured data from specialist documents may justify fine-tuning a smaller open model for both accuracy and cost.
Open vs closed models and key risks
Closed models, accessed through APIs, usually lead on capability and are simple to use, but data leaves your environment and the provider controls pricing and model changes. Open-weight models can be self-hosted, customized and run in your own region, at the cost of operating the infrastructure. Many organizations use both, choosing per use case.
Licensing differs too. Some open-weight models allow broad commercial use, while others restrict use above certain user numbers or for specific purposes, so a review of the license belongs in model selection alongside benchmark scores, cost and hosting requirements.
Risks apply to both: hallucinations, bias inherited from training data, prompt injection, copyright questions about training data and outputs, and dependence on a single vendor. Nexzem builds model-agnostic applications during generative AI development, so clients can switch foundation models as better or cheaper options appear.