RAG definition
Retrieval-augmented generation (RAG) is a technique that improves a large language model's answers by first retrieving relevant information from an external source, such as company documents or a database, and adding it to the prompt. The model then answers using that retrieved context, which makes responses more accurate, current and traceable to sources.
How does RAG work?
A RAG system has two pipelines. The ingestion pipeline prepares knowledge ahead of time, and the query pipeline uses it each time a user asks something. The model itself is unchanged: RAG works by putting the right information in front of the model at the right moment, which is why it can answer questions about documents written after the model was trained.
- Ingest: parse documents, split them into chunks and attach metadata such as source and permissions.
- Embed and index: convert chunks to vectors and store them in a vector database, often alongside a keyword index.
- Retrieve: embed the user's question and fetch the most relevant chunks with hybrid search.
- Rerank: reorder candidates with a reranking model and keep the best few.
- Generate: send the question and chunks to the LLM, instructing it to cite sources and admit when the answer is missing.
RAG vs fine-tuning
RAG supplies knowledge; fine-tuning shapes behavior. If the model needs to know your policies, product catalog or case history, use RAG: updates take effect as soon as a document is re-indexed, answers can cite sources, and per-user permissions can be enforced at retrieval time. If the model needs a consistent format, tone or specialized skill, fine-tuning helps. Many production systems use both, with retrieval for facts and a tuned model for style.
Why RAG systems fail
Most poor RAG answers are retrieval problems, not model problems. If the right passage never reaches the prompt, no model can answer correctly. Debugging therefore starts by inspecting what was retrieved for failing questions before any prompt is changed or a bigger model is tried.
- Bad parsing: tables, scanned PDFs and slides converted into garbled text.
- Poor chunking: answers split across chunks or separated from their headings.
- Vocabulary mismatch: users and documents use different terms, fixed with hybrid keyword and vector search.
- Stale indexes: documents updated at the source but not re-ingested.
- Missing reranking: loosely related chunks crowd out the passage that actually answers.
- Permission leaks: retrieval ignoring who is allowed to see which document.
Example: an internal policy assistant
An HR team indexes its handbook, leave policies and regional benefit documents, each tagged by country. When an employee in Germany asks how much parental leave they get, the system retrieves only German and global policy chunks, and the model answers with a link to the exact section. Questions the documents do not cover receive a clear "I could not find this" and a route to an HR contact.
How to evaluate a RAG system
Evaluate retrieval and generation separately. For retrieval, measure whether the correct passage appears in the top results for a set of known questions. For generation, measure faithfulness to the retrieved text, answer relevance and citation accuracy, using human review plus tools such as Ragas or TruLens. Nexzem builds this test set with the client before launch and reruns it after every change to chunking, models or prompts.