Skip to content

What is Retrieval-Augmented Generation (RAG)?

Generative AI & LLMs, explained by the engineers who build it. Definition, how it works, use cases and common questions.

RAG definition

Retrieval-augmented generation (RAG) is a technique that improves a large language model's answers by first retrieving relevant information from an external source, such as company documents or a database, and adding it to the prompt. The model then answers using that retrieved context, which makes responses more accurate, current and traceable to sources.

How does RAG work?

A RAG system has two pipelines. The ingestion pipeline prepares knowledge ahead of time, and the query pipeline uses it each time a user asks something. The model itself is unchanged: RAG works by putting the right information in front of the model at the right moment, which is why it can answer questions about documents written after the model was trained.

  • Ingest: parse documents, split them into chunks and attach metadata such as source and permissions.
  • Embed and index: convert chunks to vectors and store them in a vector database, often alongside a keyword index.
  • Retrieve: embed the user's question and fetch the most relevant chunks with hybrid search.
  • Rerank: reorder candidates with a reranking model and keep the best few.
  • Generate: send the question and chunks to the LLM, instructing it to cite sources and admit when the answer is missing.

RAG vs fine-tuning

RAG supplies knowledge; fine-tuning shapes behavior. If the model needs to know your policies, product catalog or case history, use RAG: updates take effect as soon as a document is re-indexed, answers can cite sources, and per-user permissions can be enforced at retrieval time. If the model needs a consistent format, tone or specialized skill, fine-tuning helps. Many production systems use both, with retrieval for facts and a tuned model for style.

Why RAG systems fail

Most poor RAG answers are retrieval problems, not model problems. If the right passage never reaches the prompt, no model can answer correctly. Debugging therefore starts by inspecting what was retrieved for failing questions before any prompt is changed or a bigger model is tried.

  • Bad parsing: tables, scanned PDFs and slides converted into garbled text.
  • Poor chunking: answers split across chunks or separated from their headings.
  • Vocabulary mismatch: users and documents use different terms, fixed with hybrid keyword and vector search.
  • Stale indexes: documents updated at the source but not re-ingested.
  • Missing reranking: loosely related chunks crowd out the passage that actually answers.
  • Permission leaks: retrieval ignoring who is allowed to see which document.

Example: an internal policy assistant

An HR team indexes its handbook, leave policies and regional benefit documents, each tagged by country. When an employee in Germany asks how much parental leave they get, the system retrieves only German and global policy chunks, and the model answers with a link to the exact section. Questions the documents do not cover receive a clear "I could not find this" and a route to an HR contact.

How to evaluate a RAG system

Evaluate retrieval and generation separately. For retrieval, measure whether the correct passage appears in the top results for a set of known questions. For generation, measure faithfulness to the retrieved text, answer relevance and citation accuracy, using human review plus tools such as Ragas or TruLens. Nexzem builds this test set with the client before launch and reruns it after every change to chunking, models or prompts.

RAG: common questions

Something else on your mind? Ask a consultant and get a reply within one business day.

Does RAG eliminate hallucinations?

No, but it reduces them substantially when retrieval works well. The model can still misread or over-generalize from the retrieved text, or fill gaps when retrieval returns nothing useful. Instructions to answer only from sources, citation requirements and faithfulness checks lower the remaining risk.

Do I need a vector database for RAG?

Usually some form of vector search, but not necessarily a dedicated database. PostgreSQL with pgvector, Elasticsearch or OpenSearch can handle many workloads alongside keyword search. Dedicated options such as Pinecone, Weaviate, Qdrant or Milvus make sense at larger scale or when advanced filtering and performance features are needed.

Is RAG still needed with long context windows?

Often yes. Placing an entire document library in every prompt is slow and expensive, and models can overlook details buried in very long inputs. RAG also enforces access permissions and keeps answers traceable. Long context windows do let RAG systems pass larger, more complete chunks, which improves answer quality.

Keep exploring the generative ai & llms glossary

Need RAG in your product?

A solutions consultant replies within one business day with next steps, a rough estimate and a suggested team.