New batches starting this week Β· Limited seats

What Is RAG? Retrieval-Augmented Generation Explained

RAG lets a large language model answer from your own documents, fetched at query time, with citations. This guide explains how RAG works step by step, the architecture choices that matter in production, how to evaluate it and when not to use it.

RAG pipeline: your documents are chunked and embedded, searched in a vector index, reranked, and passed to an LLM that answers with citations
Last updated Β· 14 min read Β· 3,180 words

RAG, or retrieval-augmented generation, is a way of building AI applications in which a large language model (LLM) answers from documents fetched at the moment you ask. Instead of relying only on what the model memorised during training, a RAG system searches your own content, puts the most relevant passages into the prompt and asks the model to answer from that evidence, ideally with citations. It is the most common architecture behind enterprise chat assistants, internal knowledge bots and document Q&A tools.

What is RAG? A plain definition

Retrieval-augmented generation combines two components. A retriever searches a knowledge source (documents, wiki pages, tickets, database rows) and returns the passages most relevant to a question. A generator, the LLM, writes an answer using those passages as context. The model is not retrained; the knowledge lives outside it, in an index you control and can update at any time.

Think of it as an open-book exam: the model looks up the right pages first, then writes its answer and says where it found it. Answer quality depends heavily on whether the right pages were found.

Why LLMs need retrieval

LLMs are excellent at language but unreliable as a database. Three limitations make retrieval necessary for most business use cases.

  • Knowledge cutoff. A model only knows what was in its training data up to a certain date. Last month's policy update, this quarter's price list or yesterday's incident report simply is not in there.
  • Private data. Your company's contracts, runbooks, HR policies and customer records were never in any public training set, and should not be. Retrieval lets the model use them at query time without baking them into the model.
  • Citations and trust. When a model answers from memory, it cannot reliably tell you where a fact came from, and it may produce a fluent answer that is simply wrong (a hallucination). When it answers from retrieved passages, you can show the source, and users and auditors can check it.

Updating a RAG index is also cheap and immediate: re-ingest a changed document and the next answer reflects it, whereas retraining is slower, costlier and harder to audit.

How RAG works, step by step

A RAG system has two phases. The indexing phase runs ahead of time (and again whenever content changes). The query phase runs every time a user asks a question.

INDEXING (offline, on content change)
 sources -> ingest -> chunk -> embed -> index
                                         |
QUERY (per question)                     v
 question -> retrieve (+filters) <- vector/keyword index
          -> rerank -> build prompt -> LLM generate
          -> answer + citations -> user
  1. Ingest. Pull content from its sources: SharePoint, Confluence, file shares, PDFs, ticketing systems, databases. Extract clean text, keep structure such as headings and tables, and capture metadata: title, owner, version, effective date, department and access permissions.
  2. Chunk. Split each document into passages small enough to retrieve precisely but large enough to make sense alone. A whole 80-page policy is too big to be useful context; a single sentence loses meaning.
  3. Embed. Convert each chunk into an embedding, a list of numbers that represents its meaning, using an embedding model. Passages with similar meaning end up close together in this vector space, even if they use different words.
  4. Index. Store the embeddings with the chunk text and metadata in a vector store, often alongside a keyword index for exact-term search.
  5. Retrieve. At query time, embed the user's question with the same model and find the nearest chunks, applying metadata filters (department, region, document status, and the user's permissions).
  6. Rerank. Pass the top candidates through a reranker, a model that scores each question-passage pair more carefully, and keep only the best few.
  7. Prompt. Build the prompt: system instructions ("answer only from the context; say you don't know if the context does not contain the answer"), the retrieved chunks with their IDs, and the user's question.
  8. Generate. The LLM writes the answer grounded in those chunks.
  9. Cite. Return the answer with links to the exact source passages. Validate in code that every cited ID was actually retrieved, so the model cannot invent a source.

The LLM is one step of nine; most of the engineering, and most failures, sit in the other eight.

RAG architecture: the design choices that matter

A tutorial RAG app takes an afternoon. A RAG system a bank or hospital will trust needs deliberate decisions at every stage.

Chunking strategy

Chunking affects answer quality more than most people expect. Common strategies:

  • Fixed-size with overlap. Split every N tokens with some overlap. Simple and a reasonable baseline, but it can cut a rule in half.
  • Structure-aware. Split on headings, sections, list items or table boundaries, so each chunk is a coherent unit such as "Section 4.2 Maternity Leave". Usually better for policies, manuals and contracts.
  • Parent-child (small-to-big). Retrieve on small, precise chunks but send the larger parent section to the model so it has surrounding context.

Whatever you choose, prepend context to each chunk (document title, section heading) so a passage like "this applies after six months" still makes sense on its own.

Embedding models

Embedding models come from the major cloud AI platforms (Amazon Bedrock, Azure OpenAI, Google Gemini and Vertex AI) and as self-hosted open-source models. Choose on retrieval quality on your content and languages (test it; a leaderboard is not enough), data residency, cost and latency. Queries and documents must use the same model, and changing it means re-embedding the corpus.

Vector stores

The vector store holds embeddings and finds nearest neighbours quickly:

  • pgvector, an extension that adds vector search to PostgreSQL. A strong default when you already run Postgres, because chunks, metadata, permissions and vectors live in one transactional database you already know how to back up and secure.
  • FAISS, an open-source library for fast similarity search. Excellent for experiments and embedded use, but it is a library, not a database: persistence, filtering and access control are your job.
  • OpenSearch, a search engine with vector support alongside mature keyword search, which makes hybrid search natural. Dedicated vector databases and managed cloud options also exist.

Vector search is good at meaning ("time off after having a baby" finds "maternity leave"). It is weaker on exact terms: policy numbers, error codes, product SKUs, people's names. Keyword search (BM25) is the opposite. Hybrid search runs both and merges the results, commonly with reciprocal rank fusion. In enterprise content full of codes and acronyms, hybrid search is usually worth it.

Reranking

First-stage retrieval is optimised for speed, so it casts a wide net. A reranker (typically a cross-encoder) reads each question-passage pair together and re-orders by true relevance. Reranking a few dozen candidates down to a handful often improves answers noticeably and sends less context to the LLM, cutting cost and latency.

Metadata filters

Metadata turns a pile of chunks into a governed knowledge base. Filter on document status (only "current", never "draft" or "superseded"), effective date, region, product line and language. Many "the AI gave a wrong answer" complaints are really "the AI retrieved last year's version".

Access control

This is the design choice that separates a demo from a production system. If the index is built with one admin account and no permissions, every user can retrieve every chunk, including salary bands and disciplinary files. Capture source-system permissions at ingestion, store them as metadata on each chunk, resolve the user's identity and groups at query time (for example from Microsoft Entra ID) and filter before retrieval results ever reach the prompt. Never rely on the prompt to "hide" restricted content.

If you want to build these pieces hands-on (chunking, pgvector, hybrid retrieval, reranking and evaluation) on real cloud AI platforms, Cloudsoft's AI, GenAI and Agentic AI course covers RAG end to end, in the Ameerpet classroom or online.

Worked example: an HR-policy assistant

Consider an illustrative case: the Hyderabad global capability centre (GCC) of a multinational insurer. HR answers the same questions every day about leave, relocation, notice periods and hybrid work, and leadership wants an HR-policy assistant inside the company chat tool.

Indexing. HR agrees one authoritative policy library; drafts and archived versions are excluded. Each policy is chunked by section heading, and each chunk carries metadata: policy name, section, effective date, applicable country and location, employee grade, and the access groups that can read it. Chunks are stored in PostgreSQL with pgvector plus a keyword index for policy codes. A scheduled job re-ingests changed documents and marks superseded versions inactive rather than deleting them, so old answers stay auditable.

A query. An employee in Hyderabad asks: "How many days of paternity leave do I get, and can I split them?"

  1. The assistant resolves the user's identity: India, a particular grade, standard employee group.
  2. Hybrid retrieval searches for the question, filtered to current policies applicable to India and to chunks this user's groups may read.
  3. A reranker places the India "Paternity leave: entitlement" and "Paternity leave: how to apply" sections on top.
  4. The prompt instructs the model to answer only from the supplied sections, to state the policy and effective date, and to say "I can't find this in the HR policies" when the context is silent.
  5. The answer states the entitlement and whether splitting is allowed exactly as the policy says, with two citations linking to the sections. The code checks both citation IDs were in the retrieved set.

Edge cases. Ask about a manager's bonus and the permission filter returns nothing usable, so the assistant declines instead of leaking a manager-only document. Ask about a policy that does not exist and it says so. Update the leave policy on Monday and Tuesday's answers reflect it.

How to evaluate a RAG system

"It looked right when I tried it" is not evaluation. RAG can fail in retrieval (wrong passages), in generation (right passages, wrong answer) or both, so measure them separately. Four metrics, used by frameworks such as Ragas, form the common vocabulary:

MetricWhat it asksWhat a low score tells you
FaithfulnessIs every claim in the answer supported by the retrieved context?The model is adding facts that are not in the sources (hallucinating).
Answer relevanceDoes the answer actually address the question asked?Answers are evasive, generic or off-topic, even if accurate.
Context precisionOf the passages retrieved, how many are relevant, and are the relevant ones ranked near the top?Retrieval is noisy; the model is wading through irrelevant context.
Context recallDid retrieval find all the information needed to answer?Key passages are missing: chunking, embeddings, filters or search need work.

Build a test set from real user questions, with reference answers and sources checked by domain experts (HR, in the example above), and include adversarial and permission test cases. Run it on every change to prompts, chunking, models or retrieval settings, and trace production requests with a tool such as LangSmith or Langfuse so you can see what was retrieved for a bad answer. For the full picture, including LLM-as-judge, human review and release thresholds, read our guide to LLM evaluation.

Common RAG failure modes

  • Stale or duplicated documents produce confident answers from outdated versions.
  • Bad chunking splits a rule from its exception, so the answer is half right.
  • Missing exact-match search means codes and names are never found.
  • No access control lets restricted content leak into answers.
  • Too much context buries the relevant passage and raises cost and latency.
  • Prompt injection through retrieved documents containing hidden instructions.
  • No evaluation, so nobody knows whether a change made things better or worse.

Each of these, with symptoms, root causes, fixes and an HR-assistant scenario that hits several at once, is covered in our catalogue of enterprise AI failure modes.

RAG vs fine-tuning, briefly

RAG and fine-tuning solve different problems. RAG changes what the model can see: it supplies current, private, citable facts at query time. Fine-tuning changes how the model behaves: its style, output format, domain vocabulary or performance on a narrow, repeated task. If the problem is "the model does not know our policies", start with RAG. If it is "the model knows enough but will not answer in our required format or tone", fine-tuning (or better prompting) may help. Many mature systems combine both. The detailed comparison, with a decision guide, is in RAG vs fine-tuning.

When not to use RAG

RAG is the wrong tool when:

  • The answer lives in structured data. "What were Hyderabad branch sales last quarter?" belongs in a SQL query or an API call, not a similarity search over text. Let the model call a tool (via function calling or the Model Context Protocol) that queries the system of record.
  • The knowledge is small and stable. If everything the model needs fits comfortably in the prompt and rarely changes, put it in the prompt and skip the retrieval infrastructure.
  • The task is about behaviour, not knowledge. Classification, extraction into a fixed schema or matching a house style are usually prompting or fine-tuning problems.
  • The user needs an action, not an answer. "Raise a laptop replacement ticket" requires an agent with tools, possibly using RAG as one of them. See what agentic AI is.

RAG tools and frameworks

LayerCommon optionsNotes
LLMs and embeddingsAmazon Bedrock, Azure OpenAI, Google Gemini / Vertex AI, open-source modelsChoose by quality on your data, data residency, cost and the cloud you already run
OrchestrationLangChain, LlamaIndex, LangGraphLangGraph suits multi-step or agentic RAG with explicit control flow
Vector storePostgreSQL + pgvector, FAISS, OpenSearch, dedicated vector databasespgvector is a pragmatic default if you already run Postgres
RerankingCross-encoder rerankers, hosted rerank APIsSmall change, often a large quality gain
EvaluationRagas, LLM-as-judge, expert-reviewed test setsRun in CI on every change
ObservabilityLangSmith, Langfuse, OpenTelemetryTrace retrieved chunks, prompts, tokens and latency per request
API and servingPython + FastAPI, Docker, KubernetesA RAG system is still a service that needs auth, scaling and monitoring

From RAG demo to enterprise outcome

The gap between a working notebook and a system employees rely on is mostly integration, identity, evaluation and operations. Taking RAG into production inside customer environments is a large part of what Forward Deployed Engineers do, and Cloudsoft's FDE PRO program builds an Enterprise Knowledge Assistant as one of its projects. If you are preparing for interviews, work through our RAG interview questions, or start with the RAG interview questions for freshers.

Frequently asked questions

What is RAG in simple terms?

RAG, or retrieval-augmented generation, is a technique where an AI system first searches your documents for passages relevant to a question and then asks a large language model to answer using those passages. It lets the model use current, private information and cite its sources without being retrained.

How does RAG work?

Documents are ingested, split into chunks, converted into embeddings and stored in an index. When a user asks a question, the system retrieves the most relevant chunks, often reranks them, places them in the prompt with instructions and asks the LLM to generate an answer grounded in that context, with citations to the source passages.

Does RAG stop hallucinations?

RAG reduces hallucinations but does not eliminate them. The model can still misread or go beyond the retrieved context, and if retrieval returns the wrong passages the answer will be wrong. Instructing the model to answer only from context, validating citations in code and measuring faithfulness with an evaluation set are what keep hallucinations under control.

What is a vector database in RAG?

A vector database stores embeddings, numeric representations of the meaning of text chunks, and finds the chunks closest in meaning to a question. Options include PostgreSQL with the pgvector extension, OpenSearch, the FAISS library and dedicated vector databases. Many enterprise teams combine vector search with keyword search for better results.

Is RAG better than fine-tuning?

Neither is better in general; they solve different problems. RAG gives a model access to current, private and citable knowledge at query time. Fine-tuning changes the model's behaviour, such as style, format or performance on a narrow task. For questions about your own documents, RAG is usually the right starting point, and some systems use both.

What is hybrid search in RAG?

Hybrid search runs semantic vector search and keyword search such as BM25 on the same question and merges the results, often with reciprocal rank fusion. Vector search finds passages with similar meaning, while keyword search reliably finds exact terms like policy numbers, error codes and names, so the combination usually retrieves better context for enterprise content.

Ready to go from understanding RAG to building it? Cloudsoft's AI, GenAI and Agentic AI training in Hyderabad takes you through LLMs, RAG pipelines, evaluation and agents with hands-on labs, in our Ameerpet classroom beside Ameerpet Metro or live online. Call +91 96660 19191 to book a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us