New batches starting this week Β· Limited seats

Agentic RAG Explained: When Retrieval Needs an Agent

Agentic RAG lets an LLM agent decide whether, where and how to retrieve, so it can answer multi-part, multi-source and ambiguous questions. This guide covers how it works, what it costs and when a simpler pipeline is the better choice.

Agentic RAG flow: question, query planning and rewriting, routing across documents, SQL and APIs, an evidence check, grounded answer
Last updated Β· 14 min read Β· 3,152 words

Agentic RAG is retrieval-augmented generation where an LLM agent controls the retrieval process instead of a fixed pipeline. Rather than embedding the question once, fetching the top-k chunks and answering, an agentic RAG system decides whether to retrieve at all, which source to query, how to rewrite or split the question, and whether what came back is good enough to answer from. It handles multi-part, multi-source and ambiguous questions far better than single-shot RAG, but it is slower, costs more per answer and is harder to evaluate, so use it only for the questions that need it.

New to the basic pipeline? Start with what RAG is and how it works.

Where single-shot RAG runs out

Classic RAG is one pass: question in, one retrieval call, one generation call, answer out. Better chunking, hybrid search and reranking make that call more accurate, but cannot help questions that need more than one. Three kinds of question cause most of the trouble.

Multi-part questions

"Is water damage from a burst pipe covered under my home policy, and what is my deductible for it?" Coverage sits in the policy wording; the deductible sits in the customer's policy schedule. One embedding of the whole sentence often retrieves good passages for one half and nothing for the other, so the model answers half and guesses the rest.

Questions that need several sources

A support question may need the product manual (text), the customer's contract tier (a database row) and an open ticket's status (an API). Single-shot RAG has one retriever wired to one store, and copying amounts, dates and statuses into a vector index is usually wrong: they belong in systems that can be queried exactly, and embedded copies go stale.

Ambiguous or underspecified questions

"What's the limit?" could mean a credit limit, a claim limit or an API rate limit. A fixed pipeline retrieves whatever is most similar and answers confidently; it has no step where it could ask which one, or use context to narrow it down.

A fourth, quieter failure: retrieval that silently returns poor results. A fixed pipeline cannot notice irrelevant top chunks, so the LLM either says "I don't know" or fills gaps from its training data.

What makes RAG "agentic"

The word "agentic" here has a specific meaning: the model, not the developer's hard-coded pipeline, decides the next retrieval step. Usually that is an LLM with tool calling in a loop, where each retriever and data source is a tool. The underlying patterns (ReAct, routing, plan-and-execute, reflection) are covered in agentic AI design patterns; here is how they apply to retrieval.

1. Deciding whether to retrieve

Not every turn needs retrieval. "Can you make that shorter?" needs none. Skipping retrieval saves latency and keeps irrelevant context out of the prompt. In regulated domains you will usually force retrieval for any factual claim, so this decision is constrained rather than free.

2. Query rewriting and decomposition

Rewriting (resolving "it" from history, expanding acronyms, keeping identifiers intact) is covered in the hybrid search article. What agentic RAG adds is decomposition, or query planning: the agent breaks a compound question into sub-questions, decides their order, and notices dependencies. "Which of my claims from last year were rejected, and what clause was cited each time?" has to fetch the claims first, because the second half depends on the result.

Plans can be made up front (plan-and-execute, easier to inspect and test) or step by step (ReAct-style, better at recovering when an early result is surprising).

3. Routing across indexes and tools

The agent picks a source per sub-question:

  • Vector or hybrid index for policy wording, manuals, FAQs and other unstructured text.
  • SQL for exact facts and aggregates: amounts, dates, counts, statuses. Use a narrow, parameterised tool or guarded text-to-SQL (see the text-to-SQL agent project for the safety discipline).
  • APIs for live state: claim status, ticket status, inventory, account details.
  • A knowledge graph when the question is about relationships across entities, as in GraphRAG.

Routing quality depends on tool descriptions. "search_documents" tells the model nothing; "search_policy_wording: published policy documents and endorsements; no customer-specific data" tells it when to use the tool and when not to. Many routing bugs are documentation bugs.

4. Iterative retrieval with self-checks

After each retrieval, the system checks what it got before moving on. Research has named several variants, such as corrective RAG and self-reflective RAG; the shared ideas are simple:

  • Relevance grading. A grader (a small LLM call, a classifier or a reranker score threshold) judges whether the retrieved passages actually address the sub-question. If not, the agent rewrites the query or tries another source instead of answering from weak evidence. This is the core of the corrective pattern.
  • Sufficiency check. Is there evidence for every part of the question? If not, retrieve again for the missing part only.
  • Groundedness check. After drafting, verify each claim against retrieved evidence; unsupported claims are removed, flagged or trigger another retrieval. This is the core of the self-reflective pattern.
  • Clarification. When ambiguity cannot be resolved from context, the right action is to ask the user a short question instead of guessing.

5. Stopping criteria

A loop needs a clear exit. Stop on whichever comes first:

  • every sub-question has graded, relevant evidence and the draft passes the groundedness check;
  • a hard cap on retrieval rounds or tool calls is reached;
  • a token or time budget is exhausted;
  • the agent is about to repeat a query with no new information;
  • the agent concludes the answer is not in any available source, and says so plainly.

"I couldn't find this in your policy documents; here is who to contact" is a correct answer, and running out of budget should produce an honest partial answer, never a confident guess.

The agentic RAG loop at a glance

User question
     |
     v
[Plan] retrieve? rewrite? split into sub-questions
     |
     v
[Route] pick tool per sub-question
   |-- vector / hybrid index (documents)
   |-- SQL tool (exact facts)
   |-- API tool (live status)
     |
     v
[Grade] relevant? sufficient? ---- no ---+
     | yes                               |
     v                     rewrite / reroute
[Draft answer]             (within budget)
     |                                   |
     v                                   |
[Check] grounded? -------- no -----------+
     | yes
     v
Answer with citations, or "not found" / ask user

In code, this maps cleanly onto a state graph: nodes for plan, retrieve, grade, draft and check, with conditional edges and a counter in state for the loop limit. Frameworks with explicit control flow, such as LangGraph, are a common choice because such graphs are easier to test than a free-running agent.

Agentic RAG vs RAG: comparison table

AspectClassic RAGAdvanced RAG (hybrid + rerank)Agentic RAGGraphRAG
Control flowFixed, one passFixed, one pass with more stagesModel-driven loopFixed or agent-driven over a graph
Retrieval calls per questionOneOne logical call (several retrievers fused)Variable, severalOne or more graph and vector queries
SourcesUsually one vector indexOne corpus, keyword + vectorMany: indexes, SQL, APIs, graphsKnowledge graph, often plus vectors
Best atSingle-passage factual questionsExact terms, IDs, precise top resultsMulti-part, multi-source, ambiguous questionsMulti-hop relationships, corpus-wide themes
Self-correctionNoneNoneGrades evidence, retries, verifiesLimited unless combined with an agent
Latency and costLowestLow to moderateHighest and variableHigh build cost, moderate query cost
EvaluationRetrieval + answer metricsPer-stage retrieval + answer metricsTrajectory + answer metricsRetrieval, graph quality + answer metrics

These compose rather than compete: an agentic system often calls a hybrid-plus-rerank retriever as one tool and a graph query as another. Agentic RAG changes who decides which retrieval to run; the others change how well a single retrieval works.

Illustrative example: an insurer's policy assistant

Consider a general insurer building an assistant for contact-centre agents in its Hyderabad operations centre. Policy wordings and endorsements are PDFs, each customer's policy schedule (sum insured, deductibles, add-ons) sits in the policy administration database, and claim status comes from the claims API.

An agent on the phone types: "Customer says her car was damaged in waterlogging last week. Is engine damage covered, and where is her claim?"

Single-shot RAG retrieves general flood-damage paragraphs and says flood damage is covered. It misses that engine damage from water ingress is often excluded unless an add-on is held, cannot tell whether this customer holds it, and cannot see the claim.

The agentic version:

  1. Plan. Three sub-questions: what the wording says about water damage to the engine; whether this policy includes the add-on; claim status.
  2. Route. Wording to the hybrid index, filtered to the motor product and policy version date; add-on to a read-only SQL tool keyed by policy number; claim to the claims API.
  3. Grade. The first search returns general flood text. The grader rejects it, so the agent rewrites the query with likely exclusion terms ("consequential loss", "water ingress", "engine protection") and finds the exclusion and the add-on endorsement.
  4. Combine. The schedule shows no add-on; the claim is registered and awaiting a surveyor.
  5. Check and answer. Every statement maps to a clause, schedule field or API result. The answer cites them and notes that the final decision rests with the claims team.

Note two choices: the assistant informs the human agent and never decides coverage, and the wording search is filtered to the version in force when the policy was issued. Agentic retrieval depends on good metadata even more, because the agent builds the filters.

Want to build systems like this end to end? Cloudsoft's AI, GenAI and Agentic AI course covers RAG, LangGraph agents, tool calling and evaluation with hands-on labs.

Permissions across multiple sources

Once an agent reaches several sources, access control becomes the central design problem. The rule: the agent must never retrieve anything the end user could not see directly.

  • Propagate the user's identity to every tool. Each call carries the authenticated user or a user-scoped token, and the source enforces its own permissions: row-level security in SQL, document ACLs as index filters, scoped API tokens. Never give the agent one all-seeing service account.
  • Enforce filters in code, not in the prompt. The model chooses the query; the tool implementation injects tenant, region and role filters it cannot remove.
  • Scope tools narrowly. "get_policy_schedule(policy_number)" that checks the caller is assigned to that customer is safer than a general SQL tool. Read-only by default; anything that writes needs explicit approval.
  • Watch for leakage through combination. Two permitted sources, joined, can reveal something sensitive. Mask personal data before it enters the context.
  • Treat retrieved text as untrusted. Documents can contain planted instructions. Retrieved content is data, never commands, and permissions must hold even if the model is fooled.
  • Log every tool call with the user, arguments and result size so you can audit what the agent saw.

For delegated access and scoped tokens in depth, see identity and access for AI agents.

The costs: latency, tokens, evaluation and loops

Agentic RAG buys answer quality with resources.

  • Latency. Each planning, grading and checking step is an extra LLM call in series. Run independent sub-questions in parallel, use a small fast model for grading and routing, stream the final answer and cache frequent retrievals; see LLM latency optimization.
  • Token cost. Context accumulates as each call re-sends the conversation, plan and earlier results. Summarise or drop graded evidence and pass only what the next step needs.
  • Variable behaviour. The same question can take different paths on different runs, which makes latency, debugging and capacity planning harder.
  • Loops. Without hard limits, an agent can keep rewriting a query for an answer that does not exist. Cap rounds, detect repeated queries, and treat "not found" as a valid exit.
  • Harder evaluation. A correct answer reached by calling the wrong tool, or by reading data it should not have, is still a failure. You have to judge the path as well as the result.

Evaluating agentic RAG: trajectory plus answer metrics

Standard RAG evaluation metrics still score the final output; trajectory metrics score the path.

Answer metrics (per question):

  • Correctness against a reference answer reviewed by a domain expert.
  • Faithfulness or groundedness: every claim supported by retrieved evidence.
  • Completeness: all parts of a multi-part question answered.
  • Appropriate refusal: "not found" or a clarifying question when that is the right response.

Trajectory metrics (per run):

  • Routing accuracy: did each sub-question go to the right source?
  • Decomposition quality: were the sub-questions complete and non-redundant?
  • Retrieval quality per step: context precision and recall for each retrieval call, not only the final context.
  • Grader accuracy: how often the relevance grader agrees with human judgement. False rejections cause needless loops; false acceptances defeat the point.
  • Efficiency: number of steps, tool calls, tokens and time to answer.
  • Policy compliance: zero calls outside the user's permissions, zero writes without approval, no hits on the loop limit for answerable questions.

Build the test set from real questions tagged by type (single-source, multi-part, ambiguous, unanswerable), plus adversarial cases. Capture traces with LangSmith, Langfuse or OpenTelemetry so each step can be graded. Also run the test set through your simpler pipeline: if agentic RAG does not clearly beat it, you do not need it. The full methodology for agent test suites, graders and CI gates is in AI agent evaluation.

When not to use agentic RAG

  • Most questions are single-source lookups. Improve chunking and add hybrid search and reranking first; it is cheaper, faster and easier to evaluate.
  • Latency budgets are tight. Voice assistants and in-app search often cannot afford several sequential LLM calls.
  • The steps are always the same. If every question needs "look up customer, then search policy", write that as a fixed workflow. Use agentic decisions only where the path really varies.
  • You cannot evaluate it yet. Without a labelled test set and tracing, you cannot tell whether the agent helps. Build evaluation first.
  • Retrieval itself is broken. An agent that retries a bad index just fails more slowly and more expensively. Fix data quality first.

A pragmatic middle ground is a router in front of a fast path: simple questions go straight through classic RAG, and only multi-part, multi-source or low-confidence questions are escalated to the agentic loop. Taking designs like this into production for enterprise customers, with real data, identity and evaluation, is the daily work of Forward Deployed Engineers.

Frequently asked questions

What is agentic RAG in simple terms?

It is RAG where an LLM agent decides how to retrieve: whether to search, which sources to use, how to rephrase or split the question, and whether the results are good enough. Classic RAG runs one fixed retrieval and answers.

What is the difference between agentic RAG and RAG?

Classic RAG retrieves once, then generates. Agentic RAG runs a model-controlled loop that can retrieve several times, route sub-questions to indexes, databases or APIs, check the evidence and retry. It handles harder questions but is slower and costs more.

What is corrective RAG?

A pattern where retrieved documents are graded for relevance before use. If they are weak, the system rewrites the query or searches another source instead of answering from poor evidence.

What is query planning in RAG?

Breaking a complex question into sub-questions, choosing a source for each, and ordering them when one depends on another, so multi-part questions get fully answered.

How do you stop an agentic RAG system from looping forever?

Set hard limits on retrieval rounds, tool calls, tokens and time; detect repeated queries; and treat "the answer is not in the available sources" as a valid result. When a limit is hit, return an honest partial answer instead of a guess.

Is agentic RAG better than GraphRAG?

They solve different problems and often work together. GraphRAG improves retrieval for relationship and multi-hop questions by using a knowledge graph. Agentic RAG decides which retrieval to run and when, and can use a graph as a tool.

How do you evaluate agentic RAG?

Score the answer for correctness, groundedness and completeness, and the trajectory for routing accuracy, per-step retrieval quality, grader accuracy, steps, cost and permission compliance. Compare against a simpler pipeline on the same test set.

Which tools are used to build agentic RAG?

Common choices are LangGraph or LangChain for orchestration, PostgreSQL with pgvector for vectors, SQL and API tools exposed via function calling or MCP, models from Amazon Bedrock, Azure OpenAI or Gemini, and Ragas, LangSmith or Langfuse for evaluation and tracing.

Getting agentic RAG right is mostly engineering: tool design, permissions, stopping rules and evaluation, not clever prompts. For guided, hands-on practice building and measuring RAG pipelines and LangGraph agents, explore Cloudsoft's GenAI and Agentic AI training in Hyderabad, in our Ameerpet classroom or live online. Call +91 96660 19191 for a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us