Agentic RAG: When Should It Retrieve?
Naive RAG retrieves on every turn, even when it shouldn't. How agentic RAG lets the model route, grade its own context, and re-retrieve, plus the tradeoffs.
The first RAG (retrieval-augmented generation) system you build retrieves on every turn. A question comes in, you embed it, pull the top-k chunks, staple them to the prompt, and generate. It works well enough in a demo that nobody questions its shape. Then real users show up and the cracks appear:
- Someone types “thanks, that helped,” and the pipeline dutifully runs a vector search for gratitude.
- Someone asks a follow-up that only makes sense given the previous answer, and retrieval fires on three words with no context.
- Someone asks a two-part question, and the single search only covers half of it.
The problem is that retrieval is wired in as a reflex instead of a decision. Naive RAG treats “fetch documents” as a fixed step. It pays the cost every time and injects whatever the index returns, even when the returned chunks are noise. Agentic RAG moves that decision into the model, which gets to ask three questions. Do I need to retrieve at all? Is what came back actually useful? Should I search again with a better query before I answer?
This post is for engineers who have shipped a basic RAG pipeline, watched it misfire on exactly these cases, and want to understand what “agentic” buys you and what it costs. I hit all of this building Archi, a RAG copilot that answers questions from CERN’s operations docs, JIRA tickets, and logbooks. Operators ask a mix of things. Some need a specific runbook, some are follow-ups on a running conversation, and some are chit-chat around a real request. One retrieval strategy for all of them is the wrong tool.
What “agentic” actually changes
Classic RAG, in the sense of the original retrieval-augmented generation paper, is a straight line from query to answer. Agentic RAG is a loop with a couple of decision points inside it. The word “agentic” gets thrown around loosely, so here is the concrete version:
- the model can call retrieval as a tool, rather than having retrieval called for it, and
- a control loop around that tool inspects the results before committing to an answer.
Two capabilities do most of the work:
- Routing decides whether this turn needs external documents at all.
- Grading looks at what retrieval returned and judges whether it answers the question, before you spend a generation on it.
Everything else, including the query rewriting and the retries, hangs off those two.
None of this is new machinery. Letting a model interleave reasoning with tool calls is the ReAct pattern, and treating retrieval as one of those tools is the natural extension. What agentic RAG adds is structure around the loop, so it terminates and stays cheap.
Routing: skip retrieval when you don’t need it
The cheapest retrieval is the one you never run. A surprising fraction of turns in a real chat don’t need the index. Acknowledgements, clarifying questions, requests to rephrase the last answer: the model can handle anything like that from the conversation it already has. Running a vector search for those turns wastes latency. Worse, it drags in irrelevant chunks that can pull the answer off course.
The router is a small classification step in front of retrieval. You can do it with a cheap model and a tight prompt. The prompt below asks for a one-word decision, and route maps anything containing RETRIEVE to a lookup:
ROUTER_PROMPT = """You decide if a user turn needs a documentation lookup.
Answer with one word: RETRIEVE or ANSWER.
RETRIEVE: questions about specific systems, procedures, errors, or facts
that live in the docs.
ANSWER: greetings, thanks, chit-chat, or follow-ups the prior context
already covers.
Conversation so far:
{history}
User: {question}
Decision:"""
def route(question: str, history: str) -> str:
out = cheap_model.complete(ROUTER_PROMPT.format(
history=history, question=question))
return "retrieve" if "RETRIEVE" in out.upper() else "answer"The honest tradeoff is that the router is another model call, and it can be wrong. Route a real question to ANSWER and you get a confident hallucination with no sources. So bias the router toward retrieval when unsure, keep the classes coarse, and log the decisions so you can see what it’s skipping. On Archi I’d rather retrieve a few times too often than answer an operational question from the model’s imagination.
Grading: don’t trust the top-k
Vector search always returns something. Ask about a system that isn’t documented and you still get the k nearest chunks; they’re just nearest to nothing useful. Naive RAG hands those to the generator anyway. The model, trying to be helpful, writes an answer grounded in the wrong document. This failure mode erodes trust fastest, because the answer looks sourced.
Grading puts a check between retrieval and generation. After you pull chunks, a grader judges whether they’re actually relevant to the question. This is the core idea behind Corrective RAG and the self-reflection in Self-RAG: the system critiques its own retrieval and acts on the verdict.
A grader can be as simple as a per-chunk yes/no. This version asks the cheap model about each chunk in turn and keeps only the ones it marks YES:
GRADER_PROMPT = """Question: {question}
Chunk: {chunk}
Does this chunk contain information that helps answer the question?
Reply YES or NO."""
def keep_relevant(question: str, chunks: list[str]) -> list[str]:
kept = []
for c in chunks:
verdict = cheap_model.complete(
GRADER_PROMPT.format(question=question, chunk=c))
if "YES" in verdict.upper():
kept.append(c)
return keptIf nothing survives grading, that’s a signal, not an error. It usually means the query was phrased differently from the documents, and that is where reformulation earns its place. Rewrite the question and search again. It is the same move I wrote about in query rewriting for retrieval, except now the loop decides when to do it instead of doing it unconditionally.
The loop needs a retry budget
This is where agentic RAG bites people who build it without guardrails. Grade, reformulate, retry, grade again: with nothing stopping it, a query the index genuinely can’t answer will loop until it times out or empties your token budget. The grader keeps saying “not relevant” because the documents really aren’t there. The loop keeps trying to fix a problem that has no fix.
So the loop is always bounded. Cap the retries, and when you hit the cap, stop and answer honestly. The function below ties routing, grading, and reformulation together; note the fallback to generate_with_caveat once the attempts run out:
def agentic_rag(question, history, max_retries=2):
if route(question, history) == "answer":
return generate(question, history, chunks=[])
query = question
for attempt in range(max_retries + 1):
chunks = keep_relevant(question, retrieve(query))
if chunks:
return generate(question, history, chunks)
query = reformulate(question, history) # try a different phrasing
# Budget spent, still nothing relevant. Say so instead of guessing.
return generate_with_caveat(question, history)Two retries is usually plenty. If the second reformulation still comes back empty, more attempts rarely help, and “I couldn’t find this in the docs” is a far better answer than a fabricated one. That caveat path matters as much as the happy path. An operator who is told the runbook isn’t indexed goes and asks a human. An operator handed a confident wrong answer acts on it.
What agentic RAG costs
Agentic RAG is strictly more expensive than the naive version. Name the costs before you commit.
Every extra decision is a model call. A single turn can now involve a router call, one or more grading passes, a reformulation, and the final generation, where naive RAG did one retrieval and one generation. That is latency the user feels and tokens you pay for. Use a small, cheap model for the router and grader, keep their prompts short, and grade chunks in parallel rather than in a loop when you can.
The system becomes non-deterministic in a new way. With naive RAG, the same question retrieves the same chunks. With a loop, the path depends on model judgments that can vary. That makes bugs harder to reproduce and makes tracing essential. Instrument every hop (the route decision, the grades, the reformulations) so a weird answer can be replayed. I lean on OpenTelemetry tracing across the RAG pipeline for exactly this.
The graders can be wrong in the same ways the generator is. A grader that rubber-stamps everything gives you naive RAG with extra steps and extra cost. One that’s too strict throws away good chunks and pushes you into needless retries. You tune the grader against a labeled set, the same way you’d tune retrieval itself. That is the whole reason to have a way to measure retrieval quality before you start adding loops on top of it.
When not to bother
Agentic RAG is not the default. Suppose your traffic is narrow, every question needs the docs, and the queries already match the document phrasing. Then naive RAG is simpler, faster, cheaper, and easier to debug. Reach for the loop when:
- the traffic is mixed,
- retrieval quality is uneven, or
- a wrong-but-confident answer carries a real cost.
That’s the situation on Archi. Operators mix chit-chat with high-stakes questions during an incident, and an ungrounded answer at the wrong moment is worse than no answer.
The instinct underneath all of it is the same one I bring to the WMCore and Unified operations stack: don’t do work you don’t need to, and don’t trust a result just because a system produced it. Naive RAG violates both. It retrieves when it shouldn’t and believes whatever the index hands back. Agentic RAG is the version that asks first and checks after, and for anything past a demo, that’s usually the version worth the extra calls.
Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.