Skip to content
Ishaan Reddy

Blog · architecture

Three Levels of RAG

2026-08-2510 minRAGretrievalembeddingsrerankingagentsLLMs

Most teams do not design a RAG pipeline in one pass. They ship the simplest version that answers questions, watch it fail on real queries, and add a stage to fix the specific failure they saw. Do that enough times and the pipeline accumulates a shape nobody planned up front.

I have gone through that progression twice now: once on a personal project, once on a schools platform at AILACDS where teachers retrieve notes, exam papers, and administrative records from a shared corpus. Three stages showed up in a fixed order, because each one only becomes necessary once the previous one runs out of room. Two more things get talked about as if they belonged on that same ladder, and neither one does. Both are features you turn on for the queries that need them, not stages the whole system graduates through.

1. Basic RAG

Chunk documents, embed the chunks, store the vectors, embed the query, retrieve the top-k by similarity, hand them to the LLM as context. This is the version every RAG tutorial starts with (a PDF chatbot, a "chat with your docs" demo), and it is a reasonable place to start. It takes an afternoon to build and it works on small, homogeneous corpora.

It breaks down for a boring reason: one vector cannot represent a query with more than one requirement. "What are the days with leaves in May?" carries a category (leave records) and a time range (May) folded into a single embedding, and nothing in the pipeline can unfold them again. A semantically strong match for "leaves" can outrank the correct May record without the system ever being wrong about similarity. It answered the question it was actually asked, which was "what's close to this vector."

You know you are still at Basic RAG if retrieval quality tracks corpus size inversely: fine at a few hundred documents, noticeably worse once the corpus mixes document types.

2. Advanced RAG

Advanced RAG keeps the same single retrieve-then-generate pass but tunes what goes into and comes out of it:

  • Better chunking: respect document structure (a full section, a full Q&A pair, a full record) instead of slicing at a fixed character count that might cut a table in half.
  • Metadata filtering: tag chunks with subject, class, month, or document type at ingestion time, so retrieval becomes a filtered search instead of a search over the whole index.
  • Hybrid search: combine vector similarity with keyword search (BM25). Vector search misses exact identifiers (a student ID, a document number, an exact date string) that keyword search catches trivially.
  • Query rewriting: expand or clarify the query before embedding it, so "leaves in May" becomes something closer to "leave records for the month of May."
  • Reranking: pull a generous candidate set (top 20-50) with a fast first pass, then have a slower model, an LLM or cross-encoder, read the full query against each candidate and reorder them.
  • Better context selection: once candidates are ranked, decide how many actually go into the prompt and in what order, instead of dumping the full top-k in arbitrary order.

Strung together, those fixes still add up to one pass, just a longer one:

Each of these is a targeted fix, and you can add them independently. Reranking is the one that earns its cost most reliably: it only helps once a better first pass already puts the right chunk somewhere in the candidate set, so add it when you can point to specific queries where the right chunk exists in the candidates but ranks too low.

Advanced RAG still runs one pass per query, with a fixed sequence of fixes applied to it. It fails once the corpus has structure that no amount of tuning a single pass can walk through: the "leave records" example above is exactly that case, since the query needs its category resolved before its time range does, and one pass has no way to do the two in order.

3. Dynamic Retrieval

The third stage stops treating retrieval as one step and starts treating it as a small pipeline of steps: resolve a category, then search inside it; try a search, then decide whether the result is enough; search one source, then follow what it found into a second source.

This is a spectrum, not a binary, and where a query sits on it is a cost and complexity dial, not a separate tier to graduate into:

At the fixed end, a router decides the retrieval path up front, before any retrieval happens, based on the query type: leave records go through a metadata-filtered search, open-ended questions go through hybrid search. That decision is the same every time for a given query type, which makes it cheap to run and easy to test.

At the live end, the decision gets made one hop at a time, based on what the last hop returned, instead of being fixed in advance. Both ends are solving the same problem: a single retrieval pass can't resolve a question that depends on more than one decision. Moving from one end to the other is a dial you turn as more of your query types need the next step decided after seeing the last result, not a jump to a different capability.

The two ends look different mostly because of how many steps get decided before the first retrieval call versus during it:

The fixed-end example takes one decision and one search. The live-end example takes four decisions, each one shaped by what the last search returned, and the model would not have known to search that region or that operations log before seeing the sales reports first.

Reach for Dynamic Retrieval when you can point to queries where a single-step pipeline genuinely can't produce the right retrieval path, because the corpus has categories a query needs to walk through in order, or because the right next search depends on what the last one returned.

Additional improvements

Live, per-hop decision-making and self-reflective checking both get described as if they were a fourth and fifth stage above Dynamic Retrieval. Neither one is. Both are opt-in per query type or per use case, and you can turn either one on without having built the other, or without having pushed Dynamic Retrieval anywhere near its live end.

Live/agentic decision-making

This is the far end of the Dynamic Retrieval dial, called out on its own because treating it as a system-wide upgrade is the mistake teams actually make. It is worth turning on only for the specific queries where the next retrieval step genuinely cannot be decided until the last one's result is in hand: a multi-hop question like "Find why our sales dropped in Q2," where the model has to search sales reports, notice the numbers point to one region, search that region's documents specifically, and cross-check against a separate operations log before it can answer.

It does not replace anything below it. A live decision loop still has to chunk well, filter well, and rerank well underneath, or every hop it takes inherits the same retrieval quality problems Basic and Advanced RAG already had to solve. The costs are specific to the live end of the dial: an unbounded number of retrieval calls per query, failures that vary run to run instead of failing the same way every time (stopping too early, looping on a rewritten query that never improves, searching the wrong source), and an evaluation problem that shifts from a single test case to a behavior you have to check across many runs.

Self-reflective / corrective check

This is a verification layer, not a retrieval technique, and it bolts onto any of the three stages above, including Basic RAG with no router and no agent loop at all: grade the retrieved evidence before generating, and fall back to a broader search if the grade is low; after generating, verify that a citation actually says what the answer claims it says, instead of citing a chunk that was merely nearby in the ranking; detect thin evidence and retrieve again rather than answering from it.

The question that decides whether to add it has nothing to do with which of the three stages a system is running. It is whether a wrong answer here is expensive: cited somewhere, acted on, or shipped without a human reading the source before it matters. A pipeline stuck at Basic RAG with a grading check in front of every answer can be a better bet than a fully dynamic pipeline with no check at all, if the two systems are answering questions with very different costs of being wrong.

Picking what to build

Basic, Advanced, and Dynamic Retrieval are a real ladder: each one fixes a failure the one before it can't, and each adds a cost (latency, evaluation surface, or unpredictability) that isn't worth paying without a measured failure to justify it. Climb it by looking at what's actually going wrong: corpus size and mixed document types push you from Basic to Advanced; a corpus with real internal structure, or a question whose next search depends on the last one's result, pushes you into Dynamic Retrieval.

The live end of that dial and the self-reflective check are a different kind of decision. Turn on live, per-hop retrieval for the specific queries that need more than one dependent search, not for the whole system. Turn on the reflective check for the specific answers where being wrong is expensive, whatever stage the retrieval underneath happens to be at. Neither one is something you earn by finishing the stage before it.