Skip to content
Ishaan Reddy

Blog · architecture

Advanced RAG in Practice: Why Hierarchical Retrieval Beats Flat Search

2026-08-115 minRAGretrievalembeddingsrerankingLLMs

Advanced RAG is a loose name for the work that happens once basic vector search stops being good enough. You still retrieve context before asking a language model to answer, but retrieval becomes a pipeline: route the query, narrow the search space, fetch candidates, rerank them, and only then build the model's context.

I ran into this while building a retrieval layer at AILACDS for a schools platform. Teachers could upload notes, old exam papers, reference textbooks, and administrative records, then retrieve the relevant material while preparing lessons or answering questions. We designed it around a possible 8,000 students and a few hundred teachers sharing a growing corpus. Those numbers were a capacity target, not a measured count of concurrent users.

At that size, retrieval quality was the harder problem. A fast response is not useful if it pulls the wrong page from the wrong kind of document.

Where flat retrieval starts to struggle

A basic RAG retriever usually does three things:

  1. Turn the query into an embedding.
  2. Compare that vector with every indexed chunk.
  3. Return the top-k chunks with the highest similarity scores.

That approach works well for a small collection of similar documents. The embedding has one job: find chunks that mean roughly the same thing as the query.

A school corpus has more structure. A chunk can belong to a subject, class, month, document type, or teacher. A query can contain several of those constraints at once. Flat retrieval compresses all of them into one vector and asks one similarity score to decide which constraint matters most.

Take this query:

What are the days with leaves in May?

The word “leaves” carries the category: leave records. “May” carries the time range. A flat search compares the complete query against textbook pages, biology notes, attendance records, calendars, and every other embedded chunk. Semantically strong matches for “leaves” can outrank the records for May.

Comparing Naive RAG to Hierarchical RAG

Hierarchical retrieval gives each part of the query a smaller job. The first level identifies the broad category. The second level searches within that category for the finer constraint.

For the example above, the first match routes the query to leave records. The second match looks for May inside that subset. Biology chapters and unrelated school documents never enter the second search.

The diagram shows the retrieval stages, not a required database layout or embedding model. The hierarchy could come from separate vector indexes, metadata filters, parent-child records, or a routing layer in front of one index. The useful part is the order of operations: resolve the coarse constraint before scoring the fine-grained records.

This changes the competition at the second stage. A May leave record now competes with other leave records instead of every chunk in the corpus. The embedding still performs approximate matching, but it works inside a cleaner candidate set.

There is another practical benefit. The two stages are easier to inspect. If a result is wrong, you can ask whether the category router chose the wrong branch or whether the search inside the correct branch ranked the wrong record. A single similarity score across the whole corpus gives you less information about where retrieval failed.

Rerank the candidates before generation

The hierarchical pass narrows the search space, but its similarity scores still come from an approximate first pass. I added an LLM reranker after retrieval.

The retriever returns 20 candidate chunks. The reranker reads the full query and scores each candidate for how well it answers that query, then orders the candidates again. Only the highest-ranked chunks move into the generation context.

This split gives the two stages different jobs. Vector search finds a small candidate set without reading every document at query time. The LLM spends more computation on those 20 candidates, where a closer reading can improve the final order.

For “what are the days with leaves in May?”, the reranker can prefer a record that lists exact May dates over a general leave-policy document. Both may be semantically related to the query, but only one contains the requested answer.

Reranking does add latency and model cost. It also needs a stable scoring prompt and its own evaluation set. Sending hundreds of candidates to the reranker would throw away much of the speed gained during retrieval, which is why the candidate count stays bounded.

What the hierarchy costs

Hierarchical retrieval depends on structure you can trust. Documents need useful metadata or a reliable way to infer their category. If the first stage routes a query to the wrong branch, the correct chunk may never become a candidate.

The system also has more parts to evaluate. A single end-to-end accuracy number cannot tell you which stage needs work. I would track at least three questions separately:

  • Did the router choose the correct category?
  • Did retrieval include the answer-bearing chunk in its 20 candidates?
  • Did the reranker place that chunk near the top?

Each question maps to a different fix. Routing errors call for better labels or category descriptions. Candidate-recall errors point to the chunking, embeddings, filters, or retrieval depth. Ranking errors point to the LLM prompt or scoring method.

Hierarchy is also unnecessary for some corpora. If you have a few hundred similar documents and flat retrieval already returns the correct context, another retrieval layer adds work without solving a measured problem.

Use hierarchical retrieval when the corpus already has meaningful divisions and your evaluation shows that flat search mixes them together. Start with the structure people use to browse the data themselves, keep the first-stage branches understandable, and measure whether the answer-bearing chunk survives each step.