Skip to content

Blog

Give the agent the retrieval controls

Two papers this year handed the model the retrieval interface. One gave it Lucene. The other gave it grep, embed, and read. Both beat GraphRAG. The index algorithm is no longer the story.

GraphRAG had a good year in demos. Build a knowledge graph, retrieve subgraphs, stuff entities and relations into the prompt, watch multi-hop QA look smarter. The pitch is that the missing ingredient was structure — that frozen retrieve-and-stuff fails because chunks do not know who is related to whom.

Two papers this year gave the model something else: the retrieval controls. One handed it Lucene. The other handed it grep, embed, and read. Both beat GraphRAG. That is the story, not another embedding bake-off, and not another graph construction bill.

They agree on the diagnosis. Static retrieve-and-stuff is the wrong default. Frozen IRCoT and FLARE workflows are the wrong default. The retrieval interface is the product, not the index algorithm. More retrieved text is not more evidence.

They disagree on what the hands should be. I have shipped one camp in production search. I have watched agents need the other. The lean is situational, not tribal.

Camp 1: Lucene is enough

LogicalRAG (Zeng, Deng, Wan, Jiang, Zheng, Huang; May 2026) keeps the same agent loop as a strong agentic hybrid baseline and changes only the action space. The model writes Lucene-style Boolean, phrase, and field queries. OpenSearch’s inverted index decides membership. BM25 ranks inside the matched set, not instead of the constraints.

That last sentence is the Solr instinct. Boolean is not a soft suggestion. It is membership. Rankers reorder what is in; they do not quietly admit what the query excluded. If you have ever debugged a regulated search stack where “almost relevant” is a liability, you already know why that matters.

On KILT-scale Wikipedia, LLM-judge accuracy is 0.717 against 0.716 for the hybrid. Index build is 1.27 hours versus 52.02 — about 41× cheaper, because you skip corpus-wide embeddings and FAISS. At concurrency 16 they report 152.5 QPS against 66.6, and 74.9 ms mean latency against 230.5. When gold passages are removed, refusal goes to 0.828 from 0.767 and hallucination drops to 0.083 from 0.128. A failed Boolean search looks like a miss. A near-miss embedding often does not.

Two caveats belong in the body. The headline metric is an LLM judge, not exact match — so treat the absolute numbers as comparative under one judge, not as ground truth. And the design needs a capable model: at Qwen3.5-4B, LogicalRAG trails hybrid by about three points; at Qwen3.5-Plus they are at parity. A weak model with Lucene loses. The interface is not magic; the agent has to be able to write good Boolean.

They did not run graph backends at KILT scale. Graph construction at that size was outside budget — tens of billions of tokens of entity and relation extraction. That absence is itself a result. If your GraphRAG story assumes a graph you cannot afford to build, the comparison was never fair.

The OpenSearch-native reading is the one I recognize from production. If the agent can write Boolean, you may not need a vector index. Exact intent, fail closed when the answer is not there. That is the right default for a lot of regulated search — and for any corpus where a confident near-miss is worse than an honest empty result.

Camp 2: grep, embed, then read

A-RAG (Du, Xu, Zhu, Wang, Wang, Wang, Mao; 2026) keeps dense search in the toolkit. Three tools: keyword_search, semantic_search, chunk_read. The index is light: ~1,000-token chunks, sentence embeddings, no pre-built inverted index, no graph. Keyword here is runtime exact match, not Lucene. ReAct, one tool at a time, refuses to re-read chunks it has already seen.

With GPT-5-mini, Full A-RAG beats GraphRAG, HippoRAG2, LinearRAG, MA-RAG, and RAGentA on HotpotQA, 2Wiki, MuSiQue, and GraphRAG-Bench. The embarrassment is Naive A-RAG: one embedding tool, still agentic, already beating most of the graph and workflow methods. The graph was not the missing ingredient. Letting the model decide when to retrieve was.

Granularity is the actual lever. Full A-RAG retrieves 2,737 tokens on HotpotQA against Naive’s 27,455, at higher accuracy. Snippets first, full chunk only when the agent asks. More context was the bug. That should sound familiar if you have watched agents drown in stuffed passages and then “reason” over noise.

They reviewed 100 MuSiQue errors. 82% were reasoning-chain failures, not “couldn’t retrieve.” Entity confusion is the top secondary mode. The documents were in the room. The model mixed up who was who. Retrieval was not the bottleneck on those failures. Judgment was.

A-RAG is a multi-hop QA study, not a production search stack. Do not install it as your serving layer. Keyword-as-runtime-exact-match will not survive a large corpus the way a real inverted index will. The point is the interface: keyword / sentence / chunk-read, chosen by the model, on a messy corpus where the query has to be discovered rather than declared up front.

What they share, and where I lean

Both papers are arguing that GraphRAG lost to a model that could do something with retrieval, not to a smarter graph. LogicalRAG’s something is precise Boolean over an inverted index. A-RAG’s something is a small set of tools at different granularities. Production teams will recognize both. I have shipped the first. I have watched agents need the second.

The lean, not a dunk: Boolean is the right default when intent is exact and an unanswerable query must fail closed. Hierarchical tools win when the corpus is messy and the model has to discover the query. Hybrid search is a program is the third pole if you need SQL and vectors in one action space — ProRetrieval’s DSL rather than Lucene alone or grep/embed/read alone. Three interfaces, one principle: the agent has to be able to express intent.

There is a quieter industry implication. A lot of 2025–2026 spend went into graph construction, embedding everything, and workflow scaffolding that decides retrieval for the model. These papers push the spend toward the interface: what can the agent say, what does a miss look like, how much text does it pull before it asks for more. Index algorithms still matter. They are no longer the whole story.

What not to claim

Do not flatten this into “agentic RAG is good.” A weak model with Lucene loses to hybrid. A keyword tool that is not a real inverted index will not survive a large corpus. An LLM judge can flatter a paraphrase — LogicalRAG’s headline accuracy sits on that judge. Graph absence at KILT scale is a cost result, not a proof that graphs never help. A-RAG’s wins are on multi-hop QA benchmarks, not on your ticket corpus with your latency SLO.

Measure whether a miss looks like a miss, on your schema, with a fallback when the plan does not compile. Pair that with the EAHR question from earlier in the week: even when the agent can express intent, truncated fusion can still compute a different ranking than the one you named. Interface and ranker are separate contracts.

The question is no longer which index is smarter. It is whether the agent can express intent, and whether a miss looks like a miss.

Papers: LogicalRAG (Zeng et al., 2026). A-RAG (Du et al., 2026); code.