Stop ranking on questions humans typed
In agentic RAG the first-stage retriever sees agent reformulations, not TREC-style human queries. Q2D-Web is the first large-corpus benchmark built for that pairing.
If you pick a first-stage retriever the way most teams still do, you run a bake-off on a public IR collection. BEIR tasks. MS MARCO. Maybe a TREC track. The queries are written by people — search-log users, crowdworkers, domain experts. The winner is the model with the best nDCG@10 or Recall@100 on those strings. You ship it under an agent.
That bake-off answered the wrong question.
In an agentic RAG stack the first-stage retriever often never sees the user’s words. It sees what the agent typed into the search tool: a reformulated primary query, then a sequence of support queries that decompose, clarify, or chase adjacent entities. Those strings differ in length, vocabulary, and intent shape from classic TREC-style human queries. Penha and colleagues showed years ago that even intent-preserving reformulations cut pipeline nDCG@10 by about 20% on average. Production agents do this on every turn. Your leaderboard usually does not.
Schall, Eslami, Krimmel, Chaffin, Milliken, Wang, and Bykov (Perplexity, September 2026) built Q2D-Web to close that gap (arXiv:2609.08887). The claim is infrastructural, not architectural. You cannot honestly rank first-stage models for agentic RAG on human-query benchmarks over tiny corpora. You need the production pairing: a large web index, agent-reformulated queries drawn from real traffic, and deep enough judgments that false negatives do not decide the bake-off.
What existing benches miss
Public IR collections usually fail at least one dimension of that pairing. Large corpora come with few test queries. Large query sets sit on only millions of documents. Relevance labels are often shallow — one click, one gold passage — so unlabeled relevant documents are scored as misses. And almost every collection evaluates human-written queries, while the first stage in a deployed agentic pipeline serves machine-written reformulations.
That matters because the first stage bounds what the agent can read and cite. Later rerankers only reorder the candidate set. If the retriever never recovered the evidence, no amount of listwise polish fixes it. Last week’s post was about decoding a ranking without pretending it is prose. This one is about whether the candidate set was built for the queries your production system actually issues.
What Q2D-Web is
Q2D-Web (Query2Doc-Web) is a ~190M-document web corpus paired with ~70k agentic search queries in ten languages, reformulated from nine months of PII-filtered production traffic. Roughly 17.7% of the queries are primary searches; the rest are support queries that follow from the same thread — on average about four support queries per primary. Mean positives per query sit near 100 under their combined labels. That depth is the point: shallow click labels systematically punish systems that retrieve relevant-but-unlabeled documents.
They publish three fixed judgment sets over the same queries and corpus. One comes from agent citations. One comes from production rankings. The third unions both and adds LLM judgments on unlabeled pooled candidates to cut false negatives. You can score a retriever three ways and ask whether the choice of label source moves the ranking.
They keep the corpus, queries, and judgments private on purpose — public benches rot into training data — and run a public leaderboard for open-weight models: huggingface.co/spaces/perplexity-ai/q2d-web-leaderboard. Treat that as a gate for vendors, not as a claim that you can reproduce every number offline without their harness.
What they found
Thirteen retrievers: BM25 (tantivy), a spread of dense encoders, and late-interaction models. Primary metric is Recall@1000 — the right depth for a first stage that feeds a reranker — with Recall@100 and nDCG@10 reported for shallower cuts.
Relative orderings are largely stable across the three judgment sets. That is reassuring: citation labels, production rankings, and the LLM-augmented combined set do not reshuffle the field. The rankings do diverge across topical domains, query languages, and query types. Aggregate leaderboard wins are not portable. A model that looks strong overall can lose a domain slice or a language slice you actually care about.
Query type is load-bearing. Neural retrievers show lower recall on support queries than on primary queries. BM25 shows the opposite trend — slightly higher recall on support — while remaining weakest on aggregate Recall@1000. Uniqueness tells a related story: the number of relevant documents a retriever recovers that no other system finds is not ordered by aggregate recall. BM25 contributes by far the largest exclusive positive set. Family matters independently of scale. If you only keep the top dense model from a bake-off, you throw away a recall channel the leaderboard undervalues.
They also study cheaper evaluation. Retaining about a third of the corpus (~60M docs), selected by reciprocal rank fusion over pooled runs with a fixed distractor budget, preserves full-corpus model ranking under the combined judgments while raising absolute Recall@1000 only about 3–7 points. That is a practical two-stage protocol: score new models on the RRF subcorpus first, then spend full-corpus compute on the ones that survive. It is also another reminder that truncated hybrid fusion is a different ranker — here RRF is used as a sampling policy for eval, not as the production fuse.
Why this matters if you ship agentic RAG
Most teams still optimize first-stage choice against human-query distributions. That is rational when the product is a search box. It is a category error when the product is an agent that rewrites, decomposes, and iterates. You can “win” BEIR and still mis-rank systems for the query distribution your rewriter actually emits.
The build implication is concrete. Separate the bake-offs. Keep a human-query suite if you still have a classical search surface. Add an agent-query suite — even a private one — sampled from your own tool-call logs, with judgments that are not a single gold doc. Score Recall at the depth your reranker actually consumes. Slice by primary vs support, by language, by domain. Watch unique recall by family, not only the aggregate table. And if you cannot afford full-corpus runs every time, use a pooled RRF subcorpus as a screen — then verify winners on the full index.
This sits next to posts already on this site without repeating them. Hybrid search as a program argues the fuse should be executable. Giving the agent retrieval controls argues the agent should be able to express what it needs. Q2D-Web argues the eval has to match the queries those controls produce. Otherwise you are tuning the wrong distribution and calling it science.
What not to claim
Do not install Perplexity’s production rewriter as everyone’s agent. Their reformulation distribution is theirs; yours may differ in length, language mix, and decomposition style. The combined judgment set uses LLM labels on pooled docs — useful against false negatives, still a judge with the usual assessor risks. Web content is not your private enterprise corpus. They did not include learned sparse models such as SPLADE in this round. And a private corpus plus public leaderboard is an honest contamination hedge — it is also a trust boundary. You are evaluating through their harness.
What you can claim is narrower and more useful. First-stage retrieval in agentic RAG is scored, in most public work, on the wrong query origin. Q2D-Web is the first large-corpus attempt to score it on the right one. If your internal dashboard still only shows human-query BEIR, you are optimizing a distribution your production agent does not speak.
Stop ranking first-stage models on questions humans typed. Grade them on what agents actually search.