Notes
Blog
Research and working notes on AI, search, and production systems.
- Your reranker bake-off is a tie
As rerankers converge, nDCG on coarse human labels stops telling them apart. Rubric-Calibrated Preferences lets an LLM compare, a rubric anchor, and IRT put every query on one scale.
- Stop spending the cross-encoder inside the window
When the CE budget is smaller than the context capacity, scoring the obvious top is waste. BoundaryMORPH spends the budget on the k-boundary instead.
- Stop giving the agent a search box
Agent search APIs still look like string → ranked list. Algebraic Retrieval makes the query surface a composable program the agent can write, inspect, and revise.
- Stop ranking on questions humans typed
In agentic RAG the first-stage retriever sees agent reformulations, not TREC-style human queries. Q2D-Web is the first large-corpus benchmark built for that pairing.
- Don't decode a ranking as prose
Generative listwise rerankers won on quality. Autoregressive decode still bills you per token. SPD reads the ranking off the prefill and solves the assignment instead.
- The leaderboard assumed a complete issue
SWE-bench Verified scores agents on a finished GitHub issue. Real use starts mid-thought. Dialogue-SWEBench measures whether agents know when to ask — and when to stop.
- Calibrate to the assessor
SIGIR’s Best Student Paper this year is not another LLM judge prompt. It is a 26M LoRA that learned one assessor, on one topic, and beat the prompted frontier model.
- Give the agent the retrieval controls
Two papers this year handed the model the retrieval interface. One gave it Lucene. The other gave it grep, embed, and read. Both beat GraphRAG. The index algorithm is no longer the story.
- Stop summarizing the past
The default move in long-horizon agents is to shrink the past until it fits. Two papers say that is the bug. One keeps a playbook. The other keeps an index.
- Hybrid search is a program now
Weighted fusion over two frozen lists cannot express a real query. A 4B model learned to write the query plan instead — and beat GPT-5.5 at it.
- Your hybrid top-k is a different ranker
Truncating dense and sparse lists before fusion does not approximate complete-list RRF. It computes a different ranking — and the cutoff that worked last quarter may not transfer.