Skip to content

Blog

Stop giving the agent a search box

Agent search APIs still look like string → ranked list. Algebraic Retrieval makes the query surface a composable program the agent can write, inspect, and revise.

Most agent–search APIs still look like a search box with aspirations. The agent gets a tool: pass a string, maybe a top-k and a filter enum, get a ranked list. That is enough for demos. It is a weak contract for real work.

An agent looking for implementation architecture usually wants more than one string. Favor architecture docs. Suppress marketing. Restrict to engineering threads. Weight by recency. After the first page of results, change the strategy — subtract an unwanted direction, tighten the pool, reweight what survived. Today those choices live in glue code, hard-coded pipelines, or a second tool call with a different undocumented JSON schema. The retrieval program is invisible to the agent that is supposed to be driving it.

Damian Delmas (September 2026) calls the alternative Algebraic Retrieval (arXiv:2609.19482). The claim is narrow and useful. Give the agent a mathematical query surface over Programmatic Embedding Modulation (PEM): compose scoring directions, eligibility masks, and per-record weights in one expression; execute it; inspect the scored relation; revise the expression. Evaluate whether that contract is executable and portable — not whether it wins nDCG, and not whether the agent is good at writing the algebra yet.

What the search-box API actually hides

In classical IR, composition is not exotic. You fuse channels, apply filters, boost by metadata, rerank a candidate pool. PyTerrier made pipeline composition a first-class experiment language years ago. Vespa lets you write ranking expressions. What agent runtimes usually expose is the collapsed form: one opaque call that returns ten docs. The agent cannot see which stages ran, cannot subtract a direction it does not want, and cannot reuse an intermediate scored relation as an operand for the next turn.

That is a systems mismatch with how agents already behave. They inspect, branch, and revise. A search box forces them to encode strategy in natural language (“find architecture docs but not marketing”) and hope the retriever’s query understanding does the rest. Sometimes that works. Often it produces a bag of on-topic fragments that still fail the decision — the same failure mode we keep writing about when the leaderboard assumed a complete issue or when first-stage eval ignored agent reformulations.

What Algebraic Retrieval is

PEM exposes the embedding matrix and score arrays for arithmetic at query time. Algebraic Retrieval puts a notation on top. Scoring directions combine with + and - over similarity scores (or, equivalently, over query vectors before one matrix multiply). A mask removes ineligible identities. A weight relation scales each surviving score. top(k, S) selects. The agent discovers available operands through orient — which matrices, query vectors, masks, and weights the runtime has bound — then submits something like:

top(10, w ⊙ (m ▷ ((E @ q1) − 0.5 (E @ q2))))

Read that as: within candidate pool m, favor direction q1, suppress direction q2 at half strength, apply per-record weights w, keep ten. Changing a coefficient or adding a mask changes the strategy through the same call surface. Intermediate results stay scored relations the agent can inspect or feed into SQL joins for text and metadata.

That is the product thesis. Retrieval is not a chat completion. It is a short program over known operators, with a discovery protocol for what those operators are bound to.

What they actually evaluated

Delmas is explicit about what this paper is not. It does not claim better retrieval quality. It does not measure how accurately agents write algebraic queries. The fixture uses deterministic signed-hash vectors on the public Vaswani collection (11,429 documents), not a learned embedding bake-off. The evaluation is execution parity.

Three programs, each implemented three ways — Algebra, SQL via sqlite-vec, and PyTerrier transformers — are compared on the same fixture bindings. Contrastive scoring: subtract one similarity search from another. Candidate-pool reranking: score only a recorded BM25 pool of 52 docs. Full composition: contrastive scores inside the mask, then BM25-normalized weights, then top-10.

All three paths select the same document sets. Score differences sit below 1e-6 (typically ~1e-8). One tied pair flips order between Algebra/PyTerrier and the SQL path because float32 cosine rounding in sqlite-vec separates scores that NumPy treats as exact ties. The paper treats that as a feature of the claim, not a bug: score tolerance is not rank agreement. If you care about order under near-ties, you need an explicit tie policy and a gap check, not only “scores matched within epsilon.”

That is PyTerrier-era IR craft meeting agent runtimes. The point of the experiment is that the algebra is not vapor — it compiles to the same operations serious IR stacks already run.

Why this matters if you ship agents

We have been arguing two related points on this blog. Hybrid search is a program: fusion should be executable, not a fixed weight you forgot to tune. Give the agent the retrieval controls: Boolean structure or tool-level grep/embed/read beats dumping a black-box top-k. Algebraic Retrieval is the next sentence. If the agent is supposed to own the retrieval strategy across turns, the API has to expose composable score arithmetic — directions, masks, weights — and return relations it can revise after looking.

The build implication is an interface redesign, not a new embedder. Expose orient (what can I compose?). Accept expressions, not only strings. Keep scored identities through every stage so the agent can inspect and reuse them. Make SQL or pipeline equivalents the executable spec of the algebra, and test parity the way Delmas does — including ties. Separate the questions “can the runtime execute the program?” and “can the agent write a good program?” The first is a contract. The second is an eval you still owe.

There is also a quieter ops win. When retrieval is an expression, diffs become readable. You can log the program, unit-test it against fixture bindings, and argue about a coefficient instead of arguing about prompt poetry. That is how search teams already debug PyTerrier and Vespa ranking expressions. Agents should get the same affordance.

What not to claim

Do not sell Algebraic Retrieval as “better search.” This paper measures whether three implementations of the same program agree, on hash vectors, on one classic collection. It does not show that agents choose better masks than your current glue code. It does not replace ANN recall studies, learned sparse models, or hybrid fusion quality work. PEM and the algebra assume you have a place to bind matrices, masks, and weights; a flat blob store with no eligibility relations still needs that plumbing. And the float32 tie flip is a reminder: shipping “scores matched” dashboards without rank checks will bite you in production.

What you can claim is sharper. The search-box tool is a lossy projection of retrieval. Agents need a query surface that matches how they work — compose, inspect, revise. Algebraic Retrieval is a concrete contract for that surface, with executable counterparts in SQL and PyTerrier and a clear parity test.

Stop giving the agent a search box. Give it an algebra — and let it revise the program after it sees the first page.