Skip to content

Blog

Hybrid search is a program now

Weighted fusion over two frozen lists cannot express a real query. A 4B model learned to write the query plan instead — and beat GPT-5.5 at it.

The last post was about a quieter failure: your hybrid top_k is not an optimization of RRF. It is a different ranking function. Your hybrid top-k is a different ranker argued that truncated fusion and complete-list fusion are different objects, and that production Elasticsearch and Azure cutoffs quietly ship the truncated one.

This one is about the next step. Even if you get complete-list fusion right, the thing most production teams still ship is weighted fusion over two frozen lists. Real queries are not two lists. They are Boolean over filters plus meaning — last week’s emails from marketing about the Q3 budget, not the archived ones; a Sony or Bose headphone under $200 with strong noise-cancelling reviews, in matte black. That is a program. A 4B open model learned to write it.

I have been writing that program by hand for a long time — Solr filters, OpenSearch bool queries, the algebra of must / should / must_not with a vector clause bolted on. The new claim is that the model can author the plan, not just embed the bag of words.

Fusion is not an action space

RRF is a weighted disjunction of ranks. LangChain-style self-query is a conjunction of key–value filters plus one semantic clause. Neither can say “this brand or that one, under this price, not archived, and about this topic.” Search-R1 and DeepRetrieval train a model to write a better query, but still against a single backend. The orchestration — which constraints are exact, which are semantic, how the candidate sets combine — stays outside the model, frozen by whoever designed the pipeline.

That freeze is comfortable. Pipelines are reviewable. Weights are tunable. The cost is expressiveness. Users do not speak in two ranked lists. They speak in constraints, disjunctions, negations, and “aboutness.” If your action space cannot say what they meant, no amount of embedding quality will invent the missing operators.

You, Sun, Hu, Zhou et al. (August 2026) call the alternative ProRetrieval. The model is not a reranker. It is a retrieval orchestrator. Given a natural-language query, it emits a hybrid DSL: SQL over structured fields, plus text and image vector primitives, fused by placeholder injection (id IN <text_0>). SQL is the algebra. Solr and OpenSearch people have been living in that algebra for a long time. The new part is that the model authors the plan.

They train Qwen3-4B with supervised fine-tuning, then GRPO and DAPO, under a four-term reward: format, execution, Hit@1, length. Syntax first, then “did it run,” then “did it retrieve the right document,” then brevity. That ordering matters. A pretty plan that does not execute is not a retrieval system. A plan that executes and retrieves the wrong document is not a win either. The reward stack matches how you would debug a query DSL in production.

A 4B model beat GPT-5.5 at writing the plan

On two new benchmarks — Amazon products (structured + text + image) and Enron email (structured + text) — Qwen3-4B with DAPO hits 0.808 and 0.909 Hit@1. GPT-5.5, given the schema and few-shot DSL examples, hits 0.693 and 0.855. Claude Opus 4.7 is behind that. Dedicated training on a task-specific action space beat a much larger model that only understands the syntax.

That result should make teams pause before the next “just prompt GPT to write the filter” demo. Prompting is not the same as learning an executable action space. The frontier model sees the schema and the examples. The 4B model was trained to emit plans that run and hit. The gap is not surprising once you say it that way. It is still uncomfortable if your roadmap assumes bigger general models will absorb the orchestration job for free.

The interesting failure is in the ablation. Under SFT alone, the full DSL (SQL + text + image) underperforms the smaller SQL+text action space: 0.680 vs 0.752 Hit@1 on e-commerce. Three modalities at once is multi-task interference. Imitation cannot jointly learn the routing. RL recovers it. Image queries are the tell: SFT on the full DSL gets 0.323 Hit@1; GRPO gets 0.716, a 39.3-point jump. The harder action space is exactly why you need reward, not more demonstrations.

Out-of-distribution drop is at most 3 points. End-to-end latency is about 50 ms, and about 80% of that is vector retrieval, not DSL generation. The program is cheap. The index is still the work. That matches every production search system I have shipped: query planning is not where the milliseconds go. The index is.

What this is not

It is not a serving stack. Put that in the body, not a footnote.

The DSL is schema-bound. No schema, no plan. Nesting in the training distribution stops at two levels. The corpora are 3,000 products and 5,000 emails, not millions. Each query is a single turn — the model does not look at intermediate results and rewrite. There is no production fallback when generation fails. Execution success after training is over 99%, and the residual errors are mostly unescaped apostrophes in brand names, which is a tokenizer problem, not a ranking problem. Still: if the program does not compile, you need a cascade, and they did not ship one.

Those limits are load-bearing. A 4B orchestrator on a known schema is a different object from “the model will figure out your warehouse.” Do not install Hit@1 on a few-thousand-doc, single-turn, schema-bound benchmark as a production SLO. Do not claim the model invents joins you never trained. Do not skip the fallback because execution success looked high in the paper.

What to do with it

If you already run hybrid retrieval, three implications follow from the EAHR post and this one together.

Name the algebra. EAHR’s point was that truncated RRF is a different function than complete-list RRF. ProRetrieval’s point is that even complete-list RRF cannot express the query. If the user said “Sony or Bose, under $200, not refurbished, about noise cancelling,” fusion over two lists is the wrong type. SQL (or a filter language with disjunction, negation, and grouping) is the type. OpenSearch and Solr already have it. Use it. The model’s job is then to author a plan in an algebra you already trust, not to invent a new fusion weight that approximates the missing operators.

Give the model an action space that matches the query. Prompting GPT-5.5 to emit the DSL was not enough. Training a 4B model on executable plans was. The win is not a bigger embedding. It is letting the model do the composition instead of approximating it with weights. If your queries need filters and meaning, train (or constrain) the model to emit both — then measure whether the plan is the plan you would have written.

Do not skip the eval. Hit@1 on a schema-bound, single-turn, few-thousand-doc benchmark is not a production SLO. Measure whether the emitted plan is the plan you would have written, on your schema, with a fallback when it is not. Evals before opinions still holds when the ranker is a program. Pair that with the Monday checklist from the EAHR post: name the ranker, stop treating depth as portable, measure ranking agreement. A correct program that feeds truncated fusion is still the wrong ranking function.

Context engineering is a search problem. The last post was about returning the ranking you actually specified. This one is about specifying a ranking that can say what the user meant. Fusion over two lists was always a compromise. Sometimes it is still the right compromise. When it is not, write a program — and decide whether a human or a trained 4B model authors it.

Paper: ProRetrieval: Learning to Orchestrate Hybrid Search via Executable Program Synthesis (You, Sun, Hu, Zhou et al., 2026).