Skip to content

Blog

Stop summarizing the past

The default move in long-horizon agents is to shrink the past until it fits. Two papers say that is the bug. One keeps a playbook. The other keeps an index.

Long-horizon agents run into a context problem that looks like an engineering inconvenience and behaves like a product failure. The trajectory grows. The window does not. Something has to give.

The default move is to shrink the past until it fits. Summarize the trajectory. Rewrite the prompt. Compact the window. Drop the “unimportant” turns. Two papers say that is the bug. One keeps a playbook. The other keeps an index.

Compression and rewrite decide what to keep before you know what you’ll need. That is fine for a meeting note. It is a bad contract for an agent that will face a new subtask tomorrow with a different information need. ACE refuses that at the prompt. Scroll refuses it at the log. Both are search instincts wearing agent clothes.

The long-horizon context problem

If you have shipped production search, you already know the failure mode. Write-time digests are attractive: smaller indexes, cleaner snippets, lower serving cost. They also throw away the evidence you did not know you would need. Query-time projection keeps the record and builds the view when the query arrives. Serving systems treat indexes that way. Result windows are views. The index is the record.

Agents mostly do the opposite. They rewrite the past into a shorter past. Prompt optimizers chase brevity. Memory modules store summaries. Compaction jobs run when the window fills. The agent then reasons over a digest that cannot be expanded back into the events that produced it. When the next task needs a detail the digest dropped, there is nowhere to look.

That is the setup ACE and Scroll are answering — from different angles, with different artifacts.

ACE: the playbook that does not collapse

Agentic Context Engineering (Zhang, Hu, Upasani et al.; Stanford, SambaNova, Berkeley; October 2025) is still the 2026 baseline later work positions itself against. Prompt optimizers such as GEPA and MIPROv2 suffer brevity bias: they drop domain heuristics for a shorter instruction. Full rewrites suffer context collapse. On AppWorld, an 18,282-token context at 66.7% accuracy became a 122-token rewrite at 57.1 — worse than the unadapted baseline. Shorter was not smarter. Shorter erased the strategies that made the agent work.

ACE treats context as an itemized playbook. Each bullet is a strategy, a failure mode, a tool heuristic. A Generator produces trajectories. A Reflector distills lessons. A Curator applies incremental deltas. Grow-and-refine de-duplicates. No monolithic rewrite. Offline (system prompt) and online (agent memory). Execution feedback can substitute for labels.

On AppWorld with DeepSeek-V3.1, ReAct+ACE averages 59.4 against ReAct’s 42.4. Online it hits 59.5, matching IBM CUGA on GPT-4.1 (60.3) with a smaller open model, and beating it on the harder split. Finance (FiNER/Formula): +12.8 offline with labels. Adaptation latency is 82.3% lower than GEPA. 91.8% of evaluation input tokens are cacheable — which matters if you are paying for context on every turn.

The caveat belongs in the body, not a footnote. Without reliable feedback, ACE and Dynamic Cheatsheet can degrade — spurious lessons pollute the playbook. That is evals-before-opinions in paper form. Not every task wants a rich playbook either: HotpotQA and Game of 24 often want one rule, not a growing manual. A playbook is curated memory. Curated memory still needs a judgment signal. Garbage in, policy out.

Your agent does not need a shorter prompt. It needs a versioned playbook that does not collapse when the LLM rewrites it.

Scroll: the log you can query

Scroll (Lin, Ang, Zhu, Ding, Zhou; Alibaba / Columbia; August 2026) keeps an append-only Event Log — SQLite, BM25 by default, not embeddings — plus a persistent Python kernel. The model execs to search, expand, and compute. Only print() enters the next context. Eviction changes the view, not the record. An eviction index maps headlines to Event Log sequence addresses so the agent can navigate what left the window without searching the whole log.

That architecture should feel familiar if you have lived in Solr or OpenSearch. BM25 over a lossless log. Query-time projection instead of a write-time digest. The agent does not get a summary of what happened last Tuesday. It gets tools to find Tuesday when Tuesday matters.

With Qwen3.8-Max they report 94.8% on LongMemEvalS; 73.1 on BEAM10M (they cite Exabase M-1 at 68.0 as the best published); 86.7 on LOCA 256K against 49.3 for the best published ReAct they cite. The BEAM ablation is the argument: lossy summarization at ingestion falls to 19.9 overall. Drop the REPL, −7.3. Drop the eviction index, −1.8. Median input on BEAM10M is about 105K tokens, roughly 1% of the corpus. The model is not stuffing the past. It is retrieving from it.

Two caveats. BEAM is not a fully controlled bake-off — different backbones, different judges. And weaker models can use the interface and still fail the long trajectory: Qwen3.6-35B-A3B scores 22.7 at LOCA 256K against 86.7 for Qwen3.8-Max. The hands are there. The planning is not. Giving a small model BM25 and a REPL does not make it a long-horizon agent if it cannot decide what to ask for.

Stop summarizing the agent’s past. Index it, and let the model write the query. Serving systems already treat indexes this way. Result windows are views. The index is the record.

A playbook and a log

A playbook is curated memory. A log is retrievable memory. Production search already thinks in both: you keep the inverted index, and you keep the runbook of what broke last time. Context management is retrieval over an event-sourced log, not prompt-tuning.

The retrieval interface is the product. ACE’s interface is incremental bullets with a curator. Scroll’s is search / expand / print over BM25. If the next task needs SQL and vectors in one action space, that is a program, not a summary — ProRetrieval’s point that fusion over two frozen lists cannot express real constraints. Different problems, same through-line: do not collapse what you will need to query later.

What not to claim

Do not sell “just add memory.” ACE without evals gets worse — spurious lessons stick. Scroll’s headline numbers are not a clean bake-off across identical backbones and judges. A small model can hold the tools and still miss the long hop. Lossy summarization’s 19.9 on BEAM is a warning about ingestion-time digests, not a proof that every summary is poison. Some tasks still want one rule. Some products still need compaction for cost. The claim is narrower: if you do not know what the next task will need, do not summarize it away. Write it down in a form you can search.

If you don’t know what the next task will need, don’t summarize it away. Write it down in a form you can search.

Papers: ACE (Zhang, Hu, Upasani et al., 2025). Scroll (Lin, Ang, Zhu, Ding, Zhou, 2026).