· Reading notes
Reading ConvMem: Long-Context Reasoning as a Summary Tree
Why this mattersHow a system condenses a long document determines whether the facts needed for an answer survive.
ConvMem1 answers questions over documents of up to 896k tokens with a frozen Qwen2.5-32B and no training. It reads the document in overlapping 8,000-token windows, summarizes each window with respect to the question, concatenates the summaries and repeats until the text fits in one window. The authors describe this as a convolution with the LLM as the kernel. Their main comparison is MemAgent,2 which reads the same document one chunk at a time and rewrites a fixed-size memory after each chunk.
With 8,000-token chunks, a fact in the first chunk of a million-token document goes through about 125 memory rewrites in a sequential agent before the answer is written. In ConvMem it goes through 3 or 4 summarization layers. That is the "linear chain to logarithmic tree" from the abstract. Below are notes on the method and results, followed by ideas for making the path shorter than logarithmic.
Method
The paper's notation: kernel size (tokens per window), stride , depth , and channels, one per sub-question. The experiments use and , so every token is read by 5 windows. All calls run at temperature 0.7.
Before reading, the model splits the question into sub-questions, usually 2 to 4. Each sub-question gets a channel and scans the whole document on its own.
At layer 0, each window gets a relevance score and a bullet-list summary for the sub-question. Windows scored 2 ("definitely contains the answer") are copied verbatim into a buffer that goes directly to the answer step. The paper calls this a skip connection.
In the upper layers, the previous layer's summaries are concatenated, cut into overlapping windows and summarized again until the concatenation is shorter than . The appendix reports 2 to 3 layers for up to 128k tokens and 3 to 4 layers for up to 1M.
Each channel then answers its sub-question from its top-level summary and its buffer, and one last call combines the sub-answers into the final answer.
Within a layer, all windows and channels are independent. With enough parallel capacity, latency grows with the number of layers, not with document length.
Results
The benchmarks are RULER-HotpotQA,3 which MemAgent was trained on, and RULER-2WikiMultiHopQA, which the authors built for this paper with the same recipe. F1 at the shortest and longest context:
| Method | HotpotQA 28k | HotpotQA 896k | 2Wiki 28k | 2Wiki 896k |
|---|---|---|---|---|
| Qwen2.5-32B, full context | 61.9 | 16.8 | 56.5 | 20.9 |
| MemAgent without RL | 64.0 | 58.7 | 60.5 | 49.7 |
| MemAgent (RL) | 75.6 | 68.8 | 60.9 | 58.4 |
| ConvMem | 67.4 | 63.1 | 72.3 | 59.1 |
ConvMem beats MemAgent without RL at every length on 2Wiki and at 5 of 6 lengths on HotpotQA (not at 112k), by 4 to 9 points at 896k. It is below RL-trained MemAgent at every length on HotpotQA. On 2Wiki it is 11 points ahead at 28k and within one point at 896k.
The paper reads the 2Wiki numbers as MemAgent overfitting its training data. The direct evidence is one example: the dataset label contains a typo ("Levni Yilmaz" for "Lev Yilmaz"), and MemAgent outputs the typo while the context says "Lev". Measured as gain over the base model at 28k, MemAgent gains 13.7 points on HotpotQA and 4.4 on 2Wiki, and ConvMem gains 5.5 and 15.8. At 896k the two are tied on 2Wiki. The gap is at short contexts, and the two datasets also differ in difficulty for every method.
Other things I noticed:
- Most Sub-EM scores are multiples of 1/128, so each context length has 128 questions. At 70% accuracy that is a standard error of about 4 points, which covers several of the differences in the table. A few cells (82.47, 73.31) are not multiples of 1/128.
- No variance over runs is reported, although every call samples at temperature 0.7.
- Kernel size, stride and number of channels were ablated on RULER-2WikiMultiHopQA, the same set used for the out-of-distribution claim.
- The text caps sub-questions at 5, the hyperparameter table at 10.
- RULER-HotpotQA, as built for MemAgent, drops questions the model can answer without the context. The description of the new 2Wiki set does not mention that filter.
- The paper argues for logarithmic latency but reports no latency or token counts. The limitations section says total compute is higher than for sequential reading.
Cost
At 896k tokens, layer 0 has 556 windows. With 3 channels, and if scoring and summarizing are separate calls as the prompt templates suggest, that is about 3,300 calls of up to 8,000 tokens, or roughly 27 million input tokens before the upper layers. Reading the document once is 0.9 million. The 5× overlap and the channels multiply the work by about 30.
What the logarithm counts
The paper uses "reasoning path" for three different quantities:
| Quantity | Sequential memory | ConvMem |
|---|---|---|
| Sequential steps (latency with unlimited parallel calls) | ||
| Lossy rewrites between a fact and the answer | up to | , or 2 through the skip buffer |
| Calls at layer 0 |
The depth depends on how much each layer shrinks the text. If a window of tokens yields a summary of tokens, the next layer is tokens long, and the number of layers is
The reported 3 to 4 layers at 1M tokens fit summaries of a few hundred tokens per window. Because of the overlap, the base of the logarithm is instead of , five times smaller.
Getting below logarithmic
Assume each lossy LLM rewrite keeps a needed fact with probability , independently. A path of rewrites keeps it with probability .
| Path | Rewrites | ||
|---|---|---|---|
| Sequential, fact in the first of 125 chunks | 125 | 0.002 | 0.28 |
| Sequential, fact at a random position | 1 to 125 | 0.15 | 0.57 |
| ConvMem, | 5 | 0.77 | 0.95 |
| ConvMem, | 6 | 0.74 | 0.94 |
| Constant path | 2 | 0.90 | 0.98 |
Most of the accuracy gain comes from going from linear to logarithmic. Going from 5 rewrites to 2 adds 13 points at and 3 points at . Below logarithmic, the larger savings are in cost and latency.
A tree in which every token passes through question-conditioned LLM calls, each reading at most inputs, has depth at least . There are three ways around that bound: send less of the text through the model, merge without the model, or let grow with . Each idea below uses one of them.
1. Filter, then build the tree. Layer 0 already scores every window. Dropping windows scored 0 and building the tree only over the rest makes the depth depend on the amount of relevant text instead of . In RULER-style data the relevant documents are a few paragraphs among hundreds of distractors, so the kept text would often fit in one window, and the path would be constant: score, answer, combine. The scoring step is also the natural place for a smaller model.
2. Overlap only at layer 0. Overlap protects facts that are cut at a window boundary. Above layer 0 the input is bullet lists, which can be split between bullets without cutting a fact. Setting above layer 0 raises the base from to . With , the upper layers then shrink the text 25-fold instead of 5-fold, and a 1M-token document needs 3 layers instead of 4. The duplicate bullets that overlapping layer-0 windows produce can be removed by string or embedding match instead of another model call.
3. Records and an exact merge. The kernel could emit records (entity, attribute, value, source offset) instead of prose. Merging two record sets is a union keyed on entity and attribute. It is associative and loses nothing, so the upper layers can be ordinary code. The model then touches each fact twice, once to extract it and once to answer. Conflicting values stay next to each other with their source offsets instead of being resolved inside an intermediate summary. This is ConvMem as MapReduce, with models only in the map and final steps.
4. One round per hop. ConvMem decomposes the question before reading. For "What government position was held by the woman who portrayed Corliss Archer in the film Kiss and Tell?", the channel for the government position does not know the woman is Shirley Temple, so it has to keep every government position it finds. With one parallel scan per hop, the first round finds the actress and the second scans for her. The number of sequential steps equals the number of hops in the question, 2 to 4 in these datasets, regardless of document length. The second round only needs the windows that mention the bridge entity.
5. Index once per document. Every ConvMem layer is conditioned on the question, so nothing carries over to the next question about the same document. A question-independent pass, either the records from idea 3 or a summary tree as in RAPTOR,4 costs one pass per document. Each question after that is a lookup and one or two model calls.
6. Early exit. If every channel has a window scored 2 at layer 0, and the overlapping windows extract the same answer, answer from the buffer and skip the upper layers. For needle-style questions this should be the common case.
7. Larger fan-in. The depth is a logarithm with base . A model that reads 128k tokens reliably could merge about 400 summaries per call instead of 25. This moves the problem back into long-context attention, which ConvMem was built to avoid.
Ideas 1, 2 and 6 only change prompts and scheduling. Ideas 3 to 5 change what the memory is.
Sources
- Zhang, H., Gu, Z., Bai, F., et al. "ConvMem: Convolutional Memory for Long-Context Reasoning." 2026. arXiv:2609.10441
- Yu, H., Chen, T., Feng, J., et al. "MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent." 2025. arXiv:2507.02259
- Hsieh, C.-P., et al. "RULER: What's the Real Context Size of Your Long-Context Language Models?" 2024. arXiv:2404.06654
- Sarthi, P., et al. "RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval." ICLR 2024. arXiv:2401.18059
Use your identity on another device
Save an encrypted identity file and restore it in another browser. Keep the file and its passphrase private.