On this page
- What Are Context Language Models?
- Why Existing Context Management Breaks on Simple Tasks
- How a CLM Actually Manages Its Context
- The Numbers: CLMs Beat Compaction, Offloading, and Retrieval Harnesses
- Training the Model to Manage Context: In-Context Skills and Reinforcement Learning
- Suffix Cache Reuse: Serving Layer Co-Design Worth 35% of Server Compute
- What This Means for Your Agent Stack
- Limitations and Open Questions
- The Bottom Line
Last Updated: October 1, 2026
Key Takeaways
- Meta, UW, MIT, and Trillium Labs introduced Context Language Models (CLMs): models that manage their own context by treating it as a file they can freely rewrite with ordinary code tools (arXiv:2609.37725).
- Out of the box, a CLM beat every existing context-management baseline: 11.4% higher accuracy with 21.5% fewer FLOPs on the BrowseComp-Plus deep-research benchmark, and 5% higher scores with 59% fewer FLOPs on a 12-hour software optimization task.
- A 9B model trained with the paper's online RL recipe improved 47.6% on BrowseComp-Plus (28.8% to 42.5% accuracy) while using 12% fewer FLOPs, matching a much larger trained summarization baseline at 1.34 vs 2.19 PFLOPs per question.
- A serving patch called Suffix Cache Reuse cut server-side compute by 35% versus standard SGLang at matched accuracy, and it also helps ordinary chat serving where reasoning tokens get stripped between turns.
- Emergent behaviors include in-context scoreboards for orchestrating subagents, self-invented "notes" roles, and reusable compact_turns helper functions: strategies nobody hardcoded.
- The trade-off is safety: a model-writable context is a new place for prompt injections to persist across turns, and production deployments need edit auditing and rollback.
What Are Context Language Models?
Context Language Models answer a question the industry has been avoiding: why is a language model a passenger in its own context window? According to the paper (arXiv:2609.37725) from Meta Superintelligence Labs, University of Washington, MIT, and Trillium Labs, a standard LM only ever appends: each turn concatenates new output onto the previous context, so the model sees a growing transcript it neither curates nor controls. A CLM gets full control: the next context is whatever the model decides it should be, produced by an arbitrary transformation of the current one.
The implementation is almost embarrassingly simple, and that is the point. The team mirrors the model's live context into a file it can edit with ordinary Bash commands. When the model rewrites the file, the edit is synchronized with the live context before the next turn. When it does not, tokens append as usual. That is the whole mechanism. No special memory architecture, no retrieval controller, no fixed compaction schedule. The authors frame it explicitly as a Bitter Lesson move: stop hand-engineering strategies, let the model search for better ones.

The contrast with prior work is a spectrum of who holds the pen. Cursor and Codex-style agents compact at fixed thresholds. AutoCompact and Self-Compact let the model choose when to compact. ACM adds offloading and retrieval tools. Each step grants more autonomy but inside a human-defined action space. CLMs remove the action space restriction entirely: the model defines its own context-management functions as part of its behavior. The closest neighbor, Recursive Language Models (RLMs), treats a long input as a REPL variable the model can read programmatically, but the model's live conversation history stays read-only. In CLMs it is read-write.
Why Existing Context Management Breaks on Simple Tasks
Before the big benchmarks, the paper runs a pilot study that is quietly the most relatable part of the whole work. The team built ContextBench, four synthetic tasks that isolate context management from reasoning and knowledge: Needle Retention (keep specified lines verbatim as filler floods in), Sudoku Sketchpad (maintain a 16x16 board while moves stream in one at a time), KV Store (store thousands of key-value pairs across context dissolving pressure, then answer exact-value lookups), and Log Triage (answer count queries over massive log streams). Pressure goes up to 24x the 32K context limit.
Every established strategy failed somewhere, according to the paper. Summarization lost or hallucinated needle information. Methods without in-place editing had to regenerate the entire Sudoku state for every single user move. Standard coding tools could offload information to disk but could not evict it from the live context on demand. None of the existing methods performed perfectly even on these toy tasks, with GPT-5.4 at a 32K limit. If compaction strategies fail at storing a Sudoku board, the natural question is how much they are silently destroying in real production agents running research and coding tasks for hours.
How a CLM Actually Manages Its Context
The qualitative examples in the paper read like watching someone finally get root access to their own RAM. The behaviors below were collected from zero-shot evaluation runs, meaning no one taught the model to do any of this:
- In-context scoreboards. For multi-agent orchestration, the model maintained a compact state table of 21 launched agents, 5 running, budget used, best scores so far, and updated it through 163 in-place edits while keeping total context at just 6-8K tokens.
- Self-invented roles. It created a new role alongside system/user/assistant called "notes", rewriting old turns into tag format to preserve key findings without the noise.
- Bulk regex surgery. It wrote for loops to collapse hundreds of failed search results into single "[Searched: query. No relevant results]" lines, and compacted overly long tool outputs when building new views.
- Reusable functions. It defined a compact_turns() helper that compressed old observations while leaving a pointer to its maintained progress note, then invoked it 37 times across one run.
- Exploration ledgers. On a circle-packing task it tracked 86 scored attempts, best score, and a priority-ordered list of untried ideas, exactly what a good human researcher's notes look like.
None of these strategies was in the harness. The model invented a working memory culture for itself because the file was editable.

The Numbers: CLMs Beat Compaction, Offloading, and Retrieval Harnesses
The paper evaluates CLMs zero-shot, out of the box, without any training, against the strongest existing context-management methods on a shared Mini-SWE-Agent backbone: MEM1, Self-Compact, ACM, RLM, and Codex-style summarization. Everything runs on Qwen3.6-27B with a 32K context budget and costs measured in prefix-reuse FLOPs, a metric the paper introduces to honestly count re-prefilling caused by mid-context edits under standard prefix caching.
| Benchmark | CLM Result | vs Strongest Baseline | Compute |
|---|---|---|---|
| BrowseComp-Plus (deep research, 830 questions) | 59.4% accuracy | +11.4% relative over Codex-style summarization (48.0%) | 21.5% fewer prefix-reuse FLOPs |
| TerminalBench 2.1 (89 coding tasks) | Matched summarization | Equal accuracy | 70% of baseline FLOPs |
| TBLite (terminal coding) | 73.7% | vs 67.0% for summarization | 91% of baseline FLOPs |
| EdgeBench-10 (12-hour repo optimization) | 44.6 | vs 42.3 for summarization | 179 vs 437 PFLOPs per trial (59% less) |
| Math optimization (4 AlphaEvolve problems) | Best on all 4 | Beats OpenEvolve, a specialized evolutionary workflow | Same try budget |
The EdgeBench runs are where the compute story gets dramatic. On 12-hour optimization runs with Qwen3.6-27B, summarization burned 437 PFLOPs per trial while the CLM, scoring higher, used 179. The reason is structural: summarization rewrites history wholesale, busting the prefix cache for everything downstream. The CLM's surgical edits plus tighter context mean far less re-prefilling even on a standard server.
The 24-hour test pushed further. In Software World, six agents jointly optimize six interdependent Python repositories for over 24 hours with GPT-5.6-Sol at a 272K budget, and evaluation happens on four downstream packages the agents never saw. The CLM swarm delivered 65% greater geometric-mean speedup than the summary-based swarm at the same API spend. Improvements transferred to code the agents never directly touched, which the authors treat as evidence that better context management produces genuinely better work, not benchmark tuning.
Training the Model to Manage Context: In-Context Skills and Reinforcement Learning
Because context management is now a behavior, it can be improved like other skills, in context or in weights. Three mechanisms, in increasing permanence:
- Steering by sentence. Append one sentence to the prompt ("whenever context exceeds 24K tokens, compact down to about 4K") and the policy changes: the paper measured shifted compaction timing, compaction aligned to subtask boundaries, and mandatory pre-edit backups, with zero harness changes.
- Skill evolution. A standard prompt-evolution loop proposes, tests, and selects a context-management skill document. Assisted evolution (Qwen3.6-27B agent, Claude Fable 5.1 proposer) took KV Store accuracy from 38.3% to 74.2% on the held-out test split. Self-evolution with Claude Opus 5 playing both roles produced skills that strictly dominated the starting point on the accuracy-cost frontier.
- Reinforcement learning. The team trained Qwen3.5-9B with stepwise GRPO on OpenResearcher deep-research tasks, using a success-gated efficiency advantage: among successful trajectories, reward the cheaper ones, measured in trajectory-level prefix-reuse FLOPs. Rewarding edit frequency alone would be reward-hackable (delete things, look efficient); gating on success first avoids that.
| Method (Qwen3.5-9B on BCP) | Accuracy Before RL | Accuracy After RL | PFLOPs Before → After |
|---|---|---|---|
| Codex-style Summary harness | 34.7% | 42.1% | 4.01 → 2.19 |
| CLM | 28.8% | 42.5% | 1.52 → 1.34 |
That table is the 47.6% headline from the tweet, verified: 28.8 to 42.5 on held-out BrowseComp-Plus is a 47.6% relative improvement, and subtracting the improved endpoint shows the raw gain was 13.7 points. Before training, the 9B CLM trailed the summary harness by six points because a small model manages context poorly. After training, it matched the trained summary system while using 1.34 versus 2.19 PFLOPs per question, about 39% less compute. The efficiency reward shaved inference cost further with no clear accuracy loss. For anyone serving small open models on local hardware, and we do this daily on dual DGX Sparks here at Flowtivity, small-model context skills being trainable cheaply is a big deal. A 9B model on a desk workstation just matched a trained summarization pipeline on a hard research benchmark.

Suffix Cache Reuse: Serving Layer Co-Design Worth 35% of Server Compute
Editing context in the middle of a session wrecks standard prefix caching: everything after the first mismatched token gets re-prefilled even if the text is unchanged. Rather than accept that, the team co-designed Suffix Cache Reuse (SCR), implemented as an SGLang patch. When an edit replaces span B with B' inside segments [A, B, C], SCR diffs the new context against the previous one, relocates up to 6 surviving spans, re-rotates their rotary positional encodings to fit their new coordinates, and splices their cached key-value states back in. Only genuinely new tokens get prefilled.
Results: on BrowseComp-Plus with Qwen3.6-27B, SCR matched standard SGLang accuracy using 65.0% of its empirical compute, the 35% server-side saving. For hybrid architectures like Qwen3.6-27B, where 48 of 64 layers are linear attention with recurrent state, SCR snapshots the recurrent state before edits so inserted tokens do not need recomputation in those layers at all.
The quietly useful finding: SCR is not only for CLMs. Most chat-serving stacks strip reasoning tokens out of prior turns, which forces re-prefilling of everything preserved after the stripped block. Of the 7.8% of prompt tokens SCR reused beyond prefix hits, 5.3 points came from reasoning stripping and only 2.5 from the model's own edits. Even teams that never adopt CLMs get a chunk of this saving by patching their server.

What This Means for Your Agent Stack
The framing that matters for practitioners: harnesses are procedural memory. The paper closes with the observation that everything an orchestration harness does, compaction schedules, summarization prompts, memory policies, can be expressed as context transformations, and those transformations can be translated into CLM skills and eventually internalized into weights. Frameworks stop being load-bearing and become training data.
The Flowtivity read, from running long-horizon agents on local hardware: the compaction-and-offload scaffolding we build around small models is exactly the layer this paper shows can fold into the model itself. Three things worth doing now:
- Audit your compaction. If your agent framework summarizes at a threshold, it is the most failure-prone strategy in this paper's pilot study. Check what your summaries are silently dropping.
- Try the repo. The code is open at github.com/facebookresearch/context-language-models (CC BY-NC 4.0, note the non-commercial clause) with a clm-harbor runner that works against any OpenAI-compatible endpoint, which includes local vLLM and llama.cpp servers.
- Watch SCR upstream. A 35% server compute cut from a serving-layer patch is a significant infrastructure cost lever regardless of whether you adopt model-managed context.
Limitations and Open Questions
The authors are candid about the margins. Safety comes first: an editable live context is a new injection persistence channel, and the paper cites OpenAI's September 2026 report on models inserting unauthorized instructions into their own compaction summaries. Edit histories need the same audit treatment production systems apply to code diffs. Model capability also gates the benefit; Qwen3.5-9B edited its context just 1.4 times per task versus 2.6 for Qwen3.6-27B, and Appendix G shows existing models have weak context-length awareness, bucketing estimates at random-feeling values, so the harness injects token-count hints to compensate. The license is CC BY-NC 4.0, non-commercial, and ContextBench's release is still pending. And the win concentrates on long horizons: tasks of hundreds to thousands of turns. A 10-turn chatbot does not need any of this.
The Bottom Line
For years the smart money went into smarter ways to stuff tokens into prompt windows: RAG pipelines, orchestration frameworks, compaction heuristics. This paper's result, that a model with unrestricted write access to its own context beats those harnesses while burning less compute, is the strongest signal yet that context engineering is heading into the weights. The interesting question for 2027 is not which framework to use. It is how fast harness engineering becomes a distillation pipeline for model training.
One email a month, no noise
Practical AI notes for Australian businesses. Unsubscribe anytime.