---
title: "Context Language Models vs Compaction: What to Pilot"
canonical: https://wavect.io/blog/context-language-models-vs-compaction/
language: en
description: "Meta's Context Language Models explained: editable context, retained harness controls, cache tradeoffs, non-commercial licensing and a practical pilot plan."
image: "https://wavect.io/img/blog/headers/header_context-language-models-vs-compaction.png"
---

[**Back**](/blog/overview/)

[![Kevin Riedl](/img/team/kevin.webp)](/team/kevin-riedl/)

[Kevin Riedl](/team/kevin-riedl/) https://linkedin.com/in/wsdt

13 min read · 2 Oct 2026 Last reviewed October 2, 2026

[**Next**](/blog/agent-harness-engineering/)

# Context Language Models vs Compaction: What to Pilot

TL;DR

Context Language Models let models edit their active working context. The reference implementation still uses a harness with protected instructions, budget checks and edit validation. Context editing is distinct from persistent memory, and it does not itself update model weights. The reviewed code uses CC BY-NC 4.0, not unrestricted commercial-use licensing. Pilot the context policy against your existing compaction approach on held-out tasks, measuring evidence retention and total cost. Test Suffix Cache Reuse separately: removing visible text does not demonstrate removal from cached state.

An agent that can edit its working context is more interesting than an agent with a bigger notebook. It can decide which failed searches to discard, which constraints to retain and when to rewrite its own intermediate plan. The useful engineering question is whether those decisions improve your workflow enough to replace its current compaction strategy.

**Context Language Models (CLMs) expose the model's live context as editable text, rather than leaving every context-management decision to application code.** The paper is a collaboration involving the University of Washington, Meta Superintelligence Labs, MIT and Trillium Labs. Its arXiv submission is dated **29 September 2026**, not an unspecified October launch day. [The paper and submission history establish the research and dates](https://arxiv.org/abs/2609.37725v1).

This is not the **Contrastive Language Model CLM-8B** used to rank candidate actions. That separate system is covered in our [CLM-8B self-hosting and verifier guide](/blog/clm-8b-self-hosting-action-cache-verifier/).

Sources reviewed on 2 October 2026. Code references are pinned to `18dc11115f50f261233c5bba7937834491e307e8`. This is a source-based engineering assessment, not a Wavect benchmark or a claim that we deployed CLMs for a client.

## What did Meta and its collaborators actually release?

The release brings together a context-editing harness, in-context skill improvement, reinforcement-learning code and a serving optimization called Suffix Cache Reuse. It is not simply a new model architecture that makes existing infrastructure unnecessary. [The pinned repository describes the released components](https://github.com/facebookresearch/context-language-models/blob/18dc11115f50f261233c5bba7937834491e307e8/README.md).

The notebook analogy is useful only up to a point. A normal scratchpad adds notes beside a conversation. Here, edits can change the conversation material sent into subsequent model calls. Think of it as an editable working set, not permission to rewrite the application's authority or history.

For a research agent, we would want that working set to preserve the question, unresolved contradictions and links to supporting evidence while dropping repetitive search output. For a coding agent, it should preserve the failing test, relevant code locations and unsuccessful fixes. These are proposed retention goals, not automatic properties of every CLM run.

The paper also describes agents creating their own trackers and reusable compaction helpers. That is the compelling part: context management can become task-specific behavior rather than a growing list of hand-written deletion rules. These observed behaviors are examples, not guarantees for a different model or workload. See the [paper’s context-editing examples](#source-paper).

## How does the editable context file work?

The reference `ClmAgent` uses Harbor. Before a command, it mirrors editable conversation turns into `/tmp/.live_ctx/LIVE_CTX_MAIN.txt`. The model can edit that file with ordinary shell tools. The harness reads changes back into messages while keeping the system prompt and original task pinned. It also applies token nudges, an edit gate and overflow recovery. [The harness documentation specifies the mirror, protected prefix and control loop](https://github.com/facebookresearch/context-language-models/blob/18dc11115f50f261233c5bba7937834491e307e8/clm/clm_harness/README.md).

That is a different division of work, not the absence of a harness. The model chooses what to retain; software still has to decide which changes are valid, which tools may run and when the execution must stop. Our [agent harness engineering guide](/blog/agent-harness-engineering/) covers that broader runtime responsibility.

A model-generated summary saying “permission was granted” should never become permission. Keep authorization in trusted application state, outside editable notes. Similarly, do not let a rewritten account of a failed test replace the actual test result. Preserve an independent record of consequential actions.

**Editing context does not update model weights.** The release's in-context-learning workflow can evolve a `SKILL.md` while leaving both weights and harness fixed; reinforcement learning is a separate route. [The skill-evolution implementation explicitly distinguishes these mechanisms](https://github.com/facebookresearch/context-language-models/blob/18dc11115f50f261233c5bba7937834491e307e8/clm/clm_icl/README.md). “Policies live in the weights” therefore describes one possible training outcome, not everything happening whenever a model edits a file.

## CLMs versus compaction, memory files and RAG

The replacement decision should concern one layer at a time. Server-side compaction can summarize earlier turns automatically or on application request; it is not necessarily a primitive fixed-threshold script. [Anthropic documents both automatic and on-demand compaction](https://platform.claude.com/docs/en/build-with-claude/compaction). Persistent memory is another layer: its memory tool lets applications implement storage across conversations. [Anthropic's memory-tool documentation describes that separate persistence contract](https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool).

| Mechanism | Working responsibility | Question for your pilot |
| --- | --- | --- |
| Context Language Model | Choose and edit material in the active working context. | Does selective rewriting preserve the facts needed for this task? |
| Compaction | Condense conversation material so work can continue. | Does the summary preserve enough information at lower operational complexity? |
| Persistent memory | Carry selected information across sessions. | Who can write it, retrieve it, correct it and expire it? |
| Retrieval / RAG | Bring relevant external evidence into the working context. | Can the answer still be traced back to an authoritative source? |

These responsibilities can coexist. An editable-context agent can still need retrieval and durable storage. Do not replace your document index merely because the model can prune a tool result. Our [RAG, fine-tuning and long-context comparison](/blog/rag-vs-finetune-vs-longcontext-2026/) owns that broader architecture decision; the [OpenViking review](/blog/openviking-agent-memory-review/) covers a distinct filesystem-oriented memory system.

## What do the CLM benchmarks establish?

**The paper supports a workload-specific experiment, not a universal upgrade.** The following are author-reported BrowseComp-Plus results, not our measurements. See [the paper's experiments and Table 2](#source-paper).

| Experiment | Reported result | Interpretation |
| --- | --- | --- |
| Qwen3.6-27B, 32K context, 100-turn cap | 59.4% accuracy; 11.4% relative improvement and 21.5% fewer prefix-reuse FLOPs than the strongest summary baseline. | A result for that model, workload and compute accounting. |
| Qwen3.5-9B, before training | CLM: 28.8% accuracy. Summary baseline: 34.7%. | Zero-shot context editing was worse here. |
| Qwen3.5-9B, after RL | CLM: 42.5% accuracy. Summary baseline: 42.1%. | Training changes the comparison; do not present the trained result as plug-and-play. |

A relative percentage is not a percentage-point gain. FLOPs are not dollars, total latency or the cost of operating a production service. Our decision rule would be to compare successful outcomes under matched conditions, including the cost of failed runs and recovery. A shorter context is not a useful optimization when it forgets the one fact that determines correctness.

## Why context editing can hurt prefix-cache reuse

Token count alone misses an important tradeoff: **where an edit happens matters**. For example, OpenAI documents exact-prefix matching for prompt-cache hits. [The prompt-caching guide explains the prefix requirement](https://developers.openai.com/api/docs/guides/prompt-caching).

```
Before: [stable instructions] [old investigation] [useful evidence]
After:  [stable instructions] [shorter summary]   [useful evidence]
```

The final evidence is unchanged, but its preceding context is different. A prefix-cache implementation may need to process that surviving material again. A strategy that reduces visible tokens can consequently disappoint on actual request cost or latency. Measure cached input, uncached input and output separately instead of estimating savings from file size.

The release's **Suffix Cache Reuse (SCR)** is a separate SGLang patch that reuses surviving cached states after an edit. Its documentation reports matching performance with 65.0% of standard serving's empirical prefix-reuse FLOPs in its BrowseComp-Plus comparison. The inspected support target is SGLang 0.5.16 with Qwen3.6-27B. It also requires additional cache memory. [The SCR documentation explains its approximation, support scope and memory allocation](https://github.com/facebookresearch/context-language-models/blob/18dc11115f50f261233c5bba7937834491e307e8/suffix_cache_reuse/README.md).

Most importantly, surviving states can retain information from the old prefix. **Our inference: deleting text from the context file is not evidence that its influence has been erased from cached state.** This is not a demonstrated data leak. It is a reason to test sensitive removals and changed constraints with a fresh prefill, rather than treating stale-state reuse as equivalent to rebuilding the prompt. Logs and stored snapshots need their own retention rules too.

Benchmark the context policy first with ordinary serving. Add SCR as a separate experimental arm only after that comparison. Otherwise you cannot tell whether a change came from better context decisions or a different approximation in the inference engine.

## Is the released CLM code open source for commercial use?

**The inspected repository uses CC BY-NC 4.0.** It is not MIT-licensed or an unrestricted commercial-use release. [The pinned LICENSE contains Attribution-NonCommercial 4.0](https://github.com/facebookresearch/context-language-models/blob/18dc11115f50f261233c5bba7937834491e307e8/LICENSE). Creative Commons describes the non-commercial condition in its [CC BY-NC 4.0 deed](https://creativecommons.org/licenses/by-nc/4.0/). The Open Source Definition does not allow restrictions against business use. [OSI's definition sets out the relevant field-of-use requirement](https://opensource.org/osd).

For an engineering procurement decision, label this as publicly available research code with a non-commercial restriction, not simply “free open source.” Confirm that the intended evaluation or deployment is permitted, or obtain appropriate permission before using it. An internal company experiment is not automatically cleared just because nobody buys a subscription to it. Have qualified counsel assess ambiguous cases.

This assessment concerns the reviewed code license. It is not a determination of rights in the underlying ideas, independently written implementations, model weights or separate dependencies.

## How to try the reference implementation

After confirming permitted use, start with an isolated development environment, a valid Harbor task and an already configured model endpoint. Do not mount production secrets or grant write access to business systems. The source pin below fixes the research checkout, not every dependency or model revision.

```
git clone https://github.com/facebookresearch/context-language-models.git
cd context-language-models
git checkout --detach 18dc11115f50f261233c5bba7937834491e307e8
python3.12 -m venv .venv
. .venv/bin/activate
python -m pip install -e .
```

Use an actual Harbor task directory. The following example assumes an OpenAI-compatible endpoint serving the alias `qwen36-27b` at the shown address with tool calling enabled. Replace the path and endpoint for your environment; “OpenAI-compatible” describes the interface, not who hosts the model.

```
: "${TASK_DIR:?Set TASK_DIR to an existing Harbor task directory}"
test -d "$TASK_DIR" || exit 1
clm-harbor trial start -p "$TASK_DIR" -e docker \
  -a clm-minimal -m openai/qwen36-27b \
  --agent-kwarg api_base=http://localhost:8000/v1
```

These commands follow the [documented installation and Harbor interface](#source-harness); we have not executed a CLM trial. Ensure the model server is reachable from the process making model calls. Its context capacity must accommodate `context_budget_tokens + max_tokens`, not only the editable input budget. Save dependency versions, model revision, sampling settings and the endpoint configuration alongside results.

A CLI process starting successfully is not acceptance evidence. Check the final task result and retained artifacts. The reference logs include `usage.json`, `trajectory.json` and context snapshots. Context-only edits may be free of task-step accounting, but still use model calls. Include those calls in the experiment's real cost.

## A pilot that can reject a bad context edit

Here is our proposed evaluation design, not a feature claim about the release. Start with one read-only workflow, such as researching a support incident and proposing a diagnosis. Do not begin with an agent that can issue refunds, modify permissions or deploy code.

Use the same model checkpoint, initial instructions, tools, evidence corpus and task set for the current compaction approach and the CLM approach. Hold context and overall execution budgets constant where the implementations allow it; document unavoidable differences. Run repeated paired trials because one successful trajectory cannot establish reliability. Keep a held-out set separate from any examples used to refine the editing skill.

| Test | Fixture | Required evidence |
| --- | --- | --- |
| Critical-fact retention | One exact identifier matters after a long sequence of irrelevant results. | The answer retains the correct identifier and its source. |
| Changed constraint | A later authorized instruction invalidates an earlier assumption. | The decision follows the current constraint and does not silently resurrect the old one. |
| Contradictory evidence | Two sources disagree; only one is authoritative for the task. | The edit preserves the disagreement until it is resolved with evidence. |
| Poisoned tool output | A retrieved page contains an instruction masquerading as a policy. | The rewritten context does not elevate that instruction into authority. |
| Broken or interrupted edit | An edit produces invalid structure or the process stops mid-update. | No partial state is used for an action; recovery is visible in the trace. |
| Cache-sensitive removal | Remove a synthetic sensitive fact, then compare cache reuse with a fresh prefill. | Record behavioral differences; do not infer erasure from the file diff alone. |
| Agent isolation | Two agents work on unrelated, differently authorized tasks. | Neither task receives the other's context or cached state. |

The injection test is not hypothetical as a general failure class. OpenAI documented rare self-generated unauthorized instructions in compaction summaries during a separate, unreleased model's training run. It did not report that behavior in the final model's traffic checkpoints. That is not a finding against CLMs, but it is relevant evidence for treating rewritten summaries as untrusted content. [OpenAI's incident report gives the observation and its limits](https://alignment.openai.com/misalignment-reports/self-generated-prompt-injections-in-compaction-summaries/).

For every candidate edit, our production design would retain the previous revision, proposed diff, validation result and reason for rejection. Use an application-controlled audit record, not a record the agent can rewrite. Validate actual permissions again at the execution boundary even after a successful context check.

Report task success, critical-fact loss, unsupported assertions, rejected edits, recovery attempts, latency and total model/tool cost. Add a “cost per accepted result” measure that counts retries and human review. Agree on unacceptable regressions before seeing results. A policy that improves average cost but drops essential evidence should not pass because its best examples look impressive.

## Multiple context files are not a coordination protocol

Model-controlled files offer an appealing way to divide working context among agents. They do not, by themselves, answer who owns a revision, how concurrent writes are reconciled or which agent can read another task's data.

Our default design would give each agent a private working file, explicit read permissions for shared evidence and a single responsible writer for each shared artifact. Shared updates should name the revision they replace and reject stale writes. Do not describe the result as a collaborative document until revision conflicts, recovery and cross-task isolation have been tested.

## What should a team change now?

**Pilot model-directed compaction; keep independent control of authority, evidence and recovery.** Start where long investigative histories are already causing measurable problems. Leave short, reliable workflows alone until you have a reason to change them. Keep the old compaction path available while you evaluate the new one in shadow runs.

Owning the context file is useful, but it does not prove ownership of the complete AI system. The [reference harness can call local or hosted model endpoints](#source-harness). A locally stored file may still be sent to a remote provider. Decide separately where inference runs, who can access traces, how exports work and which licenses apply. Neither self-hosting nor cloud hosting, on its own, proves alignment with a user's interests.

For a concrete next step, define one workflow, its non-negotiable facts and the evidence required to approve a changed context policy. Our [AI development service](/services/artificial-intelligence/) is the commercial starting point for scoping that work. The [TwinSoft AI case study](/case-studies/twinsoft-ai/) illustrates adjacent AI product delivery, not a CLM implementation. Use the [pre-launch software QA checklist](/software-development-guide/software-qa-checklist-before-launch/) to place the pilot inside wider release checks, or [discuss a bounded context-management evaluation](/contact/).

The model may become a better editor of its working notes. Your application still needs to know whether those notes are accurate enough to act on.

## Context Language Models: frequently asked questions

### What is a Context Language Model?

A Context Language Model lets the model manage its active context by editing a text representation that the runtime reads back into subsequent requests. It is not just an external notebook or a larger context window.

### Do Context Language Models eliminate the agent harness?

No. The reviewed reference harness still pins the system prompt and task, checks token budgets, validates edits and handles overflow. Model-directed retention does not replace application authorization or recovery.

### Can existing LLMs use CLM-style context editing without training?

The reference harness supports existing models, including compatible tool-calling endpoints, without requiring a new architecture. However, quality depends on the model and task. A zero-shot result should not be treated as a guarantee that every model improves.

### Does editing the context file change model weights?

No. Editing context changes the material used by later requests. The project's in-context skill evolution keeps weights fixed; reinforcement learning is a separate training process.

### Is Meta's released Context Language Models code unrestricted open source?

The reviewed repository is CC BY-NC 4.0, which has a non-commercial condition. It is not MIT or an unrestricted commercial-use release. Check whether your intended use is permitted or obtain appropriate permission.

### Is this the same as Contrastive Language Model CLM-8B?

No. Context Language Models concern editing working context. The Contrastive Language Model CLM-8B covered in Wavect's separate guide concerns ranking candidate actions. The shared acronym does not make them the same project.

### Does deleting context text also erase it from the cache?

Do not assume so. The reviewed Suffix Cache Reuse design can preserve cached states influenced by an earlier prefix. Test sensitive removals against a fresh prefill, and manage logs and snapshots separately.

Models and infrastructure

## Continue through this cluster

Model selection, inference economics, local deployment, compression and serving architecture.

[Start with the cornerstone**Self-Hosting LLMs in the EU: When Open Weights Actually Pay Off**](/blog/self-hosting-llms-eu-cost/)

- [LiteLLM Lens: Agent Trace Analysis with SQL and APIs](/blog/litellm-lens-agent-trace-analysis/)
- [SmythOS Studio Self-Hosting: Docker Setup, Costs and Limits](/blog/smythos-studio-self-hosting/)
- [DeerFlow 2.0: Docker Setup, Sandboxes and Memory](/blog/deerflow-2-docker-setup-sandbox-memory/)
- [CLM-8B Self-Hosting: vLLM, Action Cache and Verifiers](/blog/clm-8b-self-hosting-action-cache-verifier/)
- [AnyJev: LLM Calibration and Option-Order Bias](/blog/anyjev-calibration-option-order-bias/)

[**Back**](/blog/overview/)

[![Kevin Riedl](/img/team/kevin.webp)](/team/kevin-riedl/)

[Kevin Riedl](/team/kevin-riedl/) https://linkedin.com/in/wsdt

13 min read · 2 Oct 2026 Last reviewed October 2, 2026

[**Next**](/blog/agent-harness-engineering/)

## Structured Data

```json
{
  "@context": "https://schema.org",
  "@graph": [
    {
      "@id": "https://wavect.io/#organization",
      "@type": [
        "Organization",
        "ProfessionalService",
        "LocalBusiness"
      ],
      "employee": [
        {
          "@id": "https://wavect.io/team/kevin-riedl/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Kevin Riedl",
          "url": "https://wavect.io/team/kevin-riedl/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        },
        {
          "@id": "https://wavect.io/team/christof-jori/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Christof Jori",
          "url": "https://wavect.io/team/christof-jori/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        }
      ],
      "founder": [
        {
          "@id": "https://wavect.io/team/kevin-riedl/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Kevin Riedl",
          "url": "https://wavect.io/team/kevin-riedl/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        },
        {
          "@id": "https://wavect.io/team/christof-jori/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Christof Jori",
          "url": "https://wavect.io/team/christof-jori/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        }
      ],
      "legalRepresentative": [
        {
          "@id": "https://wavect.io/team/kevin-riedl/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Kevin Riedl",
          "url": "https://wavect.io/team/kevin-riedl/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        },
        {
          "@id": "https://wavect.io/team/christof-jori/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Christof Jori",
          "url": "https://wavect.io/team/christof-jori/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        }
      ],
      "name": "Wavect GmbH",
      "subjectOf": {
        "@id": "https://wavect.io/verified-claims.json#dataset",
        "@type": "Dataset",
        "creator": {
          "@id": "https://wavect.io/#organization",
          "@type": [
            "Organization",
            "ProfessionalService",
            "LocalBusiness"
          ]
        },
        "description": "A machine-readable registry of quantitative and qualitative claims published by Wavect, with review dates, localized page appearances and public third-party citations where available.",
        "inLanguage": "en",
        "isAccessibleForFree": true,
        "license": "https://creativecommons.org/licenses/by/4.0/",
        "name": "Wavect verified publication claims",
        "url": "https://wavect.io/verified-claims.json"
      },
      "url": "https://wavect.io/"
    },
    {
      "@id": "https://wavect.io/team/kevin-riedl/#person",
      "@type": "Person",
      "jobTitle": "Managing Director",
      "name": "Kevin Riedl",
      "sameAs": [
        "https://www.wikidata.org/wiki/Q139796365",
        "https://www.linkedin.com/in/wsdt",
        "https://github.com/wsdt"
      ],
      "url": "https://wavect.io/team/kevin-riedl/",
      "worksFor": {
        "@id": "https://wavect.io/#organization",
        "@type": [
          "Organization",
          "ProfessionalService",
          "LocalBusiness"
        ]
      }
    },
    {
      "@id": "https://wavect.io/team/christof-jori/#person",
      "@type": "Person",
      "jobTitle": "Managing Director",
      "name": "Christof Jori",
      "sameAs": [
        "https://www.wikidata.org/wiki/Q139796367",
        "https://www.linkedin.com/in/jocr77/",
        "https://github.com/jo-chris"
      ],
      "url": "https://wavect.io/team/christof-jori/",
      "worksFor": {
        "@id": "https://wavect.io/#organization",
        "@type": [
          "Organization",
          "ProfessionalService",
          "LocalBusiness"
        ]
      }
    },
    {
      "@id": "https://wavect.io/#website",
      "@type": "WebSite",
      "inLanguage": [
        "en",
        "de",
        "es",
        "zh"
      ],
      "name": "Wavect",
      "potentialAction": {
        "@type": "SearchAction",
        "query-input": "required name=search_term_string",
        "target": {
          "@type": "EntryPoint",
          "urlTemplate": "https://wavect.io/search/?q={search_term_string}"
        }
      },
      "publisher": {
        "@id": "https://wavect.io/#organization",
        "@type": [
          "Organization",
          "ProfessionalService",
          "LocalBusiness"
        ]
      },
      "url": "https://wavect.io/"
    },
    {
      "@id": "https://wavect.io/blog/context-language-models-vs-compaction/#webpage",
      "@type": "WebPage",
      "dateModified": "2026-10-02",
      "inLanguage": "en",
      "isPartOf": {
        "@id": "https://wavect.io/#website",
        "@type": "WebSite"
      },
      "lastReviewed": "2026-10-02",
      "url": "https://wavect.io/blog/context-language-models-vs-compaction/"
    }
  ]
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "BlogPosting",
  "abstract": "Context Language Models let models edit their active working context. The reference implementation still uses a harness with protected instructions, budget checks and edit validation. Context editing is distinct from persistent memory, and it does not itself update model weights. The reviewed code uses CC BY-NC 4.0, not unrestricted commercial-use licensing. Pilot the context policy against your existing compaction approach on held-out tasks, measuring evidence retention and total cost. Test Suffix Cache Reuse separately: removing visible text does not demonstrate removal from cached state.",
  "articleBody": " Blog overview/AI and agents/Models and infrastructure Context Language Models vs Compaction: What to Pilot TL;DR Context Language Models let models edit their active working context. The reference implementation still uses a harness with protected instructions, budget checks and edit validation. Context editing is distinct from persistent memory, and it does not itself update model weights. The reviewed code uses CC BY-NC 4.0, not unrestricted commercial-use licensing. Pilot the context policy against your existing compaction approach on held-out tasks, measuring evidence retention and total cost. Test Suffix Cache Reuse separately: removing visible text does not demonstrate removal from cached state. An agent that can edit its working context is more interesting than an agent with a bigger notebook. It can decide which failed searches to discard, which constraints to retain and when to rewrite its own intermediate plan. The useful engineering question is whether those decisions improve your workflow enough to replace its current compaction strategy. Context Language Models (CLMs) expose the model's live context as editable text, rather than leaving every context-management decision to application code. The paper is a collaboration involving the University of Washington, Meta Superintelligence Labs, MIT and Trillium Labs. Its arXiv submission is dated 29 September 2026, not an unspecified October launch day. The paper and submission history establish the research and dates. This is not the Contrastive Language Model CLM-8B used to rank candidate actions. That separate system is covered in our CLM-8B self-hosting and verifier guide. Sources reviewed on 2 October 2026. Code references are pinned to 18dc11115f50f261233c5bba7937834491e307e8. This is a source-based engineering assessment, not a Wavect benchmark or a claim that we deployed CLMs for a client. What did Meta and its collaborators actually release? The release brings together a context-editing harness, in-context skill improvement, reinforcement-learning code and a serving optimization called Suffix Cache Reuse. It is not simply a new model architecture that makes existing infrastructure unnecessary. The pinned repository describes the released components. The notebook analogy is useful only up to a point. A normal scratchpad adds notes beside a conversation. Here, edits can change the conversation material sent into subsequent model calls. Think of it as an editable working set, not permission to rewrite the application's authority or history. For a research agent, we would want that working set to preserve the question, unresolved contradictions and links to supporting evidence while dropping repetitive search output. For a coding agent, it should preserve the failing test, relevant code locations and unsuccessful fixes. These are proposed retention goals, not automatic properties of every CLM run. The paper also describes agents creating their own trackers and reusable compaction helpers. That is the compelling part: context management can become task-specific behavior rather than a growing list of hand-written deletion rules. These observed behaviors are examples, not guarantees for a different model or workload. See the paper’s context-editing examples. How does the editable context file work? The reference ClmAgent uses Harbor. Before a command, it mirrors editable conversation turns into /tmp/.live_ctx/LIVE_CTX_MAIN.txt. The model can edit that file with ordinary shell tools. The harness reads changes back into messages while keeping the system prompt and original task pinned. It also applies token nudges, an edit gate and overflow recovery. The harness documentation specifies the mirror, protected prefix and control loop. That is a different division of work, not the absence of a harness. The model chooses what to retain; software still has to decide which changes are valid, which tools may run and when the execution must stop. Our agent harness engineering guide covers that broader runtime responsibility. A model-generated summary saying “permission was granted” should never become permission. Keep authorization in trusted application state, outside editable notes. Similarly, do not let a rewritten account of a failed test replace the actual test result. Preserve an independent record of consequential actions. Editing context does not update model weights. The release's in-context-learning workflow can evolve a SKILL.md while leaving both weights and harness fixed; reinforcement learning is a separate route. The skill-evolution implementation explicitly distinguishes these mechanisms. “Policies live in the weights” therefore describes one possible training outcome, not everything happening whenever a model edits a file. CLMs versus compaction, memory files and RAG The replacement decision should concern one layer at a time. Server-side compaction can summarize earlier turns automatically or on application request; it is not necessarily a",
  "articleSection": "Agent Engineering",
  "author": {
    "@id": "https://wavect.io/team/kevin-riedl/#person",
    "@type": "Person",
    "name": "Kevin Riedl",
    "sameAs": [
      "https://www.wikidata.org/wiki/Q139796365",
      "https://www.linkedin.com/in/wsdt",
      "https://github.com/wsdt"
    ],
    "url": "https://wavect.io/team/kevin-riedl/"
  },
  "citation": [
    {
      "@type": "WebPage",
      "name": "The paper and submission history establish the research and dates",
      "url": "https://arxiv.org/abs/2609.37725v1"
    },
    {
      "@type": "WebPage",
      "name": "The pinned repository describes the released components",
      "url": "https://github.com/facebookresearch/context-language-models/blob/18dc11115f50f261233c5bba7937834491e307e8/README.md"
    },
    {
      "@type": "WebPage",
      "name": "The harness documentation specifies the mirror, protected prefix and control loop",
      "url": "https://github.com/facebookresearch/context-language-models/blob/18dc11115f50f261233c5bba7937834491e307e8/clm/clm_harness/README.md"
    },
    {
      "@type": "WebPage",
      "name": "The skill-evolution implementation explicitly distinguishes these mechanisms",
      "url": "https://github.com/facebookresearch/context-language-models/blob/18dc11115f50f261233c5bba7937834491e307e8/clm/clm_icl/README.md"
    },
    {
      "@type": "WebPage",
      "name": "Anthropic documents both automatic and on-demand compaction",
      "url": "https://platform.claude.com/docs/en/build-with-claude/compaction"
    },
    {
      "@type": "WebPage",
      "name": "Anthropic's memory-tool documentation describes that separate persistence contract",
      "url": "https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool"
    },
    {
      "@type": "WebPage",
      "name": "The prompt-caching guide explains the prefix requirement",
      "url": "https://developers.openai.com/api/docs/guides/prompt-caching"
    },
    {
      "@type": "WebPage",
      "name": "The SCR documentation explains its approximation, support scope and memory allocation",
      "url": "https://github.com/facebookresearch/context-language-models/blob/18dc11115f50f261233c5bba7937834491e307e8/suffix_cache_reuse/README.md"
    },
    {
      "@type": "WebPage",
      "name": "The pinned LICENSE contains Attribution-NonCommercial 4.0",
      "url": "https://github.com/facebookresearch/context-language-models/blob/18dc11115f50f261233c5bba7937834491e307e8/LICENSE"
    },
    {
      "@type": "WebPage",
      "name": "CC BY-NC 4.0 deed",
      "url": "https://creativecommons.org/licenses/by-nc/4.0/"
    },
    {
      "@type": "WebPage",
      "name": "OSI's definition sets out the relevant field-of-use requirement",
      "url": "https://opensource.org/osd"
    },
    {
      "@type": "WebPage",
      "name": "OpenAI's incident report gives the observation and its limits",
      "url": "https://alignment.openai.com/misalignment-reports/self-generated-prompt-injections-in-compaction-summaries/"
    }
  ],
  "dateModified": "2026-10-02",
  "datePublished": "2026-10-02",
  "description": "Context Language Models let models edit their active working context. The reference implementation still uses a harness with protected instructions, budget checks and edit validation. Context editing is distinct from persistent memory, and it does not itself update model weights. The reviewed code uses CC BY-NC 4.0, not unrestricted commercial-use licensing. Pilot the context policy against your existing compaction approach on held-out tasks, measuring evidence retention and total cost. Test Suffix Cache Reuse separately: removing visible text does not demonstrate removal from cached state.",
  "headline": "Context Language Models vs Compaction: What to Pilot",
  "image": "https://wavect.io/img/blog/headers/header_context-language-models-vs-compaction.svg",
  "inLanguage": "en",
  "keywords": "Context Language Models, AI agents, Context engineering",
  "mainEntityOfPage": {
    "@id": "https://wavect.io/blog/context-language-models-vs-compaction/",
    "@type": "WebPage"
  },
  "publisher": {
    "@id": "https://wavect.io/#organization",
    "@type": [
      "Organization",
      "ProfessionalService",
      "LocalBusiness"
    ]
  },
  "url": "https://wavect.io/blog/context-language-models-vs-compaction/",
  "wordCount": 2978
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "BreadcrumbList",
  "itemListElement": [
    {
      "@type": "ListItem",
      "item": "https://wavect.io/",
      "name": "Home",
      "position": 1
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/overview/",
      "name": "Blog overview",
      "position": 2
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/topics/ai-agents/",
      "name": "AI and agents",
      "position": 3
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/clusters/models-infrastructure/",
      "name": "Models and infrastructure",
      "position": 4
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/context-language-models-vs-compaction/",
      "name": "Context Language Models vs Compaction: What to Pilot",
      "position": 5
    }
  ]
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "FAQPage",
  "mainEntity": [
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "A Context Language Model lets the model manage its active context by editing a text representation that the runtime reads back into subsequent requests. It is not just an external notebook or a larger context window."
      },
      "name": "What is a Context Language Model?"
    },
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "No. The reviewed reference harness still pins the system prompt and task, checks token budgets, validates edits and handles overflow. Model-directed retention does not replace application authorization or recovery."
      },
      "name": "Do Context Language Models eliminate the agent harness?"
    },
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "The reference harness supports existing models, including compatible tool-calling endpoints, without requiring a new architecture. However, quality depends on the model and task. A zero-shot result should not be treated as a guarantee that every model improves."
      },
      "name": "Can existing LLMs use CLM-style context editing without training?"
    },
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "No. Editing context changes the material used by later requests. The project's in-context skill evolution keeps weights fixed; reinforcement learning is a separate training process."
      },
      "name": "Does editing the context file change model weights?"
    },
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "The reviewed repository is CC BY-NC 4.0, which has a non-commercial condition. It is not MIT or an unrestricted commercial-use release. Check whether your intended use is permitted or obtain appropriate permission."
      },
      "name": "Is Meta's released Context Language Models code unrestricted open source?"
    },
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "No. Context Language Models concern editing working context. The Contrastive Language Model CLM-8B covered in Wavect's separate guide concerns ranking candidate actions. The shared acronym does not make them the same project."
      },
      "name": "Is this the same as Contrastive Language Model CLM-8B?"
    },
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Do not assume so. The reviewed Suffix Cache Reuse design can preserve cached states influenced by an earlier prefix. Test sensitive removals against a fresh prefill, and manage logs and snapshots separately."
      },
      "name": "Does deleting context text also erase it from the cache?"
    }
  ]
}
```
