---
title: "CLM-8B Self-Hosting: vLLM, Action Cache and Verifiers"
canonical: https://wavect.io/blog/clm-8b-self-hosting-action-cache-verifier/
language: en
description: "Self-host CLM-8B with vLLM, design reusable action candidates, fix token-limit and API issues, and interpret its fine-tuned coding verifier results correctly."
image: "https://wavect.io/img/blog/headers/header_clm-8b-self-hosting-action-cache-verifier.png"
---

[**Back**](/blog/overview/)

[![Kevin Riedl](/img/team/kevin.webp)](/team/kevin-riedl/)

[Kevin Riedl](/team/kevin-riedl/) https://linkedin.com/in/wsdt

14 min read · 28 Sep 2026 Last reviewed September 28, 2026

[**Next**](/blog/llm-as-a-verifier/)

# CLM-8B Self-Hosting: vLLM, Action Cache and Verifiers

TL;DR

CLM-8B scores supplied candidates using a frozen Qwen3-8B encoder and small state/action projection heads. Self-hosting still needs the encoder, not just the head download. Reused candidate text can reduce encoding work, but cached vectors are not cached permissions or decisions. This guide covers a private two-service setup, the 2,048-token truncation boundary, request validation and verifier evaluation. The published coding scores use fine-tuned heads on held-out subsets, not the reference checkpoint zero-shot.

**CLM-8B is useful when an agent must choose among known actions, not write another answer.** Its distinctive deployment feature is independent state and action encoding: a changing state can be compared with previously encoded candidates. That makes reusable action catalogs worth investigating before replacing a generative step with yet another prompt.

The released `CLM-v0.1-8B` contains state and action projection heads tied to a frozen Qwen3-8B encoder. The heads are roughly 20 million trainable parameters each; the base encoder still runs at inference time. The reference weights are Apache 2.0. [The CLM model card describes the architecture, license and limits](https://huggingface.co/Contrastive-LM/CLM-v0.1-8B). This is a contrastive language model, not contract lifecycle management software or a general chat model.

**The engineering question is not “Did CLM beat Jev?” It is “Can this application reuse candidates, preserve decision quality and reduce measured end-to-end latency?”** Our [Jev technical review](/blog/jev-ai-decision-model-review/), [Laya versus Jev benchmark analysis](/blog/laya-vs-jev-benchmark-ai-startup-moat/) and [business workflow guide](/blog/laya-jev-business-workflows-roi/) cover those separate topics. Here we focus on running CLM, shaping its inputs and evaluating its verifier role.

Sources reviewed on 28 September 2026. Code inspection is pinned to CLM commit `bb42c6c5bf914fd449bed2f6ca65be80602cb1f7`. The examples below are proposed integration patterns, not a Wavect GPU benchmark or a claim of production deployment.

## What does CLM-8B actually replace?

**CLM replaces a scoring or selection step when the candidate set is already available.** A planner, retrieval system, deterministic rule or another model must still supply the candidates. CLM cannot choose a correct action that is absent, generate an arbitrary tool argument, execute the chosen tool or establish permission to use it.

The project provides a TypeSafe-compatible typed API and a direct ranking endpoint. It describes training on approximately 60 million question-answer pairs, 30 million synthetic hard negatives and one million agentic trajectories. State and action representations are aligned with a contrastive objective rather than trained to produce the final decision as a generated paragraph. [The pinned project README explains the training and serving design](https://github.com/Contrastive-LM/CLM/blob/bb42c6c5bf914fd449bed2f6ca65be80602cb1f7/README.md).

Do not contrast CLM with Jev as “choosing versus writing.” Jev is itself a typed decision model rather than an ordinary answer-writing chatbot. [TypeSafe's own introduction establishes that distinction](https://docs.typesafe.ai/introduction). A shared request shape does not establish interchangeable predictions, calibrated probabilities or identical operational behavior.

A narrow first integration is a read-only diagnostic assistant. Your application supplies approved diagnostic actions, CLM ranks them against an error report, and trusted code verifies the selected action before a separate executor can run anything. This is a proposed design, not evidence of CLM's accuracy on your logs.

## What do the speed and coding results actually establish?

**Read the launch numbers as author-reported experimental results, not a universal replacement verdict.** The project reports up to 9x lower latency in selected zero-shot tasks and roughly 13x speedup around 1,000 candidates. Those are different workloads and caching conditions, not two guarantees for every request. See the [reviewed model card](#source-model) for the author-reported speedups.

The T-Rex reproduction matters because it documents a physics planner that labels safe actions, multiple in-flight requests and an enabled safety shield. CLM was run locally on an RTX 4090; the compared Jev endpoint resolved to `jev-1.13.0`. Its latencies are client-side per request. Survival therefore measures the combined system, not an unassisted model's understanding of screenshots. [The T-Rex methodology identifies the planner, shield and execution setup](https://github.com/Contrastive-LM/CLM/blob/bb42c6c5bf914fd449bed2f6ca65be80602cb1f7/examples/t_rex/README.md). Do not turn a game-loop result into a production latency objective.

For coding, CLM ranks candidate solutions produced by other models. The README reports 81.6% on 38 held-out DeepSWE tasks and 87.6% on 30 held-out Terminal-Bench 2.1 tasks after task-specific head fine-tuning. These are not zero-shot results for the downloadable reference head or full-suite leaderboard scores. The stated verifier timing setup uses an H100; it is not the same setup as the T-Rex example. The [pinned results section](#source-readme) states the task counts and hardware.

The released DeepSWE head gives a more concrete interpretation: best-of-four selection, the mean of the final 12 available step scores, and 31 successes out of 38 held-out tasks. It lists pass@1 as 28/38 and an oracle ceiling of 34/38. [The DeepSWE head card publishes the task split, scoring rule and checkpoint hash](https://huggingface.co/Contrastive-LM/deepswe-clm-heads-8k). That is three additional successful selections over the stated pass@1 baseline on this set, not proof that an 8B model independently solved a full coding benchmark.

For your own comparison, hold candidate generation, task split, hardware, concurrency and network boundary constant. Publish accuracy and abstention alongside latency. A fast incorrect selection creates rework; removing that rework from the timing would manufacture an improvement.

## How to self-host CLM-8B with vLLM

**Run two services: a Qwen3-8B pooling encoder and the CLM scoring API.** A small projection-head download does not mean a small complete runtime. CLM's package declares Python 3.10 or newer and dependencies including PyTorch, vLLM, FastAPI and NumPy. [The pinned package definition records the actual dependencies](https://github.com/Contrastive-LM/CLM/blob/bb42c6c5bf914fd449bed2f6ca65be80602cb1f7/pyproject.toml).

Use an isolated, CUDA-capable Linux environment suitable for your chosen vLLM and PyTorch builds. The CLM source pin below does not pin every dependency or model-weight revision. Resolve those versions for your hardware, then save your environment lock and encoder revision. No minimum VRAM or hardware cost is asserted here.

```
python3 -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install \
  "git+https://github.com/Contrastive-LM/CLM.git@bb42c6c5bf914fd449bed2f6ca65be80602cb1f7"
python -m pip freeze > clm-environment.txt
```

Start the encoder in one activated terminal. Use the reference Qwen3-8B encoder and its expected pooling behavior, not an arbitrary embedding model with the same output dimension:

```
vllm serve Qwen/Qwen3-8B \
  --served-model-name qwen3-8b \
  --runner pooling \
  --max-model-len 2048 \
  --host 127.0.0.1 \
  --port 8090
```

In a second activated terminal, supply `CLM_API_KEY` through your secret-management process. Keep that same secret available to the client shell; never commit it. Start a private API with explicitly bounded cache memory:

```
: "${CLM_API_KEY:?Set CLM_API_KEY in this shell first}"
export CLM_API_KEY
clm-serve \
  --host 127.0.0.1 \
  --port 8700 \
  --emb-url http://127.0.0.1:8090/v1/embeddings \
  --emb-model qwen3-8b \
  --max-tokens 2048 \
  --device cpu \
  --action-cache 64MiB \
  --no-ui
```

This example intentionally puts the small heads and their vector cache on CPU, while the 8B encoder stays on GPU. It is a deployment illustration, not the configuration behind the headline speedups. Benchmark CPU and GPU heads on your workload before choosing.

The inspected server defaults to `0.0.0.0`, makes authentication conditional on `CLM_API_KEY`, and exposes a health response containing both `ok` and `embedder`. [The server implementation defines these flags, authentication and health behavior](https://github.com/Contrastive-LM/CLM/blob/bb42c6c5bf914fd449bed2f6ca65be80602cb1f7/src/clm/server.py). Explicit loopback binding matters. The CLM key does not automatically secure the separate encoder endpoint, and `--no-ui` is not authentication. Keep both services private; add authenticated ingress, request limits and workload isolation before remote use.

A successful health HTTP status alone is not sufficient readiness evidence. Check the encoder flag and advertised model too:

```
curl -fsS --connect-timeout 5 --max-time 15 http://127.0.0.1:8700/health \
  | python -c 'import json,sys; d=json.load(sys.stdin); sys.exit(0 if d.get("ok") and d.get("embedder") and "clm-latest" in d.get("models", []) else 1)'
```

## Send a bounded request without granting tool permissions

**Use stable option identifiers and self-contained descriptions.** In CLM's `Choice` implementation, the state encoder receives context plus question instructions. The action encoder receives each option's description, or its key when the description is empty. Two different IDs with the same description are therefore not meaningfully different candidates to that encoder. [The schema implementation shows exactly how states and candidates become text](https://github.com/Contrastive-LM/CLM/blob/bb42c6c5bf914fd449bed2f6ca65be80602cb1f7/src/clm/schema.py).

The following synthetic English request asks for a read-only diagnostic suggestion. The fixture remains identical in each translation so results can be compared. The `review` candidate is deliberate. Application code still needs a hard fallback because a model is not guaranteed to choose the abstention option when it should.

```
{
  "model": "clm-latest",
  "state": "A CI job fails during dependency installation. The log reports a lockfile mismatch. No production change is authorized.",
  "questions": {
    "next_step": {
      "type": "choice",
      "instructions": "Select the most useful permitted read-only diagnostic step, or request review when the evidence is insufficient.",
      "criteria": {
        "inspect_lockfile": "Read the manifest and lockfile to identify inconsistent dependency versions; make no changes.",
        "read_network_log": "Read existing dependency-download network logs to investigate connection failures; make no changes.",
        "review": "Ask a human to review because the supplied evidence is insufficient or no listed diagnostic is appropriate."
      }
    }
  }
}
```

Save it as `request.json`, then call the API from a shell with the same `CLM_API_KEY`:

```
: "${CLM_API_KEY:?Set the same CLM_API_KEY used by the API}"
curl -fsS --connect-timeout 5 --max-time 30 \
  -H "Authorization: Bearer ${CLM_API_KEY}" \
  -H "Content-Type: application/json" \
  --data-binary @request.json \
  http://127.0.0.1:8700/v1/systemone
```

Treat the response as a suggestion. Validate its model identity, question ID, candidate keys, numeric values and schema before consuming it. A timeout, malformed response, unknown option or missing policy version goes to review. Do not interpolate returned text into a shell command. Map a validated identifier to a separately authorized handler.

When replaying a TypeSafe-shaped request, compare semantic outcomes rather than merely HTTP 200 responses. Re-evaluate thresholds, label descriptions and input preparation. For ordinary business-process examples rather than this CLM-specific contract, use the existing workflow guide linked above.

## When does action caching actually help?

**Caching helps when the same candidate text is scored against new states.** A fixed set of diagnostics, product records or permitted actions can amortize its encoding cost. A fresh set of long generated solutions on every call cannot reuse those action embeddings across unrelated tasks, although state-side or other reuse may still help.

The inspected embedder caches normalized encoder vectors by exact input text in an in-process least-recently-used cache. Requests on misses use the configured pooling endpoint; the token-usage count tracks tokens spent on those misses. [The embedder source defines the exact-text cache and token accounting](https://github.com/Contrastive-LM/CLM/blob/bb42c6c5bf914fd449bed2f6ca65be80602cb1f7/src/clm/embedder.py). “Zero new encoder tokens” is not “zero compute,” zero hosting cost or proof that a new user's input was processed independently.

At the engine layer, projected state and action vectors use separate namespaces; the head identity participates in their namespace. [The engine separates raw embeddings, projected vectors and model selection](https://github.com/Contrastive-LM/CLM/blob/bb42c6c5bf914fd449bed2f6ca65be80602cb1f7/src/clm/engine.py). The device arena has a startup allocation and least-recently-used eviction. [The vector-cache implementation defines its bounded pools](https://github.com/Contrastive-LM/CLM/blob/bb42c6c5bf914fd449bed2f6ca65be80602cb1f7/src/clm/cache.py). Its memory budget does not bound the separate encoder-text cache or the entire process. Treat either cache as an implementation detail, not an authorization layer or durable memory store.

Three distinct things need different reuse rules:

| Object | Reuse opportunity | What must remain valid |
| --- | --- | --- |
| Candidate embedding | The same action description appears again | Encoder, pooling, preprocessing and exact text |
| Projected candidate vector | The same embedding is scored again | All of the above plus the projection-head identity |
| Final decision | The exact decision context recurs | State, candidate set, permissions, policy, model, temperature and freshness |

Our deployment recommendation is to keep a manifest of the encoder revision, pooling mode, head hash, text-rendering version and token limit. Restart and rewarm after changing encoder or preprocessing configuration rather than assuming a head reload invalidates everything. Version the action catalog separately, and never reuse a final decision after a permission or record change.

A shared process cache is not a tenant-isolation guarantee. Decide whether sensitive workloads need separate processes or stronger isolation; include logs, memory, cache timing and request routing in that assessment. Warm only authorized candidate catalogs. Do not embed a global list and assume a high score permits access to every member.

## Why can CLM return a confident but wrong choice?

**CLM probabilities are relative to the candidates supplied.** A softmax must assign its mass somewhere even when all options are poor. With one candidate, the distribution is necessarily 1.0; that is not evidence that the action is correct. Adding plausible alternatives can change probabilities without changing the underlying task.

The published `confidence` field is the highest probability minus the mean of the other probabilities. It is neither a general factual-verification probability nor a substitute for a calibrated acceptance policy. The formula is visible in the schema source cited above. Treat `0.9` as a model output to validate, not a portable 90% guarantee.

For this CLM integration, test missing-correct-option cases, duplicate descriptions, near-duplicate candidates, contradictory context, candidate ordering and changes in set size. Measure error at the proposed acceptance threshold separately for each relevant language and action family. This article is localized; that does not establish multilingual CLM performance.

Keep permission checks outside the candidate scorer and repeat them immediately before execution. For consequential actions, a fallback is an application decision, not simply another label the same model may ignore. Restrict the initial rollout to suggestions until held-out evidence supports more authority.

## The 2,048-token limit and common setup failures

**The reference quickstart truncates input texts at 2,048 tokens.** Raising only the encoder's maximum does not remove the CLM-side truncation. To evaluate 8K inputs, increase both `vllm serve --max-model-len 8192` and `clm-serve --max-tokens 8192`, then recheck GPU memory and quality. Longer accepted input does not establish that the head was validated for that use.

The inspected embedder sends `truncate_prompt_tokens`; the state/question rendering and action descriptions count toward what is embedded. Test instructions and critical evidence near the boundary. Reject or deliberately summarize oversized inputs in trusted application code rather than silently treating a truncated decision as complete.

| Symptom | Check first | Controlled response |
| --- | --- | --- |
| HTTP 502 from CLM | Encoder URL, model name, process and vLLM error | Keep the request for review; do not substitute approval |
| HTTP 401 | Matching CLM bearer key | Fix credentials without printing the key |
| HTTP 422 | Question type, criteria, model and temperature | Validate the request before retrying |
| Healthy API, unusable decisions | `embedder` status and an actual scoring smoke test | Gate readiness on both services |
| Slow “warm” requests | Candidate-text changes, cache eviction and fresh states | Measure cache misses instead of assuming reuse |
| Good short tests, poor long tasks | Both token limits and critical text placement | Add long-input fixtures before rollout |

Do not replace Qwen3-8B with an unrelated encoder because it appears faster. The head is tied to its training representation and pooling. Quantization, alternative runtimes and longer contexts each require separate quality and latency validation.

## Use CLM as a verifier, not the code generator

**A verifier selects among available solutions; it does not eliminate generation or testing.** A candidate pipeline can generate several solutions, run deterministic checks, score remaining candidates, and hand a selected artifact to review. Record candidate-generation cost separately from verifier time. The general architecture belongs in our [LLM-as-a-verifier guide](/blog/llm-as-a-verifier/); the CLM-specific issue is matching its head and evaluation recipe.

For the published DeepSWE experiment, the head card provides a reproducible command using stored embeddings. Run it from the pinned CLM source checkout with the documented evaluation dependencies and Hugging Face CLI available:

```
git clone https://github.com/Contrastive-LM/CLM.git clm-source
git -C clm-source checkout bb42c6c5bf914fd449bed2f6ca65be80602cb1f7
cd clm-source
hf download Contrastive-LM/deepswe-clm-heads-8k --local-dir heads/deepswe
python evaluation/bon_eval.py \
  --hf-dataset Contrastive-LM/deepswe-clm-embeddings-8k \
  --checkpoint heads/deepswe/best_head.pt \
  --tasks-file heads/deepswe/heldout_tasks.json \
  --n 4 --window 12
```

This evaluates a verifier over a published embedding dataset. It does not start a coding agent, regenerate all trajectories or establish the wall-clock cost of solving the original tasks. Verify the published checkpoint and held-out-list hashes, freeze the candidate budget and preserve task-disjoint splits.

The project's fine-tuning instructions explicitly keep the data, folds, evaluation set and best-of-N settings fixed while changing training. [The pinned fine-tuning guide defines those experimental boundaries](https://github.com/Contrastive-LM/CLM/blob/bb42c6c5bf914fd449bed2f6ca65be80602cb1f7/docs/FINETUNING.md). Fine-tuning small heads may reduce trainable-parameter cost, but data labeling, embedding generation, evaluation and serving remain real work. Do not tune on the held-out tasks and report the resulting score as untouched test performance.

Also check each artifact's own license metadata: the reference CLM weights are Apache 2.0, while the reviewed DeepSWE head card labels its artifact MIT. Do not assume every downstream checkpoint shares one license. “Open weights” does not mean free infrastructure or a warranty of suitability.

## Benchmark the full path before switching

**A tenfold faster decision component does not make the whole agent tenfold faster.** As a deliberately hypothetical budget, let one workflow spend 900 ms outside selection and 100 ms selecting. Replacing 100 ms with 10 ms gives 910 ms total instead of 1,000 ms: 9% less elapsed time, or about 1.10x overall speedup. This arithmetic is not a CLM measurement.

For a CLM pilot, compare fresh states with cold candidates, fresh states with warm candidates, repeated identical requests and changing candidate catalogs. The repeated-request test can measure a different cache path from real work. Keep results separate. Report end-to-end p50/p95, failures, memory use, accepted-decision accuracy and review rate at the same concurrency and candidate budget.

Before promotion, require a versioned encoder/head pair, explicit candidate ownership, tested error handling, monitored cache behavior, a representative held-out evaluation and a rollback path. Keep the existing selector running in shadow comparisons before giving the new service write authority. Use the broader [pilot kill-or-scale scorecard](/blog/ai-pilot-kill-or-scale-scorecard/) for the investment decision rather than creating a second rollout methodology here.

As of the review date, the reference model card describes multimodal CLM-35B as planned for early October 2026. That is a roadmap statement, not a released capability or a guaranteed delivery date. Validate any later checkpoint on its own terms instead of transferring the 8B results to it.

For an implementation assessment, Wavect's [AI integration service](/services/artificial-intelligence/) can address the selector, permissions, observability and evaluation together. The [Twinsoft AI case study](/case-studies/twinsoft-ai/) is related delivery context, not a CLM deployment claim. Bring the [pre-launch QA checklist](/software-development-guide/software-qa-checklist-before-launch/) and a representative candidate set to [discuss a bounded CLM verifier pilot](/contact/).

## CLM-8B deployment questions

### Can I run CLM-8B using only the projection-head download?

No. The reference heads depend on Qwen3-8B last-token-pooled embeddings. You still need the encoder service and its memory and compute budget. Putting the heads on CPU does not turn the whole 8B runtime into a small CPU-only model.

### Does the CLM action cache also cache tool permissions?

No. It reuses text embeddings and projected vectors, not authorization. Keep permission checks in trusted application code and repeat them before execution. Reuse of a final decision requires the full state, candidate set, policy and freshness conditions to remain valid.

### Why are CLM requests truncated even after raising the vLLM context limit?

The reference setup has two limits: vLLM max-model-len and CLM max-tokens. Raise both for a longer-input evaluation. Then check memory, critical text near the boundary and quality; accepting more tokens is not proof of validated long-context behavior.

### Is an HTTP 200 from the CLM health endpoint enough for readiness?

No. The inspected response can contain ok=true while embedder=false. Check the encoder status and available model, then run an actual scoring smoke test. Keep the embedding endpoint and the CLM API private; the CLM bearer key does not automatically secure the encoder.

### Are the published CLM coding scores zero-shot results?

No. The cited coding results use task-specific fine-tuned heads on held-out subsets: 38 DeepSWE tasks and 30 Terminal-Bench 2.1 tasks. CLM selects among generated candidates. The reference head download alone does not reproduce those scores.

### Does CLM confidence of 0.9 mean the action is 90% correct?

Not automatically. Choice probabilities are relative to the supplied candidates. The inspected confidence formula subtracts the mean of other probabilities from the top probability. Validate acceptance thresholds on representative held-out cases and retain a hard review fallback.

### Can I replace the Qwen3-8B encoder with another embedding model?

Not as an assumed compatible substitution. The reference heads depend on their training representation and last-token pooling. A different encoder, quantization or input preparation needs separate quality validation; matching vector dimensions alone is insufficient.

### Is multimodal CLM-35B available in this guide?

No. As reviewed on 28 September 2026, the model card describes it as planned for early October. The setup here targets the released CLM-v0.1-8B reference head. A future checkpoint needs its own availability, license, runtime and quality checks.

## Final thoughts

Treat CLM as a replaceable scoring component with a versioned encoder, explicit candidate contract and measured quality. Reuse vectors where the inputs genuinely repeat, but revalidate decisions and permissions. The first milestone is a reliable private selector on your own held-out tasks, not a headline speedup.

## You may also like..

[**LLM-as-a-Verifier: the broader architecture** Understand candidate generation, verification costs and deterministic gates.](/blog/llm-as-a-verifier/) [**Jev: the technical review** Explore Jev separately, without turning a CLM deployment guide into another model comparison.](/blog/jev-ai-decision-model-review/)

Models and infrastructure

## Continue through this cluster

Model selection, inference economics, local deployment, compression and serving architecture.

[Start with the cornerstone**Self-Hosting LLMs in the EU: When Open Weights Actually Pay Off**](/blog/self-hosting-llms-eu-cost/)

- [DeerFlow 2.0: Docker Setup, Sandboxes and Memory](/blog/deerflow-2-docker-setup-sandbox-memory/)
- [AnyJev: LLM Calibration and Option-Order Bias](/blog/anyjev-calibration-option-order-bias/)
- [LiteAgents SDK: Per-Turn Routing, Setup and Migration](/blog/liteagents-sdk-per-turn-model-routing/)
- [mcp-memory-service: Shared Memory for Claude Code and Cursor](/blog/mcp-memory-service-claude-code-cursor/)
- [Claude Opus 5.5: Best Uses, Prompts and Effort Settings](/blog/claude-opus-5-5-best-use-cases-workflows/)

[**Back**](/blog/overview/)

[![Kevin Riedl](/img/team/kevin.webp)](/team/kevin-riedl/)

[Kevin Riedl](/team/kevin-riedl/) https://linkedin.com/in/wsdt

14 min read · 28 Sep 2026 Last reviewed September 28, 2026

[**Next**](/blog/llm-as-a-verifier/)

## Structured Data

```json
{
  "@context": "https://schema.org",
  "@graph": [
    {
      "@id": "https://wavect.io/#organization",
      "@type": [
        "Organization",
        "ProfessionalService",
        "LocalBusiness"
      ],
      "employee": [
        {
          "@id": "https://wavect.io/team/kevin-riedl/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Kevin Riedl",
          "url": "https://wavect.io/team/kevin-riedl/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        },
        {
          "@id": "https://wavect.io/team/christof-jori/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Christof Jori",
          "url": "https://wavect.io/team/christof-jori/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        }
      ],
      "founder": [
        {
          "@id": "https://wavect.io/team/kevin-riedl/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Kevin Riedl",
          "url": "https://wavect.io/team/kevin-riedl/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        },
        {
          "@id": "https://wavect.io/team/christof-jori/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Christof Jori",
          "url": "https://wavect.io/team/christof-jori/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        }
      ],
      "legalRepresentative": [
        {
          "@id": "https://wavect.io/team/kevin-riedl/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Kevin Riedl",
          "url": "https://wavect.io/team/kevin-riedl/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        },
        {
          "@id": "https://wavect.io/team/christof-jori/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Christof Jori",
          "url": "https://wavect.io/team/christof-jori/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        }
      ],
      "name": "Wavect GmbH",
      "subjectOf": {
        "@id": "https://wavect.io/verified-claims.json#dataset",
        "@type": "Dataset",
        "creator": {
          "@id": "https://wavect.io/#organization",
          "@type": [
            "Organization",
            "ProfessionalService",
            "LocalBusiness"
          ]
        },
        "description": "A machine-readable registry of quantitative and qualitative claims published by Wavect, with review dates, localized page appearances and public third-party citations where available.",
        "inLanguage": "en",
        "isAccessibleForFree": true,
        "license": "https://creativecommons.org/licenses/by/4.0/",
        "name": "Wavect verified publication claims",
        "url": "https://wavect.io/verified-claims.json"
      },
      "url": "https://wavect.io/"
    },
    {
      "@id": "https://wavect.io/team/kevin-riedl/#person",
      "@type": "Person",
      "jobTitle": "Managing Director",
      "name": "Kevin Riedl",
      "sameAs": [
        "https://www.wikidata.org/wiki/Q139796365",
        "https://www.linkedin.com/in/wsdt",
        "https://github.com/wsdt"
      ],
      "url": "https://wavect.io/team/kevin-riedl/",
      "worksFor": {
        "@id": "https://wavect.io/#organization",
        "@type": [
          "Organization",
          "ProfessionalService",
          "LocalBusiness"
        ]
      }
    },
    {
      "@id": "https://wavect.io/team/christof-jori/#person",
      "@type": "Person",
      "jobTitle": "Managing Director",
      "name": "Christof Jori",
      "sameAs": [
        "https://www.wikidata.org/wiki/Q139796367",
        "https://www.linkedin.com/in/jocr77/",
        "https://github.com/jo-chris"
      ],
      "url": "https://wavect.io/team/christof-jori/",
      "worksFor": {
        "@id": "https://wavect.io/#organization",
        "@type": [
          "Organization",
          "ProfessionalService",
          "LocalBusiness"
        ]
      }
    },
    {
      "@id": "https://wavect.io/#website",
      "@type": "WebSite",
      "inLanguage": [
        "en",
        "de",
        "es",
        "zh"
      ],
      "name": "Wavect",
      "potentialAction": {
        "@type": "SearchAction",
        "query-input": "required name=search_term_string",
        "target": {
          "@type": "EntryPoint",
          "urlTemplate": "https://wavect.io/search/?q={search_term_string}"
        }
      },
      "publisher": {
        "@id": "https://wavect.io/#organization",
        "@type": [
          "Organization",
          "ProfessionalService",
          "LocalBusiness"
        ]
      },
      "url": "https://wavect.io/"
    },
    {
      "@id": "https://wavect.io/blog/clm-8b-self-hosting-action-cache-verifier/#webpage",
      "@type": "WebPage",
      "dateModified": "2026-09-28",
      "inLanguage": "en",
      "isPartOf": {
        "@id": "https://wavect.io/#website",
        "@type": "WebSite"
      },
      "lastReviewed": "2026-09-28",
      "url": "https://wavect.io/blog/clm-8b-self-hosting-action-cache-verifier/"
    }
  ]
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "BlogPosting",
  "abstract": "CLM-8B scores supplied candidates using a frozen Qwen3-8B encoder and small state/action projection heads. Self-hosting still needs the encoder, not just the head download. Reused candidate text can reduce encoding work, but cached vectors are not cached permissions or decisions. This guide covers a private two-service setup, the 2,048-token truncation boundary, request validation and verifier evaluation. The published coding scores use fine-tuned heads on held-out subsets, not the reference checkpoint zero-shot.",
  "articleBody": " Blog overview/AI and agents/Models and infrastructure CLM-8B Self-Hosting: vLLM, Action Cache and Verifiers TL;DR CLM-8B scores supplied candidates using a frozen Qwen3-8B encoder and small state/action projection heads. Self-hosting still needs the encoder, not just the head download. Reused candidate text can reduce encoding work, but cached vectors are not cached permissions or decisions. This guide covers a private two-service setup, the 2,048-token truncation boundary, request validation and verifier evaluation. The published coding scores use fine-tuned heads on held-out subsets, not the reference checkpoint zero-shot. CLM-8B is useful when an agent must choose among known actions, not write another answer. Its distinctive deployment feature is independent state and action encoding: a changing state can be compared with previously encoded candidates. That makes reusable action catalogs worth investigating before replacing a generative step with yet another prompt. The released CLM-v0.1-8B contains state and action projection heads tied to a frozen Qwen3-8B encoder. The heads are roughly 20 million trainable parameters each; the base encoder still runs at inference time. The reference weights are Apache 2.0. The CLM model card describes the architecture, license and limits. This is a contrastive language model, not contract lifecycle management software or a general chat model. The engineering question is not “Did CLM beat Jev?” It is “Can this application reuse candidates, preserve decision quality and reduce measured end-to-end latency?” Our Jev technical review, Laya versus Jev benchmark analysis and business workflow guide cover those separate topics. Here we focus on running CLM, shaping its inputs and evaluating its verifier role. Sources reviewed on 28 September 2026. Code inspection is pinned to CLM commit bb42c6c5bf914fd449bed2f6ca65be80602cb1f7. The examples below are proposed integration patterns, not a Wavect GPU benchmark or a claim of production deployment. What does CLM-8B actually replace? CLM replaces a scoring or selection step when the candidate set is already available. A planner, retrieval system, deterministic rule or another model must still supply the candidates. CLM cannot choose a correct action that is absent, generate an arbitrary tool argument, execute the chosen tool or establish permission to use it. The project provides a TypeSafe-compatible typed API and a direct ranking endpoint. It describes training on approximately 60 million question-answer pairs, 30 million synthetic hard negatives and one million agentic trajectories. State and action representations are aligned with a contrastive objective rather than trained to produce the final decision as a generated paragraph. The pinned project README explains the training and serving design. Do not contrast CLM with Jev as “choosing versus writing.” Jev is itself a typed decision model rather than an ordinary answer-writing chatbot. TypeSafe's own introduction establishes that distinction. A shared request shape does not establish interchangeable predictions, calibrated probabilities or identical operational behavior. A narrow first integration is a read-only diagnostic assistant. Your application supplies approved diagnostic actions, CLM ranks them against an error report, and trusted code verifies the selected action before a separate executor can run anything. This is a proposed design, not evidence of CLM's accuracy on your logs. What do the speed and coding results actually establish? Read the launch numbers as author-reported experimental results, not a universal replacement verdict. The project reports up to 9x lower latency in selected zero-shot tasks and roughly 13x speedup around 1,000 candidates. Those are different workloads and caching conditions, not two guarantees for every request. See the reviewed model card for the author-reported speedups. The T-Rex reproduction matters because it documents a physics planner that labels safe actions, multiple in-flight requests and an enabled safety shield. CLM was run locally on an RTX 4090; the compared Jev endpoint resolved to jev-1.13.0. Its latencies are client-side per request. Survival therefore measures the combined system, not an unassisted model's understanding of screenshots. The T-Rex methodology identifies the planner, shield and execution setup. Do not turn a game-loop result into a production latency objective. For coding, CLM ranks candidate solutions produced by other models. The README reports 81.6% on 38 held-out DeepSWE tasks and 87.6% on 30 held-out Terminal-Bench 2.1 tasks after task-specific head fine-tuning. These are not zero-shot results for the downloadable reference head or full-suite leaderboard scores. The stated verifier timing setup uses an H100; it is not the same setup as the T-Rex example. The pinned results section states the task counts and hardware. The released DeepSWE head gives a more concrete interpretation: best-of-four",
  "articleSection": "AI Infrastructure",
  "author": {
    "@id": "https://wavect.io/team/kevin-riedl/#person",
    "@type": "Person",
    "name": "Kevin Riedl",
    "sameAs": [
      "https://www.wikidata.org/wiki/Q139796365",
      "https://www.linkedin.com/in/wsdt",
      "https://github.com/wsdt"
    ],
    "url": "https://wavect.io/team/kevin-riedl/"
  },
  "citation": [
    {
      "@type": "WebPage",
      "name": "The CLM model card describes the architecture, license and limits",
      "url": "https://huggingface.co/Contrastive-LM/CLM-v0.1-8B"
    },
    {
      "@type": "WebPage",
      "name": "The pinned project README explains the training and serving design",
      "url": "https://github.com/Contrastive-LM/CLM/blob/bb42c6c5bf914fd449bed2f6ca65be80602cb1f7/README.md"
    },
    {
      "@type": "WebPage",
      "name": "TypeSafe's own introduction establishes that distinction",
      "url": "https://docs.typesafe.ai/introduction"
    },
    {
      "@type": "WebPage",
      "name": "The T-Rex methodology identifies the planner, shield and execution setup",
      "url": "https://github.com/Contrastive-LM/CLM/blob/bb42c6c5bf914fd449bed2f6ca65be80602cb1f7/examples/t_rex/README.md"
    },
    {
      "@type": "WebPage",
      "name": "The DeepSWE head card publishes the task split, scoring rule and checkpoint hash",
      "url": "https://huggingface.co/Contrastive-LM/deepswe-clm-heads-8k"
    },
    {
      "@type": "WebPage",
      "name": "The pinned package definition records the actual dependencies",
      "url": "https://github.com/Contrastive-LM/CLM/blob/bb42c6c5bf914fd449bed2f6ca65be80602cb1f7/pyproject.toml"
    },
    {
      "@type": "WebPage",
      "name": "The server implementation defines these flags, authentication and health behavior",
      "url": "https://github.com/Contrastive-LM/CLM/blob/bb42c6c5bf914fd449bed2f6ca65be80602cb1f7/src/clm/server.py"
    },
    {
      "@type": "WebPage",
      "name": "The schema implementation shows exactly how states and candidates become text",
      "url": "https://github.com/Contrastive-LM/CLM/blob/bb42c6c5bf914fd449bed2f6ca65be80602cb1f7/src/clm/schema.py"
    },
    {
      "@type": "WebPage",
      "name": "The embedder source defines the exact-text cache and token accounting",
      "url": "https://github.com/Contrastive-LM/CLM/blob/bb42c6c5bf914fd449bed2f6ca65be80602cb1f7/src/clm/embedder.py"
    },
    {
      "@type": "WebPage",
      "name": "The engine separates raw embeddings, projected vectors and model selection",
      "url": "https://github.com/Contrastive-LM/CLM/blob/bb42c6c5bf914fd449bed2f6ca65be80602cb1f7/src/clm/engine.py"
    },
    {
      "@type": "WebPage",
      "name": "The vector-cache implementation defines its bounded pools",
      "url": "https://github.com/Contrastive-LM/CLM/blob/bb42c6c5bf914fd449bed2f6ca65be80602cb1f7/src/clm/cache.py"
    },
    {
      "@type": "WebPage",
      "name": "The pinned fine-tuning guide defines those experimental boundaries",
      "url": "https://github.com/Contrastive-LM/CLM/blob/bb42c6c5bf914fd449bed2f6ca65be80602cb1f7/docs/FINETUNING.md"
    }
  ],
  "dateModified": "2026-09-28",
  "datePublished": "2026-09-28",
  "description": "CLM-8B scores supplied candidates using a frozen Qwen3-8B encoder and small state/action projection heads. Self-hosting still needs the encoder, not just the head download. Reused candidate text can reduce encoding work, but cached vectors are not cached permissions or decisions. This guide covers a private two-service setup, the 2,048-token truncation boundary, request validation and verifier evaluation. The published coding scores use fine-tuned heads on held-out subsets, not the reference checkpoint zero-shot.",
  "headline": "CLM-8B Self-Hosting: vLLM, Action Cache and Verifiers",
  "image": "https://wavect.io/img/blog/headers/header_clm-8b-self-hosting-action-cache-verifier.svg",
  "inLanguage": "en",
  "keywords": "AI agents, CLM-8B, Self-hosting",
  "mainEntityOfPage": {
    "@id": "https://wavect.io/blog/clm-8b-self-hosting-action-cache-verifier/",
    "@type": "WebPage"
  },
  "publisher": {
    "@id": "https://wavect.io/#organization",
    "@type": [
      "Organization",
      "ProfessionalService",
      "LocalBusiness"
    ]
  },
  "url": "https://wavect.io/blog/clm-8b-self-hosting-action-cache-verifier/",
  "wordCount": 3647
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "BreadcrumbList",
  "itemListElement": [
    {
      "@type": "ListItem",
      "item": "https://wavect.io/",
      "name": "Home",
      "position": 1
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/overview/",
      "name": "Blog overview",
      "position": 2
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/topics/ai-agents/",
      "name": "AI and agents",
      "position": 3
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/clusters/models-infrastructure/",
      "name": "Models and infrastructure",
      "position": 4
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/clm-8b-self-hosting-action-cache-verifier/",
      "name": "CLM-8B Self-Hosting: vLLM, Action Cache and Verifiers",
      "position": 5
    }
  ]
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "FAQPage",
  "mainEntity": [
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "No. The reference heads depend on Qwen3-8B last-token-pooled embeddings. You still need the encoder service and its memory and compute budget. Putting the heads on CPU does not turn the whole 8B runtime into a small CPU-only model."
      },
      "name": "Can I run CLM-8B using only the projection-head download?"
    },
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "No. It reuses text embeddings and projected vectors, not authorization. Keep permission checks in trusted application code and repeat them before execution. Reuse of a final decision requires the full state, candidate set, policy and freshness conditions to remain valid."
      },
      "name": "Does the CLM action cache also cache tool permissions?"
    },
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "The reference setup has two limits: vLLM max-model-len and CLM max-tokens. Raise both for a longer-input evaluation. Then check memory, critical text near the boundary and quality; accepting more tokens is not proof of validated long-context behavior."
      },
      "name": "Why are CLM requests truncated even after raising the vLLM context limit?"
    },
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "No. The inspected response can contain ok=true while embedder=false. Check the encoder status and available model, then run an actual scoring smoke test. Keep the embedding endpoint and the CLM API private; the CLM bearer key does not automatically secure the encoder."
      },
      "name": "Is an HTTP 200 from the CLM health endpoint enough for readiness?"
    },
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "No. The cited coding results use task-specific fine-tuned heads on held-out subsets: 38 DeepSWE tasks and 30 Terminal-Bench 2.1 tasks. CLM selects among generated candidates. The reference head download alone does not reproduce those scores."
      },
      "name": "Are the published CLM coding scores zero-shot results?"
    },
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Not automatically. Choice probabilities are relative to the supplied candidates. The inspected confidence formula subtracts the mean of other probabilities from the top probability. Validate acceptance thresholds on representative held-out cases and retain a hard review fallback."
      },
      "name": "Does CLM confidence of 0.9 mean the action is 90% correct?"
    },
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Not as an assumed compatible substitution. The reference heads depend on their training representation and last-token pooling. A different encoder, quantization or input preparation needs separate quality validation; matching vector dimensions alone is insufficient."
      },
      "name": "Can I replace the Qwen3-8B encoder with another embedding model?"
    },
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "No. As reviewed on 28 September 2026, the model card describes it as planned for early October. The setup here targets the released CLM-v0.1-8B reference head. A future checkpoint needs its own availability, license, runtime and quality checks."
      },
      "name": "Is multimodal CLM-35B available in this guide?"
    }
  ]
}
```
