Back
Kevin Riedl

14 min read · 28 Sep 2026
Last reviewed

Next
Made on your device, with no Instagram connection. We copy the post link for Instagram’s Link sticker.

CLM-8B Self-Hosting: vLLM, Action Cache and Verifiers

CLM-8B is useful when an agent must choose among known actions, not write another answer. Its distinctive deployment feature is independent state and action encoding: a changing state can be compared with previously encoded candidates. That makes reusable action catalogs worth investigating before replacing a generative step with yet another prompt.

The released CLM-v0.1-8B contains state and action projection heads tied to a frozen Qwen3-8B encoder. The heads are roughly 20 million trainable parameters each; the base encoder still runs at inference time. The reference weights are Apache 2.0. The CLM model card describes the architecture, license and limits. This is a contrastive language model, not contract lifecycle management software or a general chat model.

The engineering question is not “Did CLM beat Jev?” It is “Can this application reuse candidates, preserve decision quality and reduce measured end-to-end latency?” Our Jev technical review, Laya versus Jev benchmark analysis and business workflow guide cover those separate topics. Here we focus on running CLM, shaping its inputs and evaluating its verifier role.

Sources reviewed on . Code inspection is pinned to CLM commit bb42c6c5bf914fd449bed2f6ca65be80602cb1f7. The examples below are proposed integration patterns, not a Wavect GPU benchmark or a claim of production deployment.

What does CLM-8B actually replace?

CLM replaces a scoring or selection step when the candidate set is already available. A planner, retrieval system, deterministic rule or another model must still supply the candidates. CLM cannot choose a correct action that is absent, generate an arbitrary tool argument, execute the chosen tool or establish permission to use it.

The project provides a TypeSafe-compatible typed API and a direct ranking endpoint. It describes training on approximately 60 million question-answer pairs, 30 million synthetic hard negatives and one million agentic trajectories. State and action representations are aligned with a contrastive objective rather than trained to produce the final decision as a generated paragraph. The pinned project README explains the training and serving design.

Do not contrast CLM with Jev as “choosing versus writing.” Jev is itself a typed decision model rather than an ordinary answer-writing chatbot. TypeSafe's own introduction establishes that distinction. A shared request shape does not establish interchangeable predictions, calibrated probabilities or identical operational behavior.

A narrow first integration is a read-only diagnostic assistant. Your application supplies approved diagnostic actions, CLM ranks them against an error report, and trusted code verifies the selected action before a separate executor can run anything. This is a proposed design, not evidence of CLM's accuracy on your logs.

What do the speed and coding results actually establish?

Read the launch numbers as author-reported experimental results, not a universal replacement verdict. The project reports up to 9x lower latency in selected zero-shot tasks and roughly 13x speedup around 1,000 candidates. Those are different workloads and caching conditions, not two guarantees for every request. See the reviewed model card for the author-reported speedups.

The T-Rex reproduction matters because it documents a physics planner that labels safe actions, multiple in-flight requests and an enabled safety shield. CLM was run locally on an RTX 4090; the compared Jev endpoint resolved to jev-1.13.0. Its latencies are client-side per request. Survival therefore measures the combined system, not an unassisted model's understanding of screenshots. The T-Rex methodology identifies the planner, shield and execution setup. Do not turn a game-loop result into a production latency objective.

For coding, CLM ranks candidate solutions produced by other models. The README reports 81.6% on 38 held-out DeepSWE tasks and 87.6% on 30 held-out Terminal-Bench 2.1 tasks after task-specific head fine-tuning. These are not zero-shot results for the downloadable reference head or full-suite leaderboard scores. The stated verifier timing setup uses an H100; it is not the same setup as the T-Rex example. The pinned results section states the task counts and hardware.

The released DeepSWE head gives a more concrete interpretation: best-of-four selection, the mean of the final 12 available step scores, and 31 successes out of 38 held-out tasks. It lists pass@1 as 28/38 and an oracle ceiling of 34/38. The DeepSWE head card publishes the task split, scoring rule and checkpoint hash. That is three additional successful selections over the stated pass@1 baseline on this set, not proof that an 8B model independently solved a full coding benchmark.

For your own comparison, hold candidate generation, task split, hardware, concurrency and network boundary constant. Publish accuracy and abstention alongside latency. A fast incorrect selection creates rework; removing that rework from the timing would manufacture an improvement.

How to self-host CLM-8B with vLLM

Run two services: a Qwen3-8B pooling encoder and the CLM scoring API. A small projection-head download does not mean a small complete runtime. CLM's package declares Python 3.10 or newer and dependencies including PyTorch, vLLM, FastAPI and NumPy. The pinned package definition records the actual dependencies.

Use an isolated, CUDA-capable Linux environment suitable for your chosen vLLM and PyTorch builds. The CLM source pin below does not pin every dependency or model-weight revision. Resolve those versions for your hardware, then save your environment lock and encoder revision. No minimum VRAM or hardware cost is asserted here.

python3 -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install \
  "git+https://github.com/Contrastive-LM/CLM.git@bb42c6c5bf914fd449bed2f6ca65be80602cb1f7"
python -m pip freeze > clm-environment.txt

Start the encoder in one activated terminal. Use the reference Qwen3-8B encoder and its expected pooling behavior, not an arbitrary embedding model with the same output dimension:

vllm serve Qwen/Qwen3-8B \
  --served-model-name qwen3-8b \
  --runner pooling \
  --max-model-len 2048 \
  --host 127.0.0.1 \
  --port 8090

In a second activated terminal, supply CLM_API_KEY through your secret-management process. Keep that same secret available to the client shell; never commit it. Start a private API with explicitly bounded cache memory:

: "${CLM_API_KEY:?Set CLM_API_KEY in this shell first}"
export CLM_API_KEY
clm-serve \
  --host 127.0.0.1 \
  --port 8700 \
  --emb-url http://127.0.0.1:8090/v1/embeddings \
  --emb-model qwen3-8b \
  --max-tokens 2048 \
  --device cpu \
  --action-cache 64MiB \
  --no-ui

This example intentionally puts the small heads and their vector cache on CPU, while the 8B encoder stays on GPU. It is a deployment illustration, not the configuration behind the headline speedups. Benchmark CPU and GPU heads on your workload before choosing.

The inspected server defaults to 0.0.0.0, makes authentication conditional on CLM_API_KEY, and exposes a health response containing both ok and embedder. The server implementation defines these flags, authentication and health behavior. Explicit loopback binding matters. The CLM key does not automatically secure the separate encoder endpoint, and --no-ui is not authentication. Keep both services private; add authenticated ingress, request limits and workload isolation before remote use.

A successful health HTTP status alone is not sufficient readiness evidence. Check the encoder flag and advertised model too:

curl -fsS --connect-timeout 5 --max-time 15 http://127.0.0.1:8700/health \
  | python -c 'import json,sys; d=json.load(sys.stdin); sys.exit(0 if d.get("ok") and d.get("embedder") and "clm-latest" in d.get("models", []) else 1)'

Send a bounded request without granting tool permissions

Use stable option identifiers and self-contained descriptions. In CLM's Choice implementation, the state encoder receives context plus question instructions. The action encoder receives each option's description, or its key when the description is empty. Two different IDs with the same description are therefore not meaningfully different candidates to that encoder. The schema implementation shows exactly how states and candidates become text.

The following synthetic English request asks for a read-only diagnostic suggestion. The fixture remains identical in each translation so results can be compared. The review candidate is deliberate. Application code still needs a hard fallback because a model is not guaranteed to choose the abstention option when it should.

{
  "model": "clm-latest",
  "state": "A CI job fails during dependency installation. The log reports a lockfile mismatch. No production change is authorized.",
  "questions": {
    "next_step": {
      "type": "choice",
      "instructions": "Select the most useful permitted read-only diagnostic step, or request review when the evidence is insufficient.",
      "criteria": {
        "inspect_lockfile": "Read the manifest and lockfile to identify inconsistent dependency versions; make no changes.",
        "read_network_log": "Read existing dependency-download network logs to investigate connection failures; make no changes.",
        "review": "Ask a human to review because the supplied evidence is insufficient or no listed diagnostic is appropriate."
      }
    }
  }
}

Save it as request.json, then call the API from a shell with the same CLM_API_KEY:

: "${CLM_API_KEY:?Set the same CLM_API_KEY used by the API}"
curl -fsS --connect-timeout 5 --max-time 30 \
  -H "Authorization: Bearer ${CLM_API_KEY}" \
  -H "Content-Type: application/json" \
  --data-binary @request.json \
  http://127.0.0.1:8700/v1/systemone

Treat the response as a suggestion. Validate its model identity, question ID, candidate keys, numeric values and schema before consuming it. A timeout, malformed response, unknown option or missing policy version goes to review. Do not interpolate returned text into a shell command. Map a validated identifier to a separately authorized handler.

When replaying a TypeSafe-shaped request, compare semantic outcomes rather than merely HTTP 200 responses. Re-evaluate thresholds, label descriptions and input preparation. For ordinary business-process examples rather than this CLM-specific contract, use the existing workflow guide linked above.

When does action caching actually help?

Caching helps when the same candidate text is scored against new states. A fixed set of diagnostics, product records or permitted actions can amortize its encoding cost. A fresh set of long generated solutions on every call cannot reuse those action embeddings across unrelated tasks, although state-side or other reuse may still help.

The inspected embedder caches normalized encoder vectors by exact input text in an in-process least-recently-used cache. Requests on misses use the configured pooling endpoint; the token-usage count tracks tokens spent on those misses. The embedder source defines the exact-text cache and token accounting. “Zero new encoder tokens” is not “zero compute,” zero hosting cost or proof that a new user's input was processed independently.

At the engine layer, projected state and action vectors use separate namespaces; the head identity participates in their namespace. The engine separates raw embeddings, projected vectors and model selection. The device arena has a startup allocation and least-recently-used eviction. The vector-cache implementation defines its bounded pools. Its memory budget does not bound the separate encoder-text cache or the entire process. Treat either cache as an implementation detail, not an authorization layer or durable memory store.

Three distinct things need different reuse rules:

CLM cache layers and the conditions for safe reuse
ObjectReuse opportunityWhat must remain valid
Candidate embeddingThe same action description appears againEncoder, pooling, preprocessing and exact text
Projected candidate vectorThe same embedding is scored againAll of the above plus the projection-head identity
Final decisionThe exact decision context recursState, candidate set, permissions, policy, model, temperature and freshness

Our deployment recommendation is to keep a manifest of the encoder revision, pooling mode, head hash, text-rendering version and token limit. Restart and rewarm after changing encoder or preprocessing configuration rather than assuming a head reload invalidates everything. Version the action catalog separately, and never reuse a final decision after a permission or record change.

A shared process cache is not a tenant-isolation guarantee. Decide whether sensitive workloads need separate processes or stronger isolation; include logs, memory, cache timing and request routing in that assessment. Warm only authorized candidate catalogs. Do not embed a global list and assume a high score permits access to every member.

Why can CLM return a confident but wrong choice?

CLM probabilities are relative to the candidates supplied. A softmax must assign its mass somewhere even when all options are poor. With one candidate, the distribution is necessarily 1.0; that is not evidence that the action is correct. Adding plausible alternatives can change probabilities without changing the underlying task.

The published confidence field is the highest probability minus the mean of the other probabilities. It is neither a general factual-verification probability nor a substitute for a calibrated acceptance policy. The formula is visible in the schema source cited above. Treat 0.9 as a model output to validate, not a portable 90% guarantee.

For this CLM integration, test missing-correct-option cases, duplicate descriptions, near-duplicate candidates, contradictory context, candidate ordering and changes in set size. Measure error at the proposed acceptance threshold separately for each relevant language and action family. This article is localized; that does not establish multilingual CLM performance.

Keep permission checks outside the candidate scorer and repeat them immediately before execution. For consequential actions, a fallback is an application decision, not simply another label the same model may ignore. Restrict the initial rollout to suggestions until held-out evidence supports more authority.

The 2,048-token limit and common setup failures

The reference quickstart truncates input texts at 2,048 tokens. Raising only the encoder's maximum does not remove the CLM-side truncation. To evaluate 8K inputs, increase both vllm serve --max-model-len 8192 and clm-serve --max-tokens 8192, then recheck GPU memory and quality. Longer accepted input does not establish that the head was validated for that use.

The inspected embedder sends truncate_prompt_tokens; the state/question rendering and action descriptions count toward what is embedded. Test instructions and critical evidence near the boundary. Reject or deliberately summarize oversized inputs in trusted application code rather than silently treating a truncated decision as complete.

CLM setup symptoms, checks and controlled failure handling
SymptomCheck firstControlled response
HTTP 502 from CLMEncoder URL, model name, process and vLLM errorKeep the request for review; do not substitute approval
HTTP 401Matching CLM bearer keyFix credentials without printing the key
HTTP 422Question type, criteria, model and temperatureValidate the request before retrying
Healthy API, unusable decisionsembedder status and an actual scoring smoke testGate readiness on both services
Slow “warm” requestsCandidate-text changes, cache eviction and fresh statesMeasure cache misses instead of assuming reuse
Good short tests, poor long tasksBoth token limits and critical text placementAdd long-input fixtures before rollout

Do not replace Qwen3-8B with an unrelated encoder because it appears faster. The head is tied to its training representation and pooling. Quantization, alternative runtimes and longer contexts each require separate quality and latency validation.

Use CLM as a verifier, not the code generator

A verifier selects among available solutions; it does not eliminate generation or testing. A candidate pipeline can generate several solutions, run deterministic checks, score remaining candidates, and hand a selected artifact to review. Record candidate-generation cost separately from verifier time. The general architecture belongs in our LLM-as-a-verifier guide; the CLM-specific issue is matching its head and evaluation recipe.

For the published DeepSWE experiment, the head card provides a reproducible command using stored embeddings. Run it from the pinned CLM source checkout with the documented evaluation dependencies and Hugging Face CLI available:

git clone https://github.com/Contrastive-LM/CLM.git clm-source
git -C clm-source checkout bb42c6c5bf914fd449bed2f6ca65be80602cb1f7
cd clm-source
hf download Contrastive-LM/deepswe-clm-heads-8k --local-dir heads/deepswe
python evaluation/bon_eval.py \
  --hf-dataset Contrastive-LM/deepswe-clm-embeddings-8k \
  --checkpoint heads/deepswe/best_head.pt \
  --tasks-file heads/deepswe/heldout_tasks.json \
  --n 4 --window 12

This evaluates a verifier over a published embedding dataset. It does not start a coding agent, regenerate all trajectories or establish the wall-clock cost of solving the original tasks. Verify the published checkpoint and held-out-list hashes, freeze the candidate budget and preserve task-disjoint splits.

The project's fine-tuning instructions explicitly keep the data, folds, evaluation set and best-of-N settings fixed while changing training. The pinned fine-tuning guide defines those experimental boundaries. Fine-tuning small heads may reduce trainable-parameter cost, but data labeling, embedding generation, evaluation and serving remain real work. Do not tune on the held-out tasks and report the resulting score as untouched test performance.

Also check each artifact's own license metadata: the reference CLM weights are Apache 2.0, while the reviewed DeepSWE head card labels its artifact MIT. Do not assume every downstream checkpoint shares one license. “Open weights” does not mean free infrastructure or a warranty of suitability.

Benchmark the full path before switching

A tenfold faster decision component does not make the whole agent tenfold faster. As a deliberately hypothetical budget, let one workflow spend 900 ms outside selection and 100 ms selecting. Replacing 100 ms with 10 ms gives 910 ms total instead of 1,000 ms: 9% less elapsed time, or about 1.10x overall speedup. This arithmetic is not a CLM measurement.

For a CLM pilot, compare fresh states with cold candidates, fresh states with warm candidates, repeated identical requests and changing candidate catalogs. The repeated-request test can measure a different cache path from real work. Keep results separate. Report end-to-end p50/p95, failures, memory use, accepted-decision accuracy and review rate at the same concurrency and candidate budget.

Before promotion, require a versioned encoder/head pair, explicit candidate ownership, tested error handling, monitored cache behavior, a representative held-out evaluation and a rollback path. Keep the existing selector running in shadow comparisons before giving the new service write authority. Use the broader pilot kill-or-scale scorecard for the investment decision rather than creating a second rollout methodology here.

As of the review date, the reference model card describes multimodal CLM-35B as planned for early October 2026. That is a roadmap statement, not a released capability or a guaranteed delivery date. Validate any later checkpoint on its own terms instead of transferring the 8B results to it.

For an implementation assessment, Wavect's AI integration service can address the selector, permissions, observability and evaluation together. The Twinsoft AI case study is related delivery context, not a CLM deployment claim. Bring the pre-launch QA checklist and a representative candidate set to discuss a bounded CLM verifier pilot.

Build the product, not just the backlog

If this article maps to a real product decision, Wavect can help you scope, build, harden, or lead the software work with senior founder-level judgment.

Useful service paths:

CLM-8B deployment questions

Can I run CLM-8B using only the projection-head download?

No. The reference heads depend on Qwen3-8B last-token-pooled embeddings. You still need the encoder service and its memory and compute budget. Putting the heads on CPU does not turn the whole 8B runtime into a small CPU-only model.

Does the CLM action cache also cache tool permissions?

No. It reuses text embeddings and projected vectors, not authorization. Keep permission checks in trusted application code and repeat them before execution. Reuse of a final decision requires the full state, candidate set, policy and freshness conditions to remain valid.

Why are CLM requests truncated even after raising the vLLM context limit?

The reference setup has two limits: vLLM max-model-len and CLM max-tokens. Raise both for a longer-input evaluation. Then check memory, critical text near the boundary and quality; accepting more tokens is not proof of validated long-context behavior.

Is an HTTP 200 from the CLM health endpoint enough for readiness?

No. The inspected response can contain ok=true while embedder=false. Check the encoder status and available model, then run an actual scoring smoke test. Keep the embedding endpoint and the CLM API private; the CLM bearer key does not automatically secure the encoder.

Are the published CLM coding scores zero-shot results?

No. The cited coding results use task-specific fine-tuned heads on held-out subsets: 38 DeepSWE tasks and 30 Terminal-Bench 2.1 tasks. CLM selects among generated candidates. The reference head download alone does not reproduce those scores.

Does CLM confidence of 0.9 mean the action is 90% correct?

Not automatically. Choice probabilities are relative to the supplied candidates. The inspected confidence formula subtracts the mean of other probabilities from the top probability. Validate acceptance thresholds on representative held-out cases and retain a hard review fallback.

Can I replace the Qwen3-8B encoder with another embedding model?

Not as an assumed compatible substitution. The reference heads depend on their training representation and last-token pooling. A different encoder, quantization or input preparation needs separate quality validation; matching vector dimensions alone is insufficient.

Is multimodal CLM-35B available in this guide?

No. As reviewed on 28 September 2026, the model card describes it as planned for early October. The setup here targets the released CLM-v0.1-8B reference head. A future checkpoint needs its own availability, license, runtime and quality checks.

Final thoughts

Treat CLM as a replaceable scoring component with a versioned encoder, explicit candidate contract and measured quality. Reuse vectors where the inputs genuinely repeat, but revalidate decisions and permissions. The first milestone is a reliable private selector on your own held-out tasks, not a headline speedup.

Build the product, not just the backlog

If this article maps to a real product decision, Wavect can help you scope, build, harden, or lead the software work with senior founder-level judgment.

Useful service paths:

Inbox, without the noise

Follow the work that matters to you

Get a short email when we publish something new. Follow the whole blog or only the problems you care about.

What would you like to receive?
Choose your topics

Free, double opt-in, no tracking pixels.

Back
Kevin Riedl

14 min read · 28 Sep 2026
Last reviewed

Next

Get the next AI and agents field note

One concise email when we publish. No tracking pixels, and no inbox filler.

Free, double opt-in, no tracking pixels.