Back
Kevin Riedl

14 min read · 4 Oct 2026
Last reviewed

Next
Made on your device, with no Instagram connection. We copy the post link for Instagram’s Link sticker.

Cloudflare Clef vs Jev: Pricing, Benchmarks and Migration

Cloudflare Clef is a Jev alternative for teams that want typed AI decisions with open weights and multimodal inputs. It is not a general-purpose chatbot. Your application supplies evidence and allowed answers; the model scores those answers instead of writing a free-form response. Cloudflare announced Clef and Clef-flash on October 1, 2026. The official name is Clef, not “Clev”. Cloudflare: Clef launch

Our verdict: Clef deserves a serious pilot when images, deployment control or latency are holding back an existing decision pipeline. It is not an automatic upgrade over Jev, and neither Cloudflare variant has the lower published input-token price. The most important migration issue we found is subtler: matching JSON fields do not guarantee matching confidence semantics.

Reviewed October 4, 2026. This is a primary-source documentation and code review, not a Wavect-run benchmark or a claim that we have deployed Clef for a client. Prices and hosted contracts are a dated snapshot; recommendations below are our engineering assessment.

Clef, Clef-flash and Jev: what actually changes?

Clef’s release uses a Qwen3.8-27B backbone and a joint decision head. The head scores permitted answers across the question schema rather than generating an explanation. Clef-flash uses Qwen3.5-9B. Both Cloudflare releases publish Apache-2.0 weights. Cloudflare: Clef model card and evaluation Cloudflare: Clef-flash model card

TypeSafe currently lists jev-1.13.0, with text-only inputs and a hosted System One API. It distinguishes a 64k total request budget from a 32k budget for the state plus the longest question. Calling Jev simply “32k” and Clef “twice the context” would hide that distinction. TypeSafe: Jev models, prices and limits

Decision factorClefClef-flashJev 1.13
Published input price per million tokens$0.24$0.09$0.042
Deployment routeWorkers AI or self-hosted weightsWorkers AI or self-hosted weightsTypeSafe hosted API
Native visual evidenceYes; hosted and local contracts differYes; hosted and local contracts differNo; preprocess into text
Reason to include in a pilotImage-dependent decisions and control over deploymentLatency-sensitive, narrow decisionsExisting validated behavior and lower listed input cost

The Cloudflare rates above are from its Workers AI pricing table, not a promotional estimate. The pilot reasons are our recommendations, not benchmark guarantees. Cloudflare: Workers AI pricing

For the underlying Jev interface, use our Jev technical review. For the separate open-model and benchmark-provenance discussion, see Laya versus Jev. Here the question is whether replacing an existing Jev integration with Cloudflare Clef is worth the migration risk.

The useful shift is not “another chatbot”. Think of a returns workflow that has a customer message and a photo of the damaged package. A Clef pilot could assess that evidence together and choose among damage_review, delivery_review and manual_review. A text-only route needs an explicit image-to-text stage first. This is an illustrative design, not a measured Clef deployment. The refund itself still belongs to your application’s policy and approval flow.

The typed interface supports noul for a yes/no probability, choice for named alternatives and score for ordered levels. That makes Clef interesting as a decision step inside an agent, not a replacement for every reasoning or writing step. Published input types.

Is Cloudflare Clef cheaper than Jev?

Not at the published input-token rates. Consider one million decisions, each billed for 2,000 input tokens. That is two billion billable tokens. Applying the Cloudflare rates and TypeSafe rate gives this inference-only illustration:

ModelIllustrative input chargeRelative to Jev
Jev$841.00×
Clef-flash$1802.14×
Clef$4805.71×

This is our arithmetic, not a billing quote. It assumes the same billable input-token count, not identical tokenization of the same raw text. Include schema overhead, retries and image processing in real measurements. Free allocations, discounts, Workers execution, storage and human review are outside this example. Jev lists free output tokens; Cloudflare lists no output-token rate for these models.

A higher token price can still make sense when the new route reduces expensive errors, avoids a separate image-to-text step or finishes a time-sensitive decision reliably. Establish that from the full pipeline. Do not turn a cheaper-than-a-large-chat-model argument into a cheaper-than-Jev claim. Our decision-model workflow and ROI guide covers the broader business case.

Benchmarks: Clef wins some tasks, Jev wins others

Cloudflare’s published results support task-specific evaluation, not a universal winner. The following rows come from the Clef model card’s Decision Index 0.2.1 run. They are Cloudflare-reported results, not independent measurements by Wavect. Quality scores are percentages; higher is better. Lower latency is better.

Benchmark / metricClefClef-flashJev
BANKING77, macro F194.290.979.7
CLINC150 + OOS, macro F197.466.889.3
RAGTruth, hallucination F179.435.676.5
GPQA Diamond, accuracy48.051.078.3
When2Call MCQ, accuracy72.465.681.0
Median latency, milliseconds209.338.8524.1
p95 latency, milliseconds238.6122.4536.0

Our interpretation: include Clef in an intent-routing pilot, but test unknown intents rather than only familiar categories. Treat Flash as a different quality profile, not merely the same model running faster. Its CLINC150 + OOS and RAGTruth results are material counterexamples to a blanket replacement claim. Jev’s GPQA and When2Call results also make “Clef is better at every decision” indefensible.

The latency rows are properties of that evaluation setup. They do not establish your production response time, an SLA or the performance of a self-hosted GPU. Measure end-to-end p95 under your concurrency, request sizes, network path and fallback behavior. Do not silently relabel the benchmark’s “Jev” column as a newly tested version that the table does not identify.

The migration trap: confidence is not the same contract

Do not copy Jev confidence thresholds directly into Clef. TypeSafe defines Choice confidence as (p_max - 1/n) / (1 - 1/n), where n is the number of options. This differs from simply returning the winning probability. TypeSafe: confidence formulas

In Cloudflare’s reviewed release code, systemone_answer sets Choice confidence to the winning probability itself. We inspected repository revision a20e258. This is a finding about the published implementation, not proof that every hosted deployment uses that exact revision. Confirm the hosted behavior separately. Cloudflare: reviewed Clef release code

Our worked example uses three options with probabilities 0.80, 0.15, 0.05. Jev’s documented formula returns 0.70; the reviewed Clef helper returns 0.80. An existing confidence >= 0.75 rule would reject the former and accept the latter, even though the distribution is identical. This is a contract difference, not evidence that Clef is more certain or more accurate.

Preserve the returned probability distribution and define the decision statistic explicitly in your own adapter. Refit thresholds against labeled validation cases, then evaluate accepted errors and review coverage on an untouched test set. A consistent formula alone does not make probabilities equally calibrated. Our AnyJev calibration and option-order guide covers that separate evaluation problem.

There is another reason to test the whole request: TypeSafe describes independent question evaluation, while Clef’s release scores the schema jointly. Our resulting test recommendation is to repeat decisions with the full production question bundle and with irrelevant questions added or removed. That recommendation is an architectural inference, not a claim that we observed a Clef defect. TypeSafe: independent question evaluation

Migrating the API: keep the schema, verify the boundary

Jev’s documented endpoint is POST https://api.typesafe.ai/v1/systemone. Matching state, questions and answer IDs is useful, but a Cloudflare migration still needs new credentials, a model selector and an adapter around the actual transport. TypeSafe: System One HTTP API

The hosted Clef contract documents a 65,536-token context, 1 to 64 questions and truncation of long text state. Image input accepts up to four embedded PNG, JPEG or WebP images, not remote image URLs. The documented limits are 4 MiB and 16 megapixels per image, 8 MiB decoded in total and 13 MiB for the request body. Cloudflare: hosted Clef API contract

Clef-flash has its own route and matching model: "clef-flash" selector. Keep endpoint and selector aligned when switching variants. The documented hosted image contract should not be confused with the local release’s video-frame examples. Cloudflare: hosted Clef-flash API contract

A smoke test without business actions

Save the following as clef-request.json. The example classifies a fictional message into a review queue. It does not issue refunds, modify accounts or authorize an action.

{
  "model": "clef",
  "state": {
    "message": "The package arrived with a broken seal. I am not sure whether anything is missing."
  },
  "questions": {
    "review_queue": {
      "type": "choice",
      "instructions": "Select a review queue. Use manual_review when the evidence does not support a specific queue.",
      "criteria": {
        "damage_review": "Possible physical package damage.",
        "delivery_review": "Delivery timing or location issue.",
        "manual_review": "Missing or ambiguous evidence; a person should review."
      }
    },
    "needs_more_evidence": {
      "type": "noul",
      "instructions": "Is more information needed before a refund decision can be considered?"
    }
  }
}

Set an account ID and a server-side Cloudflare API token with Workers AI permission in your shell. Then send the request using curl 7.76 or newer:

set -eu
: "${CLOUDFLARE_ACCOUNT_ID:?Set your Cloudflare account ID}"
: "${CLOUDFLARE_API_TOKEN:?Set a server-side Workers AI token}"

curl --fail-with-body --silent --show-error \
  --connect-timeout 10 --max-time 60 \
  "https://api.cloudflare.com/client/v4/accounts/${CLOUDFLARE_ACCOUNT_ID}/ai/run/@cf/cloudflare/clef" \
  --header "Authorization: Bearer ${CLOUDFLARE_API_TOKEN}" \
  --header "Content-Type: application/json" \
  --data-binary @clef-request.json \
  --output clef-response.json

This uses the documented REST route. We checked the JSON and shell syntax locally, but did not execute a paid inference request. Inspect clef-response.json, including API-level errors. Do not wire actions directly to raw output. Test the REST response envelope separately from a Workers binding result; validate answer IDs, allowed choices and finite probabilities in the adapter.

Retain manual review for unknown labels, malformed responses, timeouts and rate limits. A failed image-based Clef request must not silently fall back to text-only Jev with the image evidence removed. Use an explicitly tested text representation, or stop for review. Likewise, reject or split oversized evidence deliberately rather than discovering after an incident that the decisive paragraph was truncated.

Open weights, self-hosting and the privacy boundary

Open weights create a deployment option, not a finished production service. The Clef release instructions load both the backbone and the joint decision head. Serving only a generic chat-generation path is not evidence that you are running the released decision model. Pin the weights, processor, head and serving code together.

The reviewed local helper defaults to max_length=16384. That is distinct from the hosted 65,536-token limit. The model card’s single-H200 test setup is a tested environment, not a stated minimum GPU requirement. For rough capacity planning, 27 billion parameters at two bytes each are about 54 GB of weights; 9 billion are about 18 GB, before runtime overhead. These are arithmetic estimates, not measured memory requirements or proof that a given quantization fits your hardware.

Also separate downloadable weights from managed customization. Cloudflare’s launch announcement describes initial fine-tuning work through its forward-deployed engineering team, with broader self-service access planned. Do not scope a customer project around an already available one-click fine-tuning product without confirming access.

Cloudflare’s Workers AI data policy says customer content is not used for training or improving models without explicit consent. That does not answer every question about your application’s persistence or processing location. Cloudflare: Workers AI data usage

For example, AI Gateway logs are enabled by default and can include prompts and responses. Review logging, storage, access and retention across your own stack. Do not turn a “not used for training” statement into an EU-only inference or zero-retention guarantee. Confirm the contractual and technical requirements for the actual route you deploy. Cloudflare: AI Gateway logging defaults

A Jev-to-Clef migration checklist that catches real regressions

Start with one reversible decision. Freeze the old implementation as the baseline and run Clef in shadow mode, without changing live actions. Define acceptance limits before inspecting the final holdout. A useful release record should answer these questions:

GateEvidence to retainDo not promote when
API and schemaActual HTTP fixtures, allowed choices, error cases and variant configurationParsing succeeds only for the happy path
Decision qualityAccepted errors, review rate and rare-class results by deployed languageAggregate accuracy hides harmful accepted decisions
Confidence contractExplicit statistic, versioned thresholds and untouched test resultsA renamed provider silently changes acceptance policy
Evidence integrityLong-state, image and full-question-bundle testsTruncation or fallback drops required evidence
OperationsEnd-to-end p95, retries, metered cost and tested rollbackThe gain disappears under representative load

Keep authorization, idempotency and transaction checks in deterministic application code. A schema-conforming answer is still a model judgment, not permission to act. During rollout, record the selected deployment, schema revision, acceptance policy and result so that a later change can be investigated without logging unnecessary personal data.

Our recommendation is to choose Clef when a measured capability or deployment benefit pays for the migration, choose Flash only when its task-specific quality is acceptable, and keep Jev when its validated behavior already meets the requirement more economically. Remaining on the existing model is a valid engineering outcome.

Wavect’s AI engineering service can help turn the comparison into an evaluation and integration plan. Use the pre-launch QA checklist to structure the release, then discuss your decision pipeline with representative, redacted examples.

For related AI delivery context, see the Twinsoft AI case study. It is not evidence of a Clef deployment.

Clef migration questions

Is Cloudflare Clev the same as Clef?

The official product is Cloudflare Clef. Clef and Clef-flash were announced on October 1, 2026. Use Clef when searching the documentation, model weights and Workers AI endpoints. Launch source.

Is Clef a drop-in replacement for Jev?

It uses a compatible typed request shape, but transport, model selection, input limits and confidence semantics still need testing. In particular, the reviewed local Clef implementation does not compute Choice confidence using Jev’s documented formula. API compatibility is not policy equivalence. Migration finding.

Is Clef cheaper than Jev?

Not per published input token on October 4, 2026: Clef is $0.24, Clef-flash $0.09 and Jev $0.042 per million input tokens. Overall economics depend on measured errors, review, retries and infrastructure, not only token price. Calculation and assumptions.

Should I use Clef or Clef-flash?

Pilot Clef for the quality-sensitive or image-dependent decision you actually need. Evaluate Flash separately for latency-sensitive work. Cloudflare’s published results show substantial task-specific differences, so Flash should not inherit Clef’s acceptance thresholds or approval. Benchmark counterexamples.

Does Clef support images and video?

The hosted schema documents embedded images with count and size limits, while the local model release also shows video-frame inputs. Those are different interfaces. Confirm the contract for your deployment instead of assuming a local video example works through the hosted image endpoint. Hosted input limits.

Does Clef have twice Jev’s context window?

That description is misleading. Hosted Clef documents 65,536 tokens. Jev documents 64k for the total request and 32k for the state plus its longest question. The reviewed local Clef helper separately defaults to 16,384 tokens. Evaluate your actual request shape. Context distinction.

Can I self-host Clef?

Cloudflare publishes Apache-2.0 weights, a joint decision head and release code. Self-hosting needs the matching components, capacity planning and an operational serving layer. The published H200 test setup is not a minimum-hardware guarantee. Deployment boundaries.

Does a high confidence score authorize an action?

No. Confidence summarizes a model output, not permissions or proven correctness. Validate accepted-decision error, retain a review path and enforce authorization and transaction rules outside the model. Migration gates.

Final thoughts

Cloudflare Clef makes the Jev-style decision interface more interesting through open weights and visual inputs. The reason to switch is a measured improvement in your pipeline, not a headline. Revalidate confidence, preserve evidence and keep a tested rollback before promoting either Clef variant.

Build the product, not just the backlog

If this article maps to a real product decision, Wavect can help you scope, build, harden, or lead the software work with senior founder-level judgment.

Useful service paths:

Inbox, without the noise

Follow the work that matters to you

Get a short email when we publish something new. Follow the whole blog or only the problems you care about.

What would you like to receive?
Choose your topics

Free, double opt-in, no tracking pixels.

Back
Kevin Riedl

14 min read · 4 Oct 2026
Last reviewed

Next

Get the next AI and agents field note

One concise email when we publish. No tracking pixels, and no inbox filler.

Free, double opt-in, no tracking pixels.