---
title: "AnyJev: LLM Calibration and Option-Order Bias"
canonical: https://wavect.io/blog/anyjev-calibration-option-order-bias/
language: en
description: "Understand AnyJev L0, L1 and L2: option-order bias, probability calibration, skewed labels and a source-checked routing evaluation with clear benchmark limits."
image: "https://wavect.io/img/blog/headers/header_anyjev-calibration-option-order-bias.png"
---

[**Back**](/blog/overview/)

[![Kevin Riedl](/img/team/kevin.webp)](/team/kevin-riedl/)

[Kevin Riedl](/team/kevin-riedl/) https://linkedin.com/in/wsdt

16 min read · 28 Sep 2026 Last reviewed September 28, 2026

[**Next**](/blog/llm-evaluation-cost-roi-production/)

# AnyJev: LLM Calibration and Option-Order Bias

TL;DR

AnyJev turns existing language models into typed decision components. L0 needs no labels and reduces order and prior bias; L1 fits confidence on labeled examples; L2 fits a per-question head without changing the backbone. The current 0.2.0 release adds optional canonical ordering and adaptive rotations. Its 1% adaptive target measures agreement with its own full-cycle answer, not real-world accuracy. This guide separates those claims, examines negative results and provides an offline calibration-to-evaluation example.

**Reviewed 28 September 2026.** Package baseline: AnyJev 0.2.0. Source snapshot: `10d5db91dda38dbde74c6abc1c075ce6463723d1`. This is a documentation and source-code review, not a GPU benchmark or a report of a Wavect production deployment. [AnyJev 0.2.0 release](https://pypi.org/project/anyjev/0.2.0/)

## What is AnyJev, and what does “no training” mean?

**AnyJev is an open-source library that reads typed decisions from an existing language model instead of asking it to write an answer. Its distinctive problem is not just speed: it addresses option-order bias and separates debiased scores from confidence calibrated on labels.** The project credits Jiamu Zhang, Tianze Yang, Yucheng Shi and Liang Wu, with Nokia and Tencent Hunyuan affiliations. It is inspired by TypeSafe AI's Jev interface, not affiliated with TypeSafe. [Project overview and attribution](https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/README.md)

The package is released under Apache 2.0. That is not a blanket license for every model or dataset used with it. [Package license metadata](#source-pypi).

The practical question is straightforward. Your agent already has a ticket, a document or a tool result. It must choose one of several permitted answers. Can you trust that choice when someone changes the order of the options? And does a confidence of 0.9 mean that comparable decisions are actually correct about 90% of the time?

Those are different problems. **A stable answer can be wrong. A normalized distribution can be overconfident.** AnyJev provides different processing levels rather than one universal fix.

“Training-free” is accurate for its label-free L0 path. L1 fits a temperature using labeled examples. L2 fits a supervised linear head, although it does not update the backbone or use gradient descent for that fit. Calling all three “no training” hides the data and validation work. The more useful distinction is **no backbone fine-tuning versus no labeled fitting**. [Level-by-level contract](https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/docs/levels.md)

For the general product landscape, use our [Jev review](/blog/jev-ai-decision-model-review/) and [Laya versus Jev benchmark analysis](/blog/laya-vs-jev-benchmark-ai-startup-moat/). This guide concentrates on AnyJev's calibration, ordering and evaluation mechanics, not another decision-model ranking.

## Raw, L0, L1 and L2: which path needs labels?

| Level | Data and mechanism | What it does not establish |
| --- | --- | --- |
| raw | No labels. Read a distribution restricted to the answer-label tokens. | Neither order robustness nor calibrated uncertainty. |
| L0 | No labels. Combine option rotations and optionally correct the label prior. | A confidence value is not a validated probability of correctness. |
| L1 | Typically 100–500 labels for the question. Fit temperature scaling above L0. | No protection against new data distributions or incorrect model knowledge. |
| L2 | Typically 100–300 labels per question. Fit a closed-form head on hidden states. | No general transfer to an unrelated question or another base model. |

The ranges are project guidance, not sample-size guarantees. A task with rare but costly mistakes can require much more evidence. Likewise, `auto` means “use the best available level,” not “automatically safe.” Check the level actually returned. The table summarizes the [documented level contract](#source-levels).

## Why rearranging the same options changes an LLM's answer

Consider three semantic answers: billing, technical support and sales. A naïve implementation maps them to A, B and C, runs a next-token readout and selects the highest score. But a preference for “A,” an early position or a neighboring option can become a preference for whichever business answer occupies that position.

This is an established research problem, not a problem first discovered by AnyJev. Zheng and colleagues examined the sensitivity of multiple-choice LLM selection to option ordering. Their work is part of the foundation for the project's permutation approach. [Research on multiple-choice order sensitivity](https://arxiv.org/abs/2309.03882)

AnyJev normally evaluates the cyclic rotations of a K-option choice. Each option appears in every position once. The default `logmean` method maps the probabilities back to semantic options, averages their logarithms, exponentiates and renormalizes. Under an idealized additive logit-position bias, the position term cancels. [Rotation and marginalization implementation](https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/anyjev/calibrate/permute.py)

**That idealized cancellation is not a proof of invariance under every real prompt permutation.** A full cycle preserves which options are neighbors; actual model responses can depend on those interactions. The previously published BANKING77 results still contain residual answer flips.

Version 0.2.0 offers a separate, opt-in remedy: `canonical_order=True` sorts options by their text before rotating. The same option set then produces the same prompt layouts regardless of caller order. This is a deterministic input-normalization property, assuming the same inference and prior state. It does not make the answer correct, eliminate numerical variation or make a running batch prior independent of request history. The [rotation-budget documentation](#source-rotation) explains why canonicalization also matters when only some rotations are read.

When testing reorder robustness, compare returned **semantic labels**, not option indices. Index zero means something different after reversing the list. Also distinguish disagreement among individual rotations from disagreement between two complete calls with reordered inputs.

## When batch-prior correction helps, and when it hurts

L0 can correct a label prior estimated from model predictions on unlabeled inputs. Conceptually, it divides each answer score by an estimate of how often that answer is favored before renormalizing. The default batch-prior strength is 0.75, with correction starting after the per-question accumulator has at least eight inputs. Before that, diagnostics report no prior correction. [Decider defaults, calibration and state handling](https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/anyjev/decider.py)

The catch: a common prediction may represent **a biased model or a genuinely common class**. Unlabeled prediction frequencies alone cannot reliably distinguish the two. For a support queue where most requests truly belong to billing, correcting that majority toward balance can damage the router.

The project's diagnostic study covers 230 model-question combinations. It reports useful gains on balanced 20-way tasks, but also substantial losses on some skewed questions. It retains `prior="none"` to isolate the rotation effect when the label marginal is known to be skewed. That is an ablation to evaluate, not a promise that switching the prior off always wins. [Positive and negative L0 results](https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/docs/when_l0_helps.md)

The underlying batch-calibration approach comes from Zhou and colleagues. AnyJev packages it with rotations and operational interfaces; it does not invent all of these methods from scratch. [Batch Calibration research](https://arxiv.org/abs/2309.17249)

There is also content-free calibration using probes such as an empty input or `N/A`, following the approach of Zhao and colleagues. But a model's answer to an empty input can itself have meaning. Treat `prior="content_free"` as another measured variant, not a universally cleaner baseline. [Calibrate Before Use](https://proceedings.mlr.press/v139/zhao21c.html)

For a routing pilot, compare raw, rotation-only L0 and default L0 on both a class-balanced diagnostic set and a time-based sample of natural traffic. The former helps expose weak classes; the latter tells you what the system will encounter.

## What L1 calibrates, and why temperature is not an accuracy upgrade

Temperature scaling changes the concentration of a fixed probability vector. Conceptually, it applies `softmax(log(p) / T)` with positive T. A larger temperature generally softens confidence; a smaller one sharpens it. AnyJev fits T by minimizing negative log likelihood on labeled calibration examples and stores the prior used for that calibration in the artifact. [Temperature-scaling implementation](https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/anyjev/calibrate/posthoc.py)

For a fixed vector, positive temperature scaling preserves the winning class. **It does not turn an incorrect class into a correct one.** The calibration literature distinguishes the quality of probability estimates from classification accuracy; Guo and colleagues provide the foundational treatment used here. [On Calibration of Modern Neural Networks](https://proceedings.mlr.press/v70/guo17a.html)

Why, then, can AnyJev's L0 and L1 accuracy columns differ slightly? The complete pipeline is not only a temperature: L1 freezes the prior estimated from its calibration states, while ordinary L0 can accumulate a different running prior. Do not attribute the difference to temperature changing the argmax. This follows from the [Decider implementation](#source-decider) and the monotonic transform above.

A useful deployment artifact therefore needs more than a temperature value. Our recommendation is to record the base-model revision, tokenizer, question wording, exact option layout, prior configuration, package version and calibration-data window. A changed question or serving configuration is a reason to revalidate, not quietly reuse a favorable confidence threshold.

## What the published benchmarks actually show

The following results are **project-reported** for Qwen3-8B on a **20-class subset of BANKING77 with 300 test items**. They are not a full 77-class benchmark, a Nokia production SLA, or a measurement of the newer opt-in canonical/adaptive configuration. [Committed benchmark tables](https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/docs/results_bench.md)

| Metric | raw | L0 | L1 |
| --- | --- | --- | --- |
| Accuracy | 74.7% | 80.3% | 80.7% |
| Expected calibration error, lower is better | 0.240 | 0.184 | 0.095 |
| Answer-flip rate under reversal | 23.0% | 7.3% | 7.7% |
| Empirical coverage at 5% error | 7.7% | 46.3% | 52.0% |

The gains are worth investigating. But the strongest qualification is in the last row: coverage is a retrospective property of this confidence-sorted test sample, not proof that a fixed threshold will automate 52% of future traffic at 5% error. At this sample size, 52% represents roughly 156 decisions. One additional error among 156 accepted decisions changes their observed error rate by about 0.64 percentage points. That arithmetic illustrates uncertainty; it does not reconstruct the unpublished error count.

There is an equally useful counterexample in the same table source. On Qwen3-8B newsgroups, L1 improves ECE from **0.309 to 0.138**, while empirical coverage at 5% error falls from **42.3% to 23.7%**. Better average calibration does not automatically create better low-risk selection. Evaluate both probability quality and the risk-coverage curve. [Benchmark source](#source-bench).

L2 results need another distinction. The repository reports 77.1% for Qwen3-8B on 2,000 typed decisions after per-question labeled fitting. In that dataset, “accuracy” is agreement with teacher-derived labels, not an independent human adjudication of every business outcome. The README marks third-party Jev comparison rows as published elsewhere, not rerun head-to-head. Do not turn those rows into an overall product winner. [Project limitations](#source-readme).

## Adaptive rotations: a 1% target is not 99% accuracy

Adaptive rotations try to stop before all K layouts are evaluated. Calibration uses unlabeled states because the reference answer is the model's own full-cycle result. A stopping threshold is selected using an upper confidence bound on disagreement with that reference. **This estimates agreement with another readout, not correctness against the real task.** [Rotation-budget experiments and limitations](https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/docs/rotation_budget.md)

In one reported Qwen2.5-7B, 18-option routing experiment on an H100 NVL with vLLM 0.7.0, full rotations delivered 16.73 decisions/s. Adaptive waves of two delivered 37.20 decisions/s using 7.28 requests per decision, a 2.22× measured throughput ratio. However, reference agreement was **98.7%**, or **1.3% disagreement**, outside the nominal 1% target. The authors disclose this result. It is not a contractual bound on future requests.

The practical sequence is to establish a full-rotation baseline, enable canonical ordering, calibrate the adaptive budget on separate representative states, then measure both reference disagreement and ground-truth errors on held-out data. Small calibration sets, long contexts and different serving engines can change the result. Do not translate fewer rotations directly into the same factor of end-to-end agent acceleration.

## A reproducible starting point for a routing experiment

The following is an illustrative offline workflow, not a measured deployment. AnyJev 0.2.0 requires Python 3.10 or newer. Install into an isolated environment and lock the resolved dependencies for your own reproduction; a package pin alone does not pin PyTorch, Transformers or model weights. [Package baseline](#source-pypi).

```
python -m venv .venv
. .venv/bin/activate
python -m pip install "anyjev[hf]==0.2.0"
```

`Question.choice` accepts two to 26 unique option strings in the current implementation. Labels supplied to calibration are integer indices into those options. Keep a stable semantic mapping, and avoid overlapping categories that make even a human annotator guess. [Typed-question validation and hashing](https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/anyjev/question.py)

For this example, use a compatible CUDA environment with enough memory for the chosen base model. The HF backend loads the model and tokenizer, accepts a model revision and does not enable remote model code by default. A small AnyJev package is not a small replacement for the backbone. [HF backend and model loading](https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/anyjev/backends/hf.py)

Set `ANYJEV_MODEL_REVISION` to the actual 40-character commit of the chosen Qwen3-8B model before running the Python blocks in order. It is a model revision, not the AnyJev source pin.

```
import os
import re
from anyjev import Decider, Question
from anyjev.backends.hf import HFBackend

## Set this to the actual Qwen model commit, not the AnyJev commit.
revision = os.environ.get("ANYJEV_MODEL_REVISION", "")
if not re.fullmatch(r"[0-9a-f]{40}", revision):
    raise ValueError("Set ANYJEV_MODEL_REVISION to a Qwen3-8B commit SHA")

backend = HFBackend(
    "Qwen/Qwen3-8B", revision=revision,
    device="cuda", dtype="bfloat16",
)
route = Question.choice(
    "Which internal queue should review this ticket? Classify only.",
    ["Billing", "Technical support", "Sales", "Manual review"],
    name="route",
)
d = Decider(
    backend, level="L0", prior="none",
    canonical_order=True, adaptive_shifts=False,
)
result = d.decide("I need a copy of my invoice.", [route])["route"]
print(result.level, result.distribution)  # Inspection only; no dispatch.
```

The example deliberately uses full rotations, canonical ordering and `prior="none"` to establish a permutation-only baseline. It does not silently rely on a warm running prior. Compare it with the default batch correction before choosing either configuration for real traffic. The English question and labels stay identical across the translations of this article so the example remains the same experiment.

Next, collect **separate** calibration and held-out evaluation files. Each non-empty JSONL line must contain `id`, `state` and a `label` matching one of the option strings. The illustrative row below documents the format; it is not a sufficient calibration dataset.

```
{"id":"ticket-example-001","state":"Please resend my invoice.","label":"Billing"}
```

```
# Continue after the setup above. Supply your own labeled JSONL files.
import json
from pathlib import Path

def load_rows(filename: str, minimum: int = 1) -> list[dict]:
    rows, ids, states = [], set(), set()
    for line_no, line in enumerate(Path(filename).read_text(encoding="utf-8").splitlines(), 1):
        if not line.strip():
            continue
        row = json.loads(line)
        if not isinstance(row, dict) or not all(
            isinstance(row.get(k), str) and row[k].strip()
            for k in ("id", "state", "label")
        ):
            raise ValueError(f"{filename}:{line_no}: id, state and label must be strings")
        if row["label"] not in route.options:
            raise ValueError(f"{filename}:{line_no}: unknown label")
        if row["id"].strip() in ids or row["state"].strip() in states:
            raise ValueError(f"{filename}:{line_no}: duplicate id or state")
        ids.add(row["id"].strip())
        states.add(row["state"].strip())
        rows.append(row)
    if len(rows) < minimum:
        raise ValueError(f"{filename}: need at least {minimum} rows for this example")
    return rows

calibration = load_rows("calibration.jsonl", minimum=100)
heldout = load_rows("heldout.jsonl")
for field in ("id", "state"):
    if {r[field].strip() for r in calibration} & {r[field].strip() for r in heldout}:
        raise ValueError(f"Calibration and held-out {field} values overlap")
if {r["label"] for r in calibration} != set(route.options):
    raise ValueError("Calibration must cover every route in this example")

artifact = d.calibrate(
    route, [r["state"] for r in calibration],
    [route.options.index(r["label"]) for r in calibration], level="L1",
)
# Exclusive creation prevents accidental overwrite of an existing artifact.
with Path("route-calibration.json").open("x", encoding="utf-8") as f:
    json.dump(artifact, f, ensure_ascii=False, indent=2)

predictions = d.decide_batch(
    [r["state"] for r in heldout], route, level="L1", require="L1",
)
for row, prediction in zip(heldout, predictions):
    print(json.dumps({
        "id": row["id"], "expected": row["label"],
        "predicted": prediction.argmax, "confidence": prediction.confidence,
        "level": prediction.level,
    }, ensure_ascii=False))
```

This code fits on one split and only scores the other. It neither chooses an automation threshold nor dispatches a ticket. The 100-row guard is an example policy, not a claim that 100 labels are statistically sufficient. Reject duplicate or overlapping inputs and check temporal or customer-level leakage as well as identifiers.

`require="L1"` prevents silent use of a weaker result; `confidence` is the maximum returned probability. Neither is a permission check or a guarantee of an acceptable business error rate. The result object also exposes the actual level and diagnostics. [Decision fields and level enforcement](https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/anyjev/result.py)

For production threshold selection, add a separate validation split. Fix the acceptance rule there and evaluate it once on the untouched test split. Report accepted counts and errors per class, not only a single aggregate confidence score.

## Serving and artifact boundaries that can break the experiment

AnyJev's vLLM adapter has two different paths. Raw, L0 and L1 use a generation server and request `max_tokens=1` with allowed answer-token IDs and log probabilities. There is no free-form answer to parse, but the adapter does request one token; “no generation” should not be read as zero token work at every API boundary. L2 instead needs unnormalized last-position hidden states from a pooling server. [Serving adapter and endpoint contracts](https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/anyjev/backends/vllm.py)

A generic embeddings endpoint is not automatically interchangeable with the hidden-state interface. An ordinary vLLM pooling server exposes the final layer of its served checkpoint; a head fitted to an earlier layer needs a matching truncated checkpoint or a backend exposing that layer. Validate model identity, tokenizer, vector shape and preprocessing together.

For L2, the project uses a per-question linear head fitted by a closed-form solve, with held-out folds guiding head and layer selection. That is supervised learning even though no backbone gradient update occurs. A successful fit on one question does not validate another question with similar-looking labels. [Closed-form head fitting](https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/anyjev/heads.py)

Two common operational mistakes are requesting L1 before loading or fitting its artifact, and fitting an artifact but continuing to call the default L0 path. Make `level` and `require` explicit. Keep the deployment configuration immutable during evaluation and record any fallback, prior warm-up or artifact mismatch as an event, not an invisible change in accuracy.

## An acceptance test for a real decision pipeline

Our recommendation is to evaluate a narrow, reversible decision first. Build the test around the cost of an incorrect accepted decision, not around the novelty of a model interface.

**Test correctness and stability separately.** Measure class accuracy, macro F1, option-reversal disagreement and several non-cyclic reorderings. Reuse an identical prior state for each comparison. A stable answer is not necessarily a correct answer; a changed running prior is not necessarily an option-order defect.

**Test confidence and coverage separately.** Record NLL, Brier score, ECE, accepted-decision error and coverage at a preselected threshold. Inspect rare classes and each deployed language. The newsgroups counterexample above is a concrete reason not to optimize ECE alone.

**Test costs at the actual boundary.** Measure prompt count, p50/p95 latency, queueing, prefix-cache behavior and fallback rate. Include model serving and human review, not just the few arithmetic operations after logits arrive. Compare against your existing router and a simple rule-based baseline.

**Keep execution controls outside the classifier.** Restrict the candidate set to authorized actions, include a review or abstention path, and let deterministic business rules enforce permissions. A confidence threshold must not authorize a refund, account change or production command. Observe drift and recalibrate before expanding scope.

The repository still lists agent-loop evaluation as future work, so its isolated-decision results are not evidence that a complete production agent is already better. [Project scope](#source-readme). For the wider rollout process, see our [LLM evaluation guide](/blog/llm-evaluation-cost-roi-production/) and [pilot kill-or-scale scorecard](/blog/ai-pilot-kill-or-scale-scorecard/). Business workflow economics remain in the [Laya/Jev ROI guide](/blog/laya-jev-business-workflows-roi/).

For the separate encoder-and-head deployment problem, see our [CLM-8B self-hosting and action-cache guide](/blog/clm-8b-self-hosting-action-cache-verifier/). It covers candidate-vector reuse and verifier setup, while the evaluation here focuses on calibration and option order.

A useful AnyJev pilot ends with a versioned question, a reproducible probability pipeline and an untouched test result. That is a stronger basis for automation than either a confident answer or a faster next-token readout.

For implementation support, explore Wavect’s [AI engineering service](/services/artificial-intelligence/). The [Twinsoft AI case study](/case-studies/twinsoft-ai/) provides separate delivery context, not an AnyJev deployment claim. Bring representative decisions and the [pre-launch QA checklist](/software-development-guide/software-qa-checklist-before-launch/) to [discuss an evaluation-first integration](/contact/).

## AnyJev calibration questions

### Is AnyJev really training-free?

L0 needs no labels or backbone fine-tuning. L1 fits a temperature on labels, and L2 fits a supervised head. No gradient update to the backbone is not the same as no labeled learning. [Level contract](#source-levels).

### Does AnyJev completely remove option-order bias?

Full rotations remove an idealized additive position effect, but real prompts can retain interaction effects. Optional canonical_order=True normalizes caller ordering in 0.2.0. Compare semantic answers under fixed inference and prior state; stability does not prove correctness. [Ordering details](#source-rotation).

### Can I calibrate AnyJev without labeled examples?

You can estimate a prior and calibrate adaptive agreement with the full-cycle readout without labels. Validating confidence against correct answers is different: L1 requires labels for the question. The adaptive 1% target is not a ground-truth error guarantee. [L0 and L1](#source-levels).

### Why can lower ECE still mean worse automation coverage?

ECE summarizes calibration, not the quality of the lowest-risk subset at every threshold. In the published Qwen3-8B newsgroups rows, L1 lowers ECE but reduces empirical coverage at 5% error. Measure both probability quality and accepted-decision risk. [Reported counterexample](#source-bench).

### Does AnyJev work with any hosted LLM API?

Not automatically. The reviewed vLLM logit path needs allowed-token control and answer-token log probabilities; L2 needs compatible hidden states. A generic chat or embeddings API is not sufficient evidence of compatibility. [Backend contract](#source-vllm).

### How many options does AnyJev support?

The current Question.choice constructor accepts 2 to 26 unique options. Ordinal score questions use 2 to 10 bins or levels. Larger candidate catalogs need another design or a later implementation, not an assumed bypass. [Question validation](#source-question).

### Does require="L1" make a decision safe to execute?

No. It enforces a minimum processing level, not permissions or validated business risk. Calibrate on representative data, choose a threshold on validation data and evaluate it on an untouched test set. Keep a separate review path and deterministic authorization. [Level enforcement](#source-result).

### Can I reuse an AnyJev L2 head for a different question?

Do not assume that a head transfers to an unrelated question or another model. AnyJev fits heads per question and model. Its same-question routing across wording or order changes still needs validation on the new inputs. Similar category names alone do not establish compatibility. [Head scope and limits](#source-levels).

## Final thoughts

Use AnyJev to test a specific decision interface, not to assume that confidence equals correctness. Separate order robustness, probability calibration and accepted-decision risk. Freeze the question and serving contract, preserve an untouched test split and keep execution permissions outside the model.

## You may also like..

[**When is an LLM evaluation worth building?** Budget for representative evidence, valid metrics and ongoing maintenance.](/blog/llm-evaluation-cost-roi-production/) [**Jev: the technical review** Understand the product that inspired the interface without mixing its claims with AnyJev’s results.](/blog/jev-ai-decision-model-review/)

Models and infrastructure

## Continue through this cluster

Model selection, inference economics, local deployment, compression and serving architecture.

[Start with the cornerstone**Self-Hosting LLMs in the EU: When Open Weights Actually Pay Off**](/blog/self-hosting-llms-eu-cost/)

- [DeerFlow 2.0: Docker Setup, Sandboxes and Memory](/blog/deerflow-2-docker-setup-sandbox-memory/)
- [CLM-8B Self-Hosting: vLLM, Action Cache and Verifiers](/blog/clm-8b-self-hosting-action-cache-verifier/)
- [LiteAgents SDK: Per-Turn Routing, Setup and Migration](/blog/liteagents-sdk-per-turn-model-routing/)
- [mcp-memory-service: Shared Memory for Claude Code and Cursor](/blog/mcp-memory-service-claude-code-cursor/)
- [Claude Opus 5.5: Best Uses, Prompts and Effort Settings](/blog/claude-opus-5-5-best-use-cases-workflows/)

[**Back**](/blog/overview/)

[![Kevin Riedl](/img/team/kevin.webp)](/team/kevin-riedl/)

[Kevin Riedl](/team/kevin-riedl/) https://linkedin.com/in/wsdt

16 min read · 28 Sep 2026 Last reviewed September 28, 2026

[**Next**](/blog/llm-evaluation-cost-roi-production/)

## Structured Data

```json
{
  "@context": "https://schema.org",
  "@graph": [
    {
      "@id": "https://wavect.io/#organization",
      "@type": [
        "Organization",
        "ProfessionalService",
        "LocalBusiness"
      ],
      "employee": [
        {
          "@id": "https://wavect.io/team/kevin-riedl/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Kevin Riedl",
          "url": "https://wavect.io/team/kevin-riedl/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        },
        {
          "@id": "https://wavect.io/team/christof-jori/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Christof Jori",
          "url": "https://wavect.io/team/christof-jori/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        }
      ],
      "founder": [
        {
          "@id": "https://wavect.io/team/kevin-riedl/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Kevin Riedl",
          "url": "https://wavect.io/team/kevin-riedl/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        },
        {
          "@id": "https://wavect.io/team/christof-jori/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Christof Jori",
          "url": "https://wavect.io/team/christof-jori/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        }
      ],
      "legalRepresentative": [
        {
          "@id": "https://wavect.io/team/kevin-riedl/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Kevin Riedl",
          "url": "https://wavect.io/team/kevin-riedl/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        },
        {
          "@id": "https://wavect.io/team/christof-jori/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Christof Jori",
          "url": "https://wavect.io/team/christof-jori/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        }
      ],
      "name": "Wavect GmbH",
      "subjectOf": {
        "@id": "https://wavect.io/verified-claims.json#dataset",
        "@type": "Dataset",
        "creator": {
          "@id": "https://wavect.io/#organization",
          "@type": [
            "Organization",
            "ProfessionalService",
            "LocalBusiness"
          ]
        },
        "description": "A machine-readable registry of quantitative and qualitative claims published by Wavect, with review dates, localized page appearances and public third-party citations where available.",
        "inLanguage": "en",
        "isAccessibleForFree": true,
        "license": "https://creativecommons.org/licenses/by/4.0/",
        "name": "Wavect verified publication claims",
        "url": "https://wavect.io/verified-claims.json"
      },
      "url": "https://wavect.io/"
    },
    {
      "@id": "https://wavect.io/team/kevin-riedl/#person",
      "@type": "Person",
      "jobTitle": "Managing Director",
      "name": "Kevin Riedl",
      "sameAs": [
        "https://www.wikidata.org/wiki/Q139796365",
        "https://www.linkedin.com/in/wsdt",
        "https://github.com/wsdt"
      ],
      "url": "https://wavect.io/team/kevin-riedl/",
      "worksFor": {
        "@id": "https://wavect.io/#organization",
        "@type": [
          "Organization",
          "ProfessionalService",
          "LocalBusiness"
        ]
      }
    },
    {
      "@id": "https://wavect.io/team/christof-jori/#person",
      "@type": "Person",
      "jobTitle": "Managing Director",
      "name": "Christof Jori",
      "sameAs": [
        "https://www.wikidata.org/wiki/Q139796367",
        "https://www.linkedin.com/in/jocr77/",
        "https://github.com/jo-chris"
      ],
      "url": "https://wavect.io/team/christof-jori/",
      "worksFor": {
        "@id": "https://wavect.io/#organization",
        "@type": [
          "Organization",
          "ProfessionalService",
          "LocalBusiness"
        ]
      }
    },
    {
      "@id": "https://wavect.io/#website",
      "@type": "WebSite",
      "inLanguage": [
        "en",
        "de",
        "es",
        "zh"
      ],
      "name": "Wavect",
      "potentialAction": {
        "@type": "SearchAction",
        "query-input": "required name=search_term_string",
        "target": {
          "@type": "EntryPoint",
          "urlTemplate": "https://wavect.io/search/?q={search_term_string}"
        }
      },
      "publisher": {
        "@id": "https://wavect.io/#organization",
        "@type": [
          "Organization",
          "ProfessionalService",
          "LocalBusiness"
        ]
      },
      "url": "https://wavect.io/"
    },
    {
      "@id": "https://wavect.io/blog/anyjev-calibration-option-order-bias/#webpage",
      "@type": "WebPage",
      "dateModified": "2026-09-28",
      "inLanguage": "en",
      "isPartOf": {
        "@id": "https://wavect.io/#website",
        "@type": "WebSite"
      },
      "lastReviewed": "2026-09-28",
      "url": "https://wavect.io/blog/anyjev-calibration-option-order-bias/"
    }
  ]
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "BlogPosting",
  "abstract": "AnyJev turns existing language models into typed decision components. L0 needs no labels and reduces order and prior bias; L1 fits confidence on labeled examples; L2 fits a per-question head without changing the backbone. The current 0.2.0 release adds optional canonical ordering and adaptive rotations. Its 1% adaptive target measures agreement with its own full-cycle answer, not real-world accuracy. This guide separates those claims, examines negative results and provides an offline calibration-to-evaluation example.",
  "articleBody": " Blog overview/AI and agents/Models and infrastructure AnyJev: LLM Calibration and Option-Order Bias TL;DR AnyJev turns existing language models into typed decision components. L0 needs no labels and reduces order and prior bias; L1 fits confidence on labeled examples; L2 fits a per-question head without changing the backbone. The current 0.2.0 release adds optional canonical ordering and adaptive rotations. Its 1% adaptive target measures agreement with its own full-cycle answer, not real-world accuracy. This guide separates those claims, examines negative results and provides an offline calibration-to-evaluation example. Reviewed 28 September 2026. Package baseline: AnyJev 0.2.0. Source snapshot: 10d5db91dda38dbde74c6abc1c075ce6463723d1. This is a documentation and source-code review, not a GPU benchmark or a report of a Wavect production deployment. AnyJev 0.2.0 release What is AnyJev, and what does “no training” mean? AnyJev is an open-source library that reads typed decisions from an existing language model instead of asking it to write an answer. Its distinctive problem is not just speed: it addresses option-order bias and separates debiased scores from confidence calibrated on labels. The project credits Jiamu Zhang, Tianze Yang, Yucheng Shi and Liang Wu, with Nokia and Tencent Hunyuan affiliations. It is inspired by TypeSafe AI's Jev interface, not affiliated with TypeSafe. Project overview and attribution The package is released under Apache 2.0. That is not a blanket license for every model or dataset used with it. Package license metadata. The practical question is straightforward. Your agent already has a ticket, a document or a tool result. It must choose one of several permitted answers. Can you trust that choice when someone changes the order of the options? And does a confidence of 0.9 mean that comparable decisions are actually correct about 90% of the time? Those are different problems. A stable answer can be wrong. A normalized distribution can be overconfident. AnyJev provides different processing levels rather than one universal fix. “Training-free” is accurate for its label-free L0 path. L1 fits a temperature using labeled examples. L2 fits a supervised linear head, although it does not update the backbone or use gradient descent for that fit. Calling all three “no training” hides the data and validation work. The more useful distinction is no backbone fine-tuning versus no labeled fitting. Level-by-level contract For the general product landscape, use our Jev review and Laya versus Jev benchmark analysis. This guide concentrates on AnyJev's calibration, ordering and evaluation mechanics, not another decision-model ranking. Raw, L0, L1 and L2: which path needs labels? AnyJev processing levels: requirements and limits LevelData and mechanismWhat it does not establish rawNo labels. Read a distribution restricted to the answer-label tokens.Neither order robustness nor calibrated uncertainty. L0No labels. Combine option rotations and optionally correct the label prior.A confidence value is not a validated probability of correctness. L1Typically 100–500 labels for the question. Fit temperature scaling above L0.No protection against new data distributions or incorrect model knowledge. L2Typically 100–300 labels per question. Fit a closed-form head on hidden states.No general transfer to an unrelated question or another base model. The ranges are project guidance, not sample-size guarantees. A task with rare but costly mistakes can require much more evidence. Likewise, auto means “use the best available level,” not “automatically safe.” Check the level actually returned. The table summarizes the documented level contract. Why rearranging the same options changes an LLM's answer Consider three semantic answers: billing, technical support and sales. A naïve implementation maps them to A, B and C, runs a next-token readout and selects the highest score. But a preference for “A,” an early position or a neighboring option can become a preference for whichever business answer occupies that position. This is an established research problem, not a problem first discovered by AnyJev. Zheng and colleagues examined the sensitivity of multiple-choice LLM selection to option ordering. Their work is part of the foundation for the project's permutation approach. Research on multiple-choice order sensitivity AnyJev normally evaluates the cyclic rotations of a K-option choice. Each option appears in every position once. The default logmean method maps the probabilities back to semantic options, averages their logarithms, exponentiates and renormalizes. Under an idealized additive logit-position bias, the position term cancels. Rotation and marginalization implementation That idealized cancellation is not a proof of invariance under every real prompt permutation. A full cycle preserves which options are neighbors; actual model responses can depend on those interactions. The previously published BANKING77",
  "articleSection": "Decision Models",
  "author": {
    "@id": "https://wavect.io/team/kevin-riedl/#person",
    "@type": "Person",
    "name": "Kevin Riedl",
    "sameAs": [
      "https://www.wikidata.org/wiki/Q139796365",
      "https://www.linkedin.com/in/wsdt",
      "https://github.com/wsdt"
    ],
    "url": "https://wavect.io/team/kevin-riedl/"
  },
  "citation": [
    {
      "@type": "WebPage",
      "name": "AnyJev 0.2.0 release",
      "url": "https://pypi.org/project/anyjev/0.2.0/"
    },
    {
      "@type": "WebPage",
      "name": "Project overview and attribution",
      "url": "https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/README.md"
    },
    {
      "@type": "WebPage",
      "name": "Level-by-level contract",
      "url": "https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/docs/levels.md"
    },
    {
      "@type": "WebPage",
      "name": "Research on multiple-choice order sensitivity",
      "url": "https://arxiv.org/abs/2309.03882"
    },
    {
      "@type": "WebPage",
      "name": "Rotation and marginalization implementation",
      "url": "https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/anyjev/calibrate/permute.py"
    },
    {
      "@type": "WebPage",
      "name": "Decider defaults, calibration and state handling",
      "url": "https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/anyjev/decider.py"
    },
    {
      "@type": "WebPage",
      "name": "Positive and negative L0 results",
      "url": "https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/docs/when_l0_helps.md"
    },
    {
      "@type": "WebPage",
      "name": "Batch Calibration research",
      "url": "https://arxiv.org/abs/2309.17249"
    },
    {
      "@type": "WebPage",
      "name": "Calibrate Before Use",
      "url": "https://proceedings.mlr.press/v139/zhao21c.html"
    },
    {
      "@type": "WebPage",
      "name": "Temperature-scaling implementation",
      "url": "https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/anyjev/calibrate/posthoc.py"
    },
    {
      "@type": "WebPage",
      "name": "On Calibration of Modern Neural Networks",
      "url": "https://proceedings.mlr.press/v70/guo17a.html"
    },
    {
      "@type": "WebPage",
      "name": "Committed benchmark tables",
      "url": "https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/docs/results_bench.md"
    },
    {
      "@type": "WebPage",
      "name": "Rotation-budget experiments and limitations",
      "url": "https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/docs/rotation_budget.md"
    },
    {
      "@type": "WebPage",
      "name": "Typed-question validation and hashing",
      "url": "https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/anyjev/question.py"
    },
    {
      "@type": "WebPage",
      "name": "HF backend and model loading",
      "url": "https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/anyjev/backends/hf.py"
    },
    {
      "@type": "WebPage",
      "name": "Decision fields and level enforcement",
      "url": "https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/anyjev/result.py"
    },
    {
      "@type": "WebPage",
      "name": "Serving adapter and endpoint contracts",
      "url": "https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/anyjev/backends/vllm.py"
    },
    {
      "@type": "WebPage",
      "name": "Closed-form head fitting",
      "url": "https://github.com/nokia-applied-research/AnyJev/blob/10d5db91dda38dbde74c6abc1c075ce6463723d1/anyjev/heads.py"
    }
  ],
  "dateModified": "2026-09-28",
  "datePublished": "2026-09-28",
  "description": "AnyJev turns existing language models into typed decision components. L0 needs no labels and reduces order and prior bias; L1 fits confidence on labeled examples; L2 fits a per-question head without changing the backbone. The current 0.2.0 release adds optional canonical ordering and adaptive rotations. Its 1% adaptive target measures agreement with its own full-cycle answer, not real-world accuracy. This guide separates those claims, examines negative results and provides an offline calibration-to-evaluation example.",
  "headline": "AnyJev: LLM Calibration and Option-Order Bias",
  "image": "https://wavect.io/img/blog/headers/header_anyjev-calibration-option-order-bias.svg",
  "inLanguage": "en",
  "keywords": "AI agents, AnyJev, LLM calibration",
  "mainEntityOfPage": {
    "@id": "https://wavect.io/blog/anyjev-calibration-option-order-bias/",
    "@type": "WebPage"
  },
  "publisher": {
    "@id": "https://wavect.io/#organization",
    "@type": [
      "Organization",
      "ProfessionalService",
      "LocalBusiness"
    ]
  },
  "url": "https://wavect.io/blog/anyjev-calibration-option-order-bias/",
  "wordCount": 3823
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "BreadcrumbList",
  "itemListElement": [
    {
      "@type": "ListItem",
      "item": "https://wavect.io/",
      "name": "Home",
      "position": 1
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/overview/",
      "name": "Blog overview",
      "position": 2
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/topics/ai-agents/",
      "name": "AI and agents",
      "position": 3
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/clusters/models-infrastructure/",
      "name": "Models and infrastructure",
      "position": 4
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/anyjev-calibration-option-order-bias/",
      "name": "AnyJev: LLM Calibration and Option-Order Bias",
      "position": 5
    }
  ]
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "FAQPage",
  "mainEntity": [
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "L0 needs no labels or backbone fine-tuning. L1 fits a temperature on labels, and L2 fits a supervised head. No gradient update to the backbone is not the same as no labeled learning. Level contract."
      },
      "name": "Is AnyJev really training-free?"
    },
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Full rotations remove an idealized additive position effect, but real prompts can retain interaction effects. Optional canonical_order=True normalizes caller ordering in 0.2.0. Compare semantic answers under fixed inference and prior state; stability does not prove correctness. Ordering details."
      },
      "name": "Does AnyJev completely remove option-order bias?"
    },
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "You can estimate a prior and calibrate adaptive agreement with the full-cycle readout without labels. Validating confidence against correct answers is different: L1 requires labels for the question. The adaptive 1% target is not a ground-truth error guarantee. L0 and L1."
      },
      "name": "Can I calibrate AnyJev without labeled examples?"
    },
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "ECE summarizes calibration, not the quality of the lowest-risk subset at every threshold. In the published Qwen3-8B newsgroups rows, L1 lowers ECE but reduces empirical coverage at 5% error. Measure both probability quality and accepted-decision risk. Reported counterexample."
      },
      "name": "Why can lower ECE still mean worse automation coverage?"
    },
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Not automatically. The reviewed vLLM logit path needs allowed-token control and answer-token log probabilities; L2 needs compatible hidden states. A generic chat or embeddings API is not sufficient evidence of compatibility. Backend contract."
      },
      "name": "Does AnyJev work with any hosted LLM API?"
    },
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "The current Question.choice constructor accepts 2 to 26 unique options. Ordinal score questions use 2 to 10 bins or levels. Larger candidate catalogs need another design or a later implementation, not an assumed bypass. Question validation."
      },
      "name": "How many options does AnyJev support?"
    },
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "No. It enforces a minimum processing level, not permissions or validated business risk. Calibrate on representative data, choose a threshold on validation data and evaluate it on an untouched test set. Keep a separate review path and deterministic authorization. Level enforcement."
      },
      "name": "Does require=\"L1\" make a decision safe to execute?"
    },
    {
      "@type": "Question",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Do not assume that a head transfers to an unrelated question or another model. AnyJev fits heads per question and model. Its same-question routing across wording or order changes still needs validation on the new inputs. Similar category names alone do not establish compatibility. Head scope and limits."
      },
      "name": "Can I reuse an AnyJev L2 head for a different question?"
    }
  ]
}
```
