In this piece
AnyJev: LLM Calibration and Option-Order Bias
Reviewed . Package baseline: AnyJev 0.2.0. Source snapshot: 10d5db91dda38dbde74c6abc1c075ce6463723d1. This is a documentation and source-code review, not a GPU benchmark or a report of a Wavect production deployment. AnyJev 0.2.0 release
What is AnyJev, and what does “no training” mean?
AnyJev is an open-source library that reads typed decisions from an existing language model instead of asking it to write an answer. Its distinctive problem is not just speed: it addresses option-order bias and separates debiased scores from confidence calibrated on labels. The project credits Jiamu Zhang, Tianze Yang, Yucheng Shi and Liang Wu, with Nokia and Tencent Hunyuan affiliations. It is inspired by TypeSafe AI's Jev interface, not affiliated with TypeSafe. Project overview and attribution
The package is released under Apache 2.0. That is not a blanket license for every model or dataset used with it. Package license metadata.
The practical question is straightforward. Your agent already has a ticket, a document or a tool result. It must choose one of several permitted answers. Can you trust that choice when someone changes the order of the options? And does a confidence of 0.9 mean that comparable decisions are actually correct about 90% of the time?
Those are different problems. A stable answer can be wrong. A normalized distribution can be overconfident. AnyJev provides different processing levels rather than one universal fix.
“Training-free” is accurate for its label-free L0 path. L1 fits a temperature using labeled examples. L2 fits a supervised linear head, although it does not update the backbone or use gradient descent for that fit. Calling all three “no training” hides the data and validation work. The more useful distinction is no backbone fine-tuning versus no labeled fitting. Level-by-level contract
For the general product landscape, use our Jev review and Laya versus Jev benchmark analysis. This guide concentrates on AnyJev's calibration, ordering and evaluation mechanics, not another decision-model ranking.
Raw, L0, L1 and L2: which path needs labels?
| Level | Data and mechanism | What it does not establish |
|---|---|---|
| raw | No labels. Read a distribution restricted to the answer-label tokens. | Neither order robustness nor calibrated uncertainty. |
| L0 | No labels. Combine option rotations and optionally correct the label prior. | A confidence value is not a validated probability of correctness. |
| L1 | Typically 100–500 labels for the question. Fit temperature scaling above L0. | No protection against new data distributions or incorrect model knowledge. |
| L2 | Typically 100–300 labels per question. Fit a closed-form head on hidden states. | No general transfer to an unrelated question or another base model. |
The ranges are project guidance, not sample-size guarantees. A task with rare but costly mistakes can require much more evidence. Likewise, auto means “use the best available level,” not “automatically safe.” Check the level actually returned. The table summarizes the documented level contract.
Why rearranging the same options changes an LLM's answer
Consider three semantic answers: billing, technical support and sales. A naïve implementation maps them to A, B and C, runs a next-token readout and selects the highest score. But a preference for “A,” an early position or a neighboring option can become a preference for whichever business answer occupies that position.
This is an established research problem, not a problem first discovered by AnyJev. Zheng and colleagues examined the sensitivity of multiple-choice LLM selection to option ordering. Their work is part of the foundation for the project's permutation approach. Research on multiple-choice order sensitivity
AnyJev normally evaluates the cyclic rotations of a K-option choice. Each option appears in every position once. The default logmean method maps the probabilities back to semantic options, averages their logarithms, exponentiates and renormalizes. Under an idealized additive logit-position bias, the position term cancels. Rotation and marginalization implementation
That idealized cancellation is not a proof of invariance under every real prompt permutation. A full cycle preserves which options are neighbors; actual model responses can depend on those interactions. The previously published BANKING77 results still contain residual answer flips.
Version 0.2.0 offers a separate, opt-in remedy: canonical_order=True sorts options by their text before rotating. The same option set then produces the same prompt layouts regardless of caller order. This is a deterministic input-normalization property, assuming the same inference and prior state. It does not make the answer correct, eliminate numerical variation or make a running batch prior independent of request history. The rotation-budget documentation explains why canonicalization also matters when only some rotations are read.
When testing reorder robustness, compare returned semantic labels, not option indices. Index zero means something different after reversing the list. Also distinguish disagreement among individual rotations from disagreement between two complete calls with reordered inputs.
When batch-prior correction helps, and when it hurts
L0 can correct a label prior estimated from model predictions on unlabeled inputs. Conceptually, it divides each answer score by an estimate of how often that answer is favored before renormalizing. The default batch-prior strength is 0.75, with correction starting after the per-question accumulator has at least eight inputs. Before that, diagnostics report no prior correction. Decider defaults, calibration and state handling
The catch: a common prediction may represent a biased model or a genuinely common class. Unlabeled prediction frequencies alone cannot reliably distinguish the two. For a support queue where most requests truly belong to billing, correcting that majority toward balance can damage the router.
The project's diagnostic study covers 230 model-question combinations. It reports useful gains on balanced 20-way tasks, but also substantial losses on some skewed questions. It retains prior="none" to isolate the rotation effect when the label marginal is known to be skewed. That is an ablation to evaluate, not a promise that switching the prior off always wins. Positive and negative L0 results
The underlying batch-calibration approach comes from Zhou and colleagues. AnyJev packages it with rotations and operational interfaces; it does not invent all of these methods from scratch. Batch Calibration research
There is also content-free calibration using probes such as an empty input or N/A, following the approach of Zhao and colleagues. But a model's answer to an empty input can itself have meaning. Treat prior="content_free" as another measured variant, not a universally cleaner baseline. Calibrate Before Use
For a routing pilot, compare raw, rotation-only L0 and default L0 on both a class-balanced diagnostic set and a time-based sample of natural traffic. The former helps expose weak classes; the latter tells you what the system will encounter.
What L1 calibrates, and why temperature is not an accuracy upgrade
Temperature scaling changes the concentration of a fixed probability vector. Conceptually, it applies softmax(log(p) / T) with positive T. A larger temperature generally softens confidence; a smaller one sharpens it. AnyJev fits T by minimizing negative log likelihood on labeled calibration examples and stores the prior used for that calibration in the artifact. Temperature-scaling implementation
For a fixed vector, positive temperature scaling preserves the winning class. It does not turn an incorrect class into a correct one. The calibration literature distinguishes the quality of probability estimates from classification accuracy; Guo and colleagues provide the foundational treatment used here. On Calibration of Modern Neural Networks
Why, then, can AnyJev's L0 and L1 accuracy columns differ slightly? The complete pipeline is not only a temperature: L1 freezes the prior estimated from its calibration states, while ordinary L0 can accumulate a different running prior. Do not attribute the difference to temperature changing the argmax. This follows from the Decider implementation and the monotonic transform above.
A useful deployment artifact therefore needs more than a temperature value. Our recommendation is to record the base-model revision, tokenizer, question wording, exact option layout, prior configuration, package version and calibration-data window. A changed question or serving configuration is a reason to revalidate, not quietly reuse a favorable confidence threshold.
What the published benchmarks actually show
The following results are project-reported for Qwen3-8B on a 20-class subset of BANKING77 with 300 test items. They are not a full 77-class benchmark, a Nokia production SLA, or a measurement of the newer opt-in canonical/adaptive configuration. Committed benchmark tables
| Metric | raw | L0 | L1 |
|---|---|---|---|
| Accuracy | 74.7% | 80.3% | 80.7% |
| Expected calibration error, lower is better | 0.240 | 0.184 | 0.095 |
| Answer-flip rate under reversal | 23.0% | 7.3% | 7.7% |
| Empirical coverage at 5% error | 7.7% | 46.3% | 52.0% |
The gains are worth investigating. But the strongest qualification is in the last row: coverage is a retrospective property of this confidence-sorted test sample, not proof that a fixed threshold will automate 52% of future traffic at 5% error. At this sample size, 52% represents roughly 156 decisions. One additional error among 156 accepted decisions changes their observed error rate by about 0.64 percentage points. That arithmetic illustrates uncertainty; it does not reconstruct the unpublished error count.
There is an equally useful counterexample in the same table source. On Qwen3-8B newsgroups, L1 improves ECE from 0.309 to 0.138, while empirical coverage at 5% error falls from 42.3% to 23.7%. Better average calibration does not automatically create better low-risk selection. Evaluate both probability quality and the risk-coverage curve. Benchmark source.
L2 results need another distinction. The repository reports 77.1% for Qwen3-8B on 2,000 typed decisions after per-question labeled fitting. In that dataset, “accuracy” is agreement with teacher-derived labels, not an independent human adjudication of every business outcome. The README marks third-party Jev comparison rows as published elsewhere, not rerun head-to-head. Do not turn those rows into an overall product winner. Project limitations.
Adaptive rotations: a 1% target is not 99% accuracy
Adaptive rotations try to stop before all K layouts are evaluated. Calibration uses unlabeled states because the reference answer is the model's own full-cycle result. A stopping threshold is selected using an upper confidence bound on disagreement with that reference. This estimates agreement with another readout, not correctness against the real task. Rotation-budget experiments and limitations
In one reported Qwen2.5-7B, 18-option routing experiment on an H100 NVL with vLLM 0.7.0, full rotations delivered 16.73 decisions/s. Adaptive waves of two delivered 37.20 decisions/s using 7.28 requests per decision, a 2.22× measured throughput ratio. However, reference agreement was 98.7%, or 1.3% disagreement, outside the nominal 1% target. The authors disclose this result. It is not a contractual bound on future requests.
The practical sequence is to establish a full-rotation baseline, enable canonical ordering, calibrate the adaptive budget on separate representative states, then measure both reference disagreement and ground-truth errors on held-out data. Small calibration sets, long contexts and different serving engines can change the result. Do not translate fewer rotations directly into the same factor of end-to-end agent acceleration.
A reproducible starting point for a routing experiment
The following is an illustrative offline workflow, not a measured deployment. AnyJev 0.2.0 requires Python 3.10 or newer. Install into an isolated environment and lock the resolved dependencies for your own reproduction; a package pin alone does not pin PyTorch, Transformers or model weights. Package baseline.
python -m venv .venv
. .venv/bin/activate
python -m pip install "anyjev[hf]==0.2.0"
Question.choice accepts two to 26 unique option strings in the current implementation. Labels supplied to calibration are integer indices into those options. Keep a stable semantic mapping, and avoid overlapping categories that make even a human annotator guess. Typed-question validation and hashing
For this example, use a compatible CUDA environment with enough memory for the chosen base model. The HF backend loads the model and tokenizer, accepts a model revision and does not enable remote model code by default. A small AnyJev package is not a small replacement for the backbone. HF backend and model loading
Set ANYJEV_MODEL_REVISION to the actual 40-character commit of the chosen Qwen3-8B model before running the Python blocks in order. It is a model revision, not the AnyJev source pin.
import os
import re
from anyjev import Decider, Question
from anyjev.backends.hf import HFBackend
# Set this to the actual Qwen model commit, not the AnyJev commit.
revision = os.environ.get("ANYJEV_MODEL_REVISION", "")
if not re.fullmatch(r"[0-9a-f]{40}", revision):
raise ValueError("Set ANYJEV_MODEL_REVISION to a Qwen3-8B commit SHA")
backend = HFBackend(
"Qwen/Qwen3-8B", revision=revision,
device="cuda", dtype="bfloat16",
)
route = Question.choice(
"Which internal queue should review this ticket? Classify only.",
["Billing", "Technical support", "Sales", "Manual review"],
name="route",
)
d = Decider(
backend, level="L0", prior="none",
canonical_order=True, adaptive_shifts=False,
)
result = d.decide("I need a copy of my invoice.", [route])["route"]
print(result.level, result.distribution) # Inspection only; no dispatch.
The example deliberately uses full rotations, canonical ordering and prior="none" to establish a permutation-only baseline. It does not silently rely on a warm running prior. Compare it with the default batch correction before choosing either configuration for real traffic. The English question and labels stay identical across the translations of this article so the example remains the same experiment.
Next, collect separate calibration and held-out evaluation files. Each non-empty JSONL line must contain id, state and a label matching one of the option strings. The illustrative row below documents the format; it is not a sufficient calibration dataset.
{"id":"ticket-example-001","state":"Please resend my invoice.","label":"Billing"}
# Continue after the setup above. Supply your own labeled JSONL files.
import json
from pathlib import Path
def load_rows(filename: str, minimum: int = 1) -> list[dict]:
rows, ids, states = [], set(), set()
for line_no, line in enumerate(Path(filename).read_text(encoding="utf-8").splitlines(), 1):
if not line.strip():
continue
row = json.loads(line)
if not isinstance(row, dict) or not all(
isinstance(row.get(k), str) and row[k].strip()
for k in ("id", "state", "label")
):
raise ValueError(f"{filename}:{line_no}: id, state and label must be strings")
if row["label"] not in route.options:
raise ValueError(f"{filename}:{line_no}: unknown label")
if row["id"].strip() in ids or row["state"].strip() in states:
raise ValueError(f"{filename}:{line_no}: duplicate id or state")
ids.add(row["id"].strip())
states.add(row["state"].strip())
rows.append(row)
if len(rows) < minimum:
raise ValueError(f"{filename}: need at least {minimum} rows for this example")
return rows
calibration = load_rows("calibration.jsonl", minimum=100)
heldout = load_rows("heldout.jsonl")
for field in ("id", "state"):
if {r[field].strip() for r in calibration} & {r[field].strip() for r in heldout}:
raise ValueError(f"Calibration and held-out {field} values overlap")
if {r["label"] for r in calibration} != set(route.options):
raise ValueError("Calibration must cover every route in this example")
artifact = d.calibrate(
route, [r["state"] for r in calibration],
[route.options.index(r["label"]) for r in calibration], level="L1",
)
# Exclusive creation prevents accidental overwrite of an existing artifact.
with Path("route-calibration.json").open("x", encoding="utf-8") as f:
json.dump(artifact, f, ensure_ascii=False, indent=2)
predictions = d.decide_batch(
[r["state"] for r in heldout], route, level="L1", require="L1",
)
for row, prediction in zip(heldout, predictions):
print(json.dumps({
"id": row["id"], "expected": row["label"],
"predicted": prediction.argmax, "confidence": prediction.confidence,
"level": prediction.level,
}, ensure_ascii=False))
This code fits on one split and only scores the other. It neither chooses an automation threshold nor dispatches a ticket. The 100-row guard is an example policy, not a claim that 100 labels are statistically sufficient. Reject duplicate or overlapping inputs and check temporal or customer-level leakage as well as identifiers.
require="L1" prevents silent use of a weaker result; confidence is the maximum returned probability. Neither is a permission check or a guarantee of an acceptable business error rate. The result object also exposes the actual level and diagnostics. Decision fields and level enforcement
For production threshold selection, add a separate validation split. Fix the acceptance rule there and evaluate it once on the untouched test split. Report accepted counts and errors per class, not only a single aggregate confidence score.
Serving and artifact boundaries that can break the experiment
AnyJev's vLLM adapter has two different paths. Raw, L0 and L1 use a generation server and request max_tokens=1 with allowed answer-token IDs and log probabilities. There is no free-form answer to parse, but the adapter does request one token; “no generation” should not be read as zero token work at every API boundary. L2 instead needs unnormalized last-position hidden states from a pooling server. Serving adapter and endpoint contracts
A generic embeddings endpoint is not automatically interchangeable with the hidden-state interface. An ordinary vLLM pooling server exposes the final layer of its served checkpoint; a head fitted to an earlier layer needs a matching truncated checkpoint or a backend exposing that layer. Validate model identity, tokenizer, vector shape and preprocessing together.
For L2, the project uses a per-question linear head fitted by a closed-form solve, with held-out folds guiding head and layer selection. That is supervised learning even though no backbone gradient update occurs. A successful fit on one question does not validate another question with similar-looking labels. Closed-form head fitting
Two common operational mistakes are requesting L1 before loading or fitting its artifact, and fitting an artifact but continuing to call the default L0 path. Make level and require explicit. Keep the deployment configuration immutable during evaluation and record any fallback, prior warm-up or artifact mismatch as an event, not an invisible change in accuracy.
An acceptance test for a real decision pipeline
Our recommendation is to evaluate a narrow, reversible decision first. Build the test around the cost of an incorrect accepted decision, not around the novelty of a model interface.
Test correctness and stability separately. Measure class accuracy, macro F1, option-reversal disagreement and several non-cyclic reorderings. Reuse an identical prior state for each comparison. A stable answer is not necessarily a correct answer; a changed running prior is not necessarily an option-order defect.
Test confidence and coverage separately. Record NLL, Brier score, ECE, accepted-decision error and coverage at a preselected threshold. Inspect rare classes and each deployed language. The newsgroups counterexample above is a concrete reason not to optimize ECE alone.
Test costs at the actual boundary. Measure prompt count, p50/p95 latency, queueing, prefix-cache behavior and fallback rate. Include model serving and human review, not just the few arithmetic operations after logits arrive. Compare against your existing router and a simple rule-based baseline.
Keep execution controls outside the classifier. Restrict the candidate set to authorized actions, include a review or abstention path, and let deterministic business rules enforce permissions. A confidence threshold must not authorize a refund, account change or production command. Observe drift and recalibrate before expanding scope.
The repository still lists agent-loop evaluation as future work, so its isolated-decision results are not evidence that a complete production agent is already better. Project scope. For the wider rollout process, see our LLM evaluation guide and pilot kill-or-scale scorecard. Business workflow economics remain in the Laya/Jev ROI guide.
For the separate encoder-and-head deployment problem, see our CLM-8B self-hosting and action-cache guide. It covers candidate-vector reuse and verifier setup, while the evaluation here focuses on calibration and option order.
A useful AnyJev pilot ends with a versioned question, a reproducible probability pipeline and an untouched test result. That is a stronger basis for automation than either a confident answer or a faster next-token readout.
For implementation support, explore Wavect’s AI engineering service. The Twinsoft AI case study provides separate delivery context, not an AnyJev deployment claim. Bring representative decisions and the pre-launch QA checklist to discuss an evaluation-first integration.
AnyJev calibration questions
Is AnyJev really training-free?
L0 needs no labels or backbone fine-tuning. L1 fits a temperature on labels, and L2 fits a supervised head. No gradient update to the backbone is not the same as no labeled learning. Level contract.
Does AnyJev completely remove option-order bias?
Full rotations remove an idealized additive position effect, but real prompts can retain interaction effects. Optional canonical_order=True normalizes caller ordering in 0.2.0. Compare semantic answers under fixed inference and prior state; stability does not prove correctness. Ordering details.
Can I calibrate AnyJev without labeled examples?
You can estimate a prior and calibrate adaptive agreement with the full-cycle readout without labels. Validating confidence against correct answers is different: L1 requires labels for the question. The adaptive 1% target is not a ground-truth error guarantee. L0 and L1.
Why can lower ECE still mean worse automation coverage?
ECE summarizes calibration, not the quality of the lowest-risk subset at every threshold. In the published Qwen3-8B newsgroups rows, L1 lowers ECE but reduces empirical coverage at 5% error. Measure both probability quality and accepted-decision risk. Reported counterexample.
Does AnyJev work with any hosted LLM API?
Not automatically. The reviewed vLLM logit path needs allowed-token control and answer-token log probabilities; L2 needs compatible hidden states. A generic chat or embeddings API is not sufficient evidence of compatibility. Backend contract.
How many options does AnyJev support?
The current Question.choice constructor accepts 2 to 26 unique options. Ordinal score questions use 2 to 10 bins or levels. Larger candidate catalogs need another design or a later implementation, not an assumed bypass. Question validation.
Does require="L1" make a decision safe to execute?
No. It enforces a minimum processing level, not permissions or validated business risk. Calibrate on representative data, choose a threshold on validation data and evaluate it on an untouched test set. Keep a separate review path and deterministic authorization. Level enforcement.
Can I reuse an AnyJev L2 head for a different question?
Do not assume that a head transfers to an unrelated question or another model. AnyJev fits heads per question and model. Its same-question routing across wording or order changes still needs validation on the new inputs. Similar category names alone do not establish compatibility. Head scope and limits.
Final thoughts
Use AnyJev to test a specific decision interface, not to assume that confidence equals correctness. Separate order robustness, probability calibration and accepted-decision risk. Freeze the question and serving contract, preserve an untouched test split and keep execution permissions outside the model.
