BENCHMARKS · V0.8.0 SOURCE AUDIT

Context is measurable. Cost savings are not proven yet.

Publish the counterexample before the claim.

Semaprax now publishes measured compact-projection tokens, task-level provider receipts, a paid harness qualification and real Rust and hot-reload application measurements. Some views shrink; small views grow. The recorded cost campaign did not qualify a new default, and the small hot-reload fixture was slower than restarting.

semaprax://benchmarksv0.8.0
entity     Semaprax
status     beta
snapshot   615e501
authority  github.com/wavect/semaprax
// STATUS

Repository status at the audited snapshot

Repository snapshot615e501 · 2026-10-06. The v0.8.0 tag resolves to the pinned source commit. Its exact-tag CI completed successfully. This does not establish blanket production or security assurance.
Full product status55 Partial · 0 Implemented · 0 Missing
Graph and project contractCanonical source, stable IDs and checked compiler representations connect laws, semantic queries, edits and execution. Feature-selected graph and Project schemas preserve their own admission, compatibility and host-authority contracts.
Published beta releasev0.8.0 · Published 6 October 2026 at 07:37 UTC. All 82 exact-tag jobs succeeded. Three toolchain archives, SHA256SUMS, per-archive attestations and signed aggregate provenance were published. The release job independently verified the signed set before publication. The archives are not notarized or claimed reproducible; offline verification does not establish current revocation state. Exact-tag CI.
Package and API previewsThe release includes useful language, law, harness and host-integration profiles. Generated Rust/npm packages, public generic ABIs and broader platform support retain separate publication and support decisions.

Is Semaprax cheaper for coding agents?

A general cost advantage is not established. Exact local cl100k_base measurements show smaller model-text payloads for two large graphs and an HTTP task context, while small contexts grow. A separate paid harness campaign measured cost per accepted task; its best saving was about 1%, below the predeclared 10% threshold. Provider usage, local token counts, execution speed and task correctness remain separate measurements.

// TOKENS

The same selected facts, counted two ways

Five rows from the committed local model-text v2 report. Counts use cl100k_base over the complete payload, including metadata. They are local tokenizer counts, not provider invoices; a higher model-text number is a regression.

Selected viewJSON tokensModel-text tokens
Banking ledger: full graph3784632486
ledger.apply: task context16791802
HTTP application: full graph161861134165
HTTP app.main: task context56025010
Calculator Project: graph25002615

Read the exact corpus, both tokenizers and replay contract

// 01

Six measured or replayed evidence tracks

Exact compact-payload tokens

With cached tiktoken 0.12.0, the large banking and HTTP graphs fall from 37,846 to 32,486 and from 161,861 to 134,165 cl100k_base tokens. The HTTP task view falls from 5,602 to 5,010. The complete model-text envelope is counted and replayed against the same selected JSON.

Small-view regressions

The ledger task context grows from 1,679 to 1,802 tokens; the calculator Project graph grows from 2,500 to 2,615. Fixed metadata can outweigh compression. Source-versus-context comparisons answer a different question from encoding the same selected facts.

Session reports and provider receipts

Token reports group compatible tokenizer fingerprints, methods and measurement boundaries. Harness receipts preserve actual provider usage, failed/truncated attempts and estimated versus reported cost. Missing counts remain unavailable; receipts are not a comparison outcome by themselves.

Paid task-cost qualification

The 5 October campaign used Claude Haiku 4.5 through a metering shim: 711 calls and USD 3.10 recorded spend. Each arm accepted 109 of 120 tasks. The best cost-per-accepted-task saving was about 1%, below the required 10%; no arm qualified and defaults stayed unchanged. The app-task path does not exercise the entire Semaprax source workflow.

Real Rust application overhead

Saved Regex/Url, Serde/iterator and reqwest/Tokio applications have clean locked offline macOS arm64 and Linux x86-64 guest receipts, each with 22 passed stages/measurements. Copy and allocation ledgers accompany the results. Adverse M1/M3 throughput remains an investigation, not a near-zero-overhead claim.

Checked hot reload versus restart

On the recorded macOS arm64 fixture, cold A-to-B reload had a 545.694 ms median versus 37.904 ms for full restart. Eleven samples and separate admission, preparation, wait and activation measurements preserve that counterevidence. It demonstrates checked continuity, not a speedup or a result for other operating systems.

// 02

What these results do not establish

  • A general advantage in provider-billed tokens or cost per accepted task across projects and models.
  • Better correctness, latency or maintenance productivity than Rust, TypeScript or another language.
  • Universal Rust interoperability without overhead, or physical Linux x86-64 performance from an emulated guest.
  • A hot-reload speed advantage, native/Wasm process swapping or broad platform performance.
  • Current release performance from historical debug-build timing or correctness from a context compression ratio.
// 03

The evidence needed before a cost claim

  1. Freeze equivalent tasks, source revisions, adapters and the acceptance grader before the experiment.
  2. Keep model identity, permissions, token and money limits, retries and stopping rules comparable.
  3. Record complete provider receipts, elapsed time, accepted-task rate and regression failures, including unsuccessful attempts.
  4. Publish raw measurements, missing evidence, exclusions and uncertainty. Keep local tokenizers separate from billed usage.
  5. Require the declared quality and cost thresholds before changing defaults; repeat on additional projects, models and hosts.

Reviewed against the v0.8.0 tag and its pinned source commit. Versioned specifications define admission and authority; older release headings in the matrix and roadmap are historical. GitHub repository.

// DOCS

Open the Semaprax handbook

The online handbook is the English guide published from main and may cover changes after 0.8.0. Use the pinned snapshot when reproducing this release.

// REF

Primary sources for this page

Reviewed against the v0.8.0 tag and its pinned source commit. Versioned specifications define admission and authority; older release headings in the matrix and roadmap are historical.

  1. Compact model-text contract and exact token table
  2. Raw local token measurement report
  3. Per-input and session token reporting
  4. Provider receipt, cost and output-cap semantics
  5. Paid harness qualification and unchanged defaults
  6. Real Rust application receipts and overhead
  7. Hot-reload timing and restart comparison
  8. Original context measurement boundary
// FAQ

Questions and practical answers

Does Semaprax use fewer tokens than Rust or C?

No general language comparison is established. The compact-context table compares two encodings of the same selected Semaprax facts and includes both reductions and regressions. A separate paid harness campaign measured cost per accepted task, but its roughly 1% best saving did not reach the 10% qualification threshold.

Why publish a benchmark that does not show savings?

It tells users when an optimization helps and when it does not. Small compact views can grow, the recorded hot-reload fixture is slower than restart, and some Rust application routes have adverse throughput. Keeping those results with the positive evidence makes default and support decisions reviewable.

Research project by Wavect: Wavect GmbH. Created by Wavect as an open-source systems research project.