BENCHMARKS · V0.5.0 SOURCE AUDIT

Context is measurable. Cost savings are not proven yet.

Publish the counterexample before the claim.

The repository now contains several measurement systems, not just the original Agent Context v1 byte comparison. Read their subject, toolchain, host and evidence status before treating them as results. A new compiler release does not refresh historical measurements.

semaprax://benchmarksv0.5.0
entity     Semaprax
status     pre-alpha research
snapshot   b9f593c
authority  github.com/wavect/semaprax
// STATUS

Repository status at the audited snapshot

Repository snapshotb9f593c · 2026-09-16. The reviewed main commit equals the v0.5.0 tag. Its exact-tag CI completed successfully; the separate branch CI was cancelled. Neither is a blanket production or security guarantee.
Full product status55 Partial · 0 Implemented · 0 Missing
Graph and project contractFeature-selected graph schemas preserve earlier contracts. Project v1 is the baseline; owned-data profiles v8, v9, v10 and v11 have distinct admission and support boundaries.
Published pre-releasev0.5.0 · Published 16 September 2026 at 09:43 UTC. Three archives: Linux x86-64, macOS Apple Silicon and Windows x86-64, with SHA256SUMS. The packaged semaprax binary is the full toolchain, not the standalone Cargo CLI. Exact-tag CI.
Package and API previewsA published toolchain does not publish its generated Rust/npm packages or promote private APIs. Project v8-v11 support decisions, public generic ownership, broader platforms and native/Wasm Agent-stage execution remain separate.

Is Semaprax cheaper for coding agents?

A general cost advantage is not established. Compact-projection tooling measures replay-checked graph/context bytes and tokens with cached cl100k_base and o200k_base tokenizers. Those are tokenizer counts, not provider invoices or accepted-task economics. The original small v1 corpus still supplies counterevidence: its structured context is larger than its source.

// 01

Five evidence tracks, five different boundaries

Original context baseline

A frozen Agent Context v1 corpus compares source bytes with deterministic structured-context bytes. It does not call a model or compare programming languages.

Compact graph and task context

The script emits text, binary and model-text projections, replays them against ordinary graph output, and records byte counts, hashes and offline tokenizer counts. The existence of the script does not establish a measured savings percentage.

Historical local performance

The committed 6 September baseline records 22 successful scenarios on one darwin-arm64 host using a debug v0.3.5 binary, with five samples per scenario. Its p50/p95 timings are advisory historical evidence, not v0.5.0 release performance.

Cross-language laboratory

Nine languages are listed; Semaprax, Rust and TypeScript have working adapters. The other six are blocked, not failed. Tasks, hidden oracles and provenance support pass/fail comparisons, but no timing measurements are committed and no language ranking follows.

Controlled model-pilot capture

v0.5.0 adds counterbalanced scheduling, isolated MCP tools for two comparison lanes, candidate source retention and exact transport archives. Captured trials remain distinct from eligible, reviewed observations. An offline two-call repair demo is not a live-provider productivity benchmark.

// 02

What the available evidence does not establish

  • A general advantage in provider-billed tokens or cost per accepted task.
  • Better correctness, latency or maintenance productivity than Rust, TypeScript or other languages.
  • v0.5.0 release performance from a historical v0.3.5 debug measurement.
  • Broad repository-scale or multi-model results from a small corpus, a harness or a captured pilot.
// 03

The evidence needed before a cost claim

  1. Freeze equivalent maintenance tasks and repository snapshots.
  2. Run the same model, harness, tool permissions, and stopping rules.
  3. Measure billed tokens, elapsed time, accepted-task rate, and regression failures.
  4. Publish raw traces, exclusions, unsuccessful runs, and confidence intervals.
  5. Repeat across repository sizes and more than one model family.

Reviewed against the pinned source commit and the v0.5.0 release record. Versioned specifications define admission and authority; older v0.4 release headings in the matrix and roadmap are historical, not the latest release identity. GitHub repository.

// REF

Primary sources for this page

Reviewed against the pinned source commit and the v0.5.0 release record. Versioned specifications define admission and authority; older v0.4 release headings in the matrix and roadmap are historical, not the latest release identity.

  1. Original Agent Context evidence contract
  2. Compact projection and offline-token measurement code
  3. Historical local v0.3.5 performance baseline
  4. Cross-language laboratory and explicit non-claims
  5. v0.5.0 release and downloadable archives
// FAQ

Questions, answered without the hype

Does Semaprax use fewer tokens than Rust or C?

No general comparison is established. The compact script can count tokens with specified offline tokenizers, while the cross-language harness runs Semaprax, Rust and TypeScript tasks without committed timing results. Neither establishes provider-billed token savings or better accepted-task cost.

Why publish a benchmark that does not show savings?

It makes the measurement boundary and counterevidence visible. Historical local timings, replay-checked token counts, harness correctness and captured model trials should remain separate until comparable reviewed observations support a narrower claim.

Research project by Wavect: Wavect GmbH. Created by Wavect as an open-source systems research project.