BENCHMARK · AGENT CONTEXT V1

Context is measurable. Cost savings are not proven yet.

Publish the counterexample before the claim.

The current benchmark remains specific to Agent Context v1. It freezes a small source corpus and compares deterministic context artifacts. It is a compiler-mechanics benchmark, not a model or repository economics result.

semaprax://benchmarksv0.2
entity     Semaprax
status     pre-alpha research
snapshot   c16348f
authority  github.com/wavect/semaprax
// STATUS

Repository status at the audited snapshot

Repository snapshotc16348f · 2026-08-29. Documentation passed, but the overall workflow failed, so this head is not labelled verified.
Full product status49 Partial · 0 Implemented · 0 Missing
Graph and project contractGraph ≤ v24 · Project v1 baseline · Project v8, v9, v10 developer preview
Promotion baselineNo exact passing promotion commit or workflow run is asserted for this snapshot.
Authored developer previewOwned-data, record, and UTF-8 project profiles, package analysis, borrowing additions, Project Agent Transport, and Revision Store are present in source but unpublished or unpromoted.

Is Semaprax cheaper for coding agents?

Not proven. Semaprax is designed to let agents request typed, bounded context, but the current Agent Context v1 benchmark does not measure model tokens, latency, answer quality, accepted-task rate, or repository-scale cost. On its small corpus, the context artifact is larger than the source.

// 01

What the current benchmark measures

Frozen input

A small versioned Semaprax corpus and maintenance task

Deterministic output

Agent Context v1 JSON with explicit byte, node, and depth limits

Comparison boundary

Source bytes versus structured context bytes for the fixed corpus

// 02

What it does not measure

  • Model input or output tokens
  • Wall-clock latency or provider price
  • Answer correctness or accepted patches
  • Repository-scale navigation and maintenance
  • A comparison with Rust, C, C++, Go, or another language
// 03

The evidence needed before a cost claim

  1. Freeze equivalent maintenance tasks and repository snapshots.
  2. Run the same model, harness, tool permissions, and stopping rules.
  3. Measure billed tokens, elapsed time, accepted-task rate, and regression failures.
  4. Publish raw traces, exclusions, unsuccessful runs, and confidence intervals.
  5. Repeat across repository sizes and more than one model family.

GitHub specifications are the normative source. This page is a dated research summary. GitHub repository.

// FAQ

Questions, answered without the hype

Does Semaprax use fewer tokens than Rust or C?

There is no credible evidence for that comparison yet. The current benchmark does not run a model or compare languages, and its small structured context artifact is larger than the source.

Why publish a benchmark that does not show savings?

Because it fixes the measurement contract and exposes counterevidence before a marketing claim. That makes later model and repository-scale results auditable.

Research project by Wavect: Wavect GmbH. Created by Wavect as an open-source systems research project.