BENCHMARK · AGENT CONTEXT V1

Context is measurable. Cost savings are not proven yet.

Publish the counterexample before the claim.

The current benchmark freezes a small source corpus and compares deterministic context artifacts. It is a compiler-mechanics benchmark, not a model or repository economics result.

semaprax://benchmarksv0.2
entity     Semaprax
status     pre-alpha research
verified   2026-08-11
authority  github.com/wavect/semaprax

Is Semaprax cheaper for coding agents?

Not proven. Semaprax is designed to let agents request typed, bounded context, but the current Agent Context v1 benchmark does not measure model tokens, latency, answer quality, accepted-task rate, or repository-scale cost. On its small corpus, the context artifact is larger than the source.

// 01

What the current benchmark measures

Frozen input

A small versioned Semaprax corpus and maintenance task

Deterministic output

Agent Context v1 JSON with explicit byte, node, and depth limits

Comparison boundary

Source bytes versus structured context bytes for the fixed corpus

// 02

What it does not measure

  • Model input or output tokens
  • Wall-clock latency or provider price
  • Answer correctness or accepted patches
  • Repository-scale navigation and maintenance
  • A comparison with Rust, C, C++, Go, or another language
// 03

The evidence needed before a cost claim

  1. Freeze equivalent maintenance tasks and repository snapshots.
  2. Run the same model, harness, tool permissions, and stopping rules.
  3. Measure billed tokens, elapsed time, accepted-task rate, and regression failures.
  4. Publish raw traces, exclusions, unsuccessful runs, and confidence intervals.
  5. Repeat across repository sizes and more than one model family.

GitHub specifications are the normative source. This page is a dated research summary. GitHub repository.

// FAQ

Questions, answered without the hype

Does Semaprax use fewer tokens than Rust or C?

There is no credible evidence for that comparison yet. The current benchmark does not run a model or compare languages, and its small structured context artifact is larger than the source.

Why publish a benchmark that does not show savings?

Because it fixes the measurement contract and exposes counterevidence before a marketing claim. That makes later model and repository-scale results auditable.

Research project by Wavect: Wavect GmbH. Created by Wavect as an open-source systems research project.