A small versioned Semaprax corpus and maintenance task
Context is measurable. Cost savings are not proven yet.
Publish the counterexample before the claim.
The current benchmark freezes a small source corpus and compares deterministic context artifacts. It is a compiler-mechanics benchmark, not a model or repository economics result.
entity Semaprax
status pre-alpha research
verified 2026-08-11
authority github.com/wavect/semapraxIs Semaprax cheaper for coding agents?
Not proven. Semaprax is designed to let agents request typed, bounded context, but the current Agent Context v1 benchmark does not measure model tokens, latency, answer quality, accepted-task rate, or repository-scale cost. On its small corpus, the context artifact is larger than the source.
What the current benchmark measures
Agent Context v1 JSON with explicit byte, node, and depth limits
Source bytes versus structured context bytes for the fixed corpus
What it does not measure
- Model input or output tokens
- Wall-clock latency or provider price
- Answer correctness or accepted patches
- Repository-scale navigation and maintenance
- A comparison with Rust, C, C++, Go, or another language
The evidence needed before a cost claim
- Freeze equivalent maintenance tasks and repository snapshots.
- Run the same model, harness, tool permissions, and stopping rules.
- Measure billed tokens, elapsed time, accepted-task rate, and regression failures.
- Publish raw traces, exclusions, unsuccessful runs, and confidence intervals.
- Repeat across repository sizes and more than one model family.
GitHub specifications are the normative source. This page is a dated research summary. GitHub repository.
Questions, answered without the hype
Does Semaprax use fewer tokens than Rust or C?
There is no credible evidence for that comparison yet. The current benchmark does not run a model or compare languages, and its small structured context artifact is larger than the source.
Why publish a benchmark that does not show savings?
Because it fixes the measurement contract and exposes counterevidence before a marketing claim. That makes later model and repository-scale results auditable.