---
title: "Caveman 3.0 for Claude Code: Input Compression"
canonical: https://wavect.io/blog/caveman-3-claude-code-input-compression/
language: en
description: "How Caveman 3.0 compresses Claude Code tool output locally, preserves exact recovery, uses Apache-2.0, and what its 33.2% input-token benchmark really shows."
image: "https://wavect.io/img/general/bak/open_graph_preview.jpg"
---

[**Back**](/blog/overview/)

[![Kevin Riedl](/img/team/kevin.webp)](/team/kevin-riedl/)

[Kevin Riedl](/team/kevin-riedl/) https://linkedin.com/in/wsdt

15 min read · 2 Oct 2026 Last reviewed October 2, 2026

[**Next**](/blog/context-language-models-vs-compaction/)

# Caveman 3.0 for Claude Code: Local Input Compression, Recovery and Benchmarks

TL;DR

Caveman 3.0 is no longer just a terse-output skill. Its local proxy can compress logs, JSON, YAML, test output and other tool results before a coding agent sends them to the model, while keeping exact originals available for recovery. The project's pinned Claude Code benchmark reports 33.2% fewer provider-reported input tokens across 18 paired runs with 18/18 exact-answer checks, but one HTML workload used 9.9% more input and the raw benchmark artifacts are not published. Version 3.0 also relicenses the whole public repository under Apache-2.0 and stabilizes its TypeScript and Python middleware at 1.0.0. This guide explains what to pilot, what to measure, and which privacy and retention boundaries still matter.

**Reviewed 2 October 2026.** Product baseline: Caveman 3.0.0, runtime `bin-v2.0.0`, SDK 1.2.0 and middleware 1.0.0. This is a source and benchmark-method review, not a Wavect production benchmark. [Caveman 3.0.0 release](https://github.com/JuliusBrussee/caveman/releases/tag/v3.0.0)

**Caveman 3.0 is interesting for a different reason than the meme that made it famous.** The original idea was simple: make a coding agent answer with fewer words. The newer architecture targets the more expensive side of long agent sessions, what the model repeatedly reads. A local proxy can compress large tool results before the next model request while retaining the original bytes for recovery. [Caveman 3.0 README](https://github.com/JuliusBrussee/caveman/blob/v3.0.0/README.md)

That is a narrower search intent than our general [coding-agent token optimization guide](/blog/smarter-token-usage-with-your-ai-coding-agent/), which covers caching, routing and context reduction as a stack. It is also narrower than our [tool-output compression guide](/blog/codag-cost-control/), which explains the category. This article is specifically about **Caveman 3.0, Claude Code input compression, recoverability, local deployment and the evidence behind its claims**.

## What changed in Caveman 3.0?

Caveman 3.0.0 was published on 30 September 2026. The release turns several previously separate product decisions into one clearer deployment story. [Release notes](#source-release).

| Area | Caveman 3.0 | Why it matters |
| --- | --- | --- |
| License | The whole public repository is Apache-2.0 from 3.0.0 onward. | The engine, proxy, CLI, SDKs and middleware can be forked, embedded and self-hosted under one permissive license. |
| Middleware | TypeScript and Python packages reach 1.0.0 and require SDK 1.2.0. | Teams can put the same compression layer inside their own agent code instead of only wrapping a terminal agent. |
| `caveman learn` | Background refresh, memory checks, trends and consent-gated fixes. | The product can show where repeated context comes from before you decide what to compress. |
| Runtime | `bin-v2.0.0` adds stronger middleware lifecycle, retention and deployment primitives. | Recovery originals can be scoped and deleted with the session instead of becoming an unmanaged side store. |

The licensing change deserves precision. Before 3.0.0, parts of the repository used MIT while engine-linked runtime code used BSL-1.1. From 3.0.0 onward the repository is Apache-2.0, with no hosted-service restriction or change date for that code. Earlier releases keep the license terms they originally shipped with. Caveman Cloud is a separate hosted product and is not covered by the repository license. [Caveman licensing notes](https://github.com/JuliusBrussee/caveman/blob/v3.0.0/LICENSING.md)

## The expensive part of an agent session is often what it rereads

A coding agent does not only pay for its final explanation. Long sessions repeatedly carry forward instructions, previous messages and tool results. A test command that emits 80 KB once can influence multiple later requests if the framework keeps it in history. The same is true for JSON payloads, YAML configuration, diffs and search output.

Caveman separates two ideas that are easy to confuse:

- **The response skill** asks the agent to write terse prose. It attacks output verbosity.
- **The proxy and middleware** transform eligible context before inference. They attack input volume.

The second is more architecturally interesting. The model does not need every INFO line from a log to diagnose the one failing stack trace, but blindly truncating the log is dangerous because the missing line might be the answer. Caveman's design is therefore **compress plus recover**, not simply discard. [Project architecture](#source-readme).

## How the local proxy works

The high-level path is straightforward:

```
Claude Code / Codex / another agent
        |
        | request + tool results
        v
local Caveman proxy
        |-- classify eligible content
        |-- replace noisy blocks with shorter representations
        |-- store exact originals locally
        |-- expose recovery handles
        v
model provider chosen by the agent
```

The proxy still forwards inference traffic to the provider the agent was already using. Running Caveman locally therefore **does not make Claude, GPT or another hosted model local**. It gives you control over the transformation and recovery layer. If the upstream model remains a cloud API, the provider still receives the resulting model request. Caveman's security documentation makes that data-flow distinction explicit. [Security and privacy model](https://github.com/JuliusBrussee/caveman/blob/v3.0.0/SECURITY.md)

For teams building their own agents, middleware 1.0 applies the same idea without replacing the framework. The adapter copies an outbound request, projects eligible tool-result text into the compressed copy, and keeps the application's original conversation intact. Compression only activates on integration paths that can register a real recovery tool. If recovery cannot be bound, the documented behavior is to pass the content through rather than silently make it lossy. [TypeScript middleware 1.0](https://github.com/JuliusBrussee/caveman/blob/v3.0.0/packages/middleware/typescript/README.md) [Python middleware 1.0](https://github.com/JuliusBrussee/caveman/blob/v3.0.0/packages/middleware/python/README.md)

## What the 33.2% Claude Code benchmark actually shows

The strongest public evidence in the repository is not a claim that every block becomes 98 or 99 percent smaller. It is a paired agent-level benchmark that measures provider-reported input tokens after the whole session behavior plays out. [CaveBench wrap benchmark](https://github.com/JuliusBrussee/caveman/blob/v3.0.0/docs/WRAP-BENCHMARK.md)

| Workload | Direct input | Caveman input | Reported change | Answer checks |
| --- | --- | --- | --- | --- |
| Log needle | 148,807 | 74,068 | -50.2% | 3/3 |
| Deployment JSON | 147,975 | 108,939 | -26.4% | 3/3 |
| Fraud CSV | 165,823 | 74,484 | -55.1% | 3/3 |
| Test output | 150,377 | 108,514 | -27.8% | 3/3 |
| Configuration YAML | 132,124 | 71,027 | -46.2% | 3/3 |
| Dashboard HTML | 140,687 | 154,641 | **+9.9%** | 3/3 |

Across the 18 direct-versus-Caveman pairs, the report totals **885,793 direct input tokens versus 591,673 with Caveman**, a 33.2 percent reduction. All 18 exact-answer checks passed. The reported case-clustered 95 percent interval is 14.6 to 48.5 percent. The negative HTML row remains in the total because no useful compression transform applied while Caveman still added overhead. [Benchmark report and method](#source-benchmark).

That red row is useful. It demonstrates why a compression ratio on one payload is not the same thing as whole-session savings. If a workload is already compact, unsupported or dominated by other context, an extra layer can cost more tokens.

**There is also an evidence limit.** The repository publishes the benchmark table, method and provenance hashes, but says the raw harness and run artifacts are not in the checkout. The authors themselves classify it as a pinned report rather than a publicly reproducible benchmark. We therefore use the 33.2 percent figure as a project-reported controlled result, not as a universal production forecast.

The launch material also circulates more dramatic examples of individual blocks shrinking by roughly 98 to 99 percent. Those may be useful demonstrations of a specific transform, but they should not be multiplied directly into a cloud bill. The repository-level benchmark above is the better basis for an engineering decision because it includes session overhead, no-op cases and recovery behavior.

## Why 98% smaller context does not mean a 98% cheaper agent

Four layers sit between a compressed log and an invoice:

1. **Eligibility.** Not every message or tool result is compressible.
2. **Session share.** System instructions, code, user prompts, prior turns and tool schemas still consume context.
3. **Provider caching.** Some repeated input may already be served from prompt cache at a different price or compute profile.
4. **Recovery and retries.** If the model needs the omitted bytes, a retrieval adds another tool step and more context.

The operational KPI should therefore be **provider-reported input tokens per accepted task**, paired with answer quality and retry rate. A local tokenizer estimate is useful for diagnosis, but it is not the same thing as provider billing. Caveman's own benchmark makes that distinction and uses provider usage counters for its primary comparison. [Measurement method](#source-benchmark).

## What `caveman learn` adds before compression

The most useful part of the 3.0 story may be diagnostic rather than compressive. `caveman learn` reads local agent history and setup files to rank repeated token sinks. It supports Claude Code, Codex, Gemini CLI, opencode and aider with documented source locations. The analysis itself runs locally and labels its estimates as `inferred`, not verified savings. [caveman learn technical documentation](https://github.com/JuliusBrussee/caveman/blob/v3.0.0/docs/technical/learn.md)

That matters because the cheapest context is often context you never needed to load. A 400-line `CLAUDE.md`, unused skill descriptions or the same procedure pasted into every session can be a bigger structural problem than one noisy test command. Learn can surface those habits before you add another optimizer.

```
caveman learn --plain --since 7d --sources claude,codex
caveman learn --all
caveman learn implement
```

Version 3.0 also adds an autopilot that can rescan in the background after sessions, at most once every six hours. The release notes say proposed changes still require consent, are re-measured and are undone if they do not make messages smaller. You can disable the background refresh with `caveman learn autopilot off`. [3.0 release notes](#source-release).

## Local does not mean zero data policy work

Caveman's local mode removes a third-party Caveman service from the inference path, but there are still three separate data boundaries to review.

### 1. The model provider still receives the model request

If Claude Code is still connected to Anthropic, the compressed request goes to Anthropic. If your application uses OpenAI, the request still goes to OpenAI. Caveman does not change the provider's role just because compression happens first. For fully local inference you need a local or self-hosted model endpoint as well.

### 2. Recovery originals are sensitive local data

The runtime stores exact tool-result originals so the agent can recover them. From `bin-v2.0.0`, middleware originals are scope-owned, retention can cover those originals, and deleting a session can delete its stored originals. Encryption at rest is available when an encryption key is configured. Without a key, storage relies on filesystem protections. The security document explicitly notes that pre-`bin-v2.0.0` lifecycle behavior was weaker, so pin the runtime when retention requirements matter. [Runtime storage lifecycle](#source-security).

### 3. CLI telemetry is on by default

The standalone skill sends nothing by itself. The 3.0 CLI and agent hooks, however, use opt-out telemetry. The release and security documentation say events include a random install identifier, token aggregates and the sender IP address. They exclude prompts, code, file paths, tool arguments and tool results. The project accurately describes this telemetry as pseudonymous rather than anonymous. CI does not send by default. [Telemetry details](#source-security).

```
caveman telemetry status
caveman telemetry off
# or
export DO_NOT_TRACK=1
```

For an enterprise pilot, decide this setting before the first production-like run instead of treating it as cleanup later.

## A practical Caveman 3.0 pilot for Claude Code

Do not start by enabling compression for every developer. Build a paired test that makes a quality regression visible.

### Step 1. Pin the version and inspect what will run

```
npm install -g @caveman-ai/cli@2.0.0
caveman setup
caveman telemetry off   # if that matches your policy
```

The 3.0 release maps CLI 2.0.0 to runtime `bin-v2.0.0`. Pinning versions matters because retention behavior, middleware contracts and adapters can change. [Version matrix](#source-release).

### Step 2. Measure your own token sinks first

```
caveman learn --plain --since 7d
```

If most waste comes from always-loaded instructions, fix that before interpreting proxy compression as the answer to everything.

### Step 3. Choose tasks with exact acceptance criteria

Good pilot tasks include: find one fatal line in a large log, identify one JSON configuration drift, isolate the failing test, or extract a specific CSV outlier. Each task should have an oracle that another script can verify. Avoid judging quality only by whether the answer sounds plausible.

### Step 4. Run direct and compressed arms

Keep the model, agent version, task fixtures, tool permissions and starting repository state fixed. Record provider-reported input tokens, cache reads, cache writes, output tokens, latency, retries and recovery calls. Rotate run order to reduce warm-cache and time-of-day effects.

### Step 5. Gate rollout on accepted-task economics

| Metric | What to require |
| --- | --- |
| Task correctness | No statistically or operationally meaningful drop on held-out tasks. |
| Provider input | Lower median and aggregate provider-reported input tokens on compressible workloads. |
| Negative cases | Unsupported or already-small payloads must remain visible, not removed from the report. |
| Recovery | Exact original can be recovered when the short view is insufficient. |
| Retries | No rise that erases token savings or increases human correction time. |
| Storage policy | Retention, deletion, encryption and telemetry settings match your data classification. |

## Using Caveman 1.0 middleware in your own agent

For custom agents, the main design decision is whether your framework integration can bind recovery. Caveman's TypeScript middleware documents certified compression paths for Vercel AI SDK, OpenAI, Anthropic and LangChain, with other adapters marked experimental. Python documents certified paths for LangChain, OpenAI, Anthropic and LiteLLM. Unsupported versions and runtime failures normally fail open, which means the original request proceeds unchanged and a reason is logged. [TypeScript compatibility contract](#source-ts) [Python compatibility contract](#source-py).

A good rollout uses three modes:

- `record` to validate integration without changing model input.
- `compress` for the treatment arm.
- `off` as the explicit rollback path.

The official quickstarts are unusually useful because their default demos make no provider request. They exercise the real local runtime, compression and exact paginated recovery first, then leave paid provider execution as an optional separate step. That is the right order for an infrastructure pilot.

## Does Apache-2.0 mean you now “own your AI”?

It gives teams materially more control over this layer. You can inspect, fork, modify and embed Caveman 3.0's public runtime without the previous BSL service restriction. That matters for internal platforms and teams that do not want a proprietary optimizer in the middle of every agent request. [License scope](#source-license).

But ownership has layers. You can own the compression proxy and still rent the model. You can self-host the model and still depend on proprietary developer tools. You can own both and still route telemetry or browser data through external services. The useful question is not “is this AI ours?” but **which parts of the inference, context, memory, tooling and observability path can we inspect, replace and operate ourselves?**

Caveman 3.0 improves that answer for the context-efficiency layer. It does not solve the rest automatically.

## When Caveman is worth testing, and when it is not

| Test Caveman when... | Start elsewhere when... |
| --- | --- |
| Agent sessions repeatedly ingest large logs, test output, JSON or YAML. | Most requests are short and already fit comfortably in context. |
| You can define exact or testable task outcomes. | You cannot tell whether compressed runs are still correct. |
| You need byte-exact local recovery rather than irreversible truncation. | Your policy does not allow local retention of tool-output originals. |
| You want an Apache-2.0 layer you can fork or embed. | Your real cost problem is frontier-model overuse, poor caching or unnecessary call count. |
| You are building an agent on a supported middleware path. | Your framework version sits outside the tested adapter range and you cannot validate it yourself. |

For the broader sequence, our [token usage playbook](/blog/smarter-token-usage-with-your-ai-coding-agent/) starts with cacheability and routing before compression. For a broader comparison of the compression category, use [our tool-output compression guide](/blog/codag-cost-control/). For model-managed context rather than proxy-managed compression, see [Context Language Models versus compaction](/blog/context-language-models-vs-compaction/).

## FAQ

### Is Caveman 3.0 fully open source?

The public repository from version 3.0.0 onward is licensed Apache-2.0, including the engine, proxy, CLI, SDKs and middleware. Earlier releases retain their original licenses. Caveman Cloud is separate commercial software and is not covered by the repository license. [License details](#source-license).

### Does Caveman reduce Claude Code input tokens?

In Caveman's published six-workload benchmark, the wrapped Claude Code arm used 33.2 percent fewer provider-reported input tokens in aggregate than direct Claude Code and passed 18/18 exact-answer checks. The result is workload-specific and the raw run artifacts are not public, so treat it as a controlled project-reported result, not a universal saving. [Benchmark](#source-benchmark).

### Does Caveman send my code to Caveman servers?

The local proxy and framework middleware do not need a Caveman account and do not send prompt or tool-result content to Caveman servers. Your selected model provider still receives the inference request. The CLI has separate opt-out telemetry that includes usage metadata and IP address but excludes prompts, code and file paths. [Security model](#source-security).

### Can the agent retrieve compressed-away content?

Yes, supported compression paths preserve exact originals and bind a recovery mechanism. Middleware paths that cannot bind recovery are documented to pass content through instead of compressing it. Recovery handles are scoped and can expire according to runtime retention. [Middleware recovery contract](#source-ts).

### Is `caveman learn` local?

The Learn analysis reads local session history and files and says it sends nothing anywhere. Its reports are local estimates. This is separate from CLI telemetry, which can be disabled with `caveman telemetry off` or `DO_NOT_TRACK=1`. [Learn documentation](#source-learn) [Telemetry documentation](#source-security).

### Is Caveman the same as prompt caching?

No. Prompt caching reuses provider-side computation for repeated prefixes. Caveman changes eligible context before the request is sent. They can complement each other, but compression can also change cache behavior, so measure provider cache-read and cache-write counters in the same pilot.

### Does Caveman middleware support LiteLLM?

The Python middleware 1.0 documentation lists LiteLLM as a certified adapter family for supported async and proxy paths, with explicit recovery requirements and provider limitations on some sync routes. Pin the tested package versions and run the compatibility checks before production. [Python middleware matrix](#source-py).

## Sources and verification notes

- [Caveman 3.0.0 release notes](https://github.com/JuliusBrussee/caveman/releases/tag/v3.0.0), versions, Apache-2.0 change, Learn autopilot and middleware 1.0.
- [LICENSING.md at v3.0.0](https://github.com/JuliusBrussee/caveman/blob/v3.0.0/LICENSING.md), repository and Cloud license boundaries.
- [CaveBench wrap benchmark](https://github.com/JuliusBrussee/caveman/blob/v3.0.0/docs/WRAP-BENCHMARK.md), provider input totals, exact-answer checks, confidence interval, negative HTML case and reproduction limits.
- [SECURITY.md at v3.0.0](https://github.com/JuliusBrussee/caveman/blob/v3.0.0/SECURITY.md), data flow, local recovery storage and telemetry.
- [caveman learn documentation](https://github.com/JuliusBrussee/caveman/blob/v3.0.0/docs/technical/learn.md), local inputs, outputs, commands and evidence labels.
- [TypeScript middleware 1.0 README](https://github.com/JuliusBrussee/caveman/blob/v3.0.0/packages/middleware/typescript/README.md), recovery and fail-open integration behavior.
- [Python middleware 1.0 README](https://github.com/JuliusBrussee/caveman/blob/v3.0.0/packages/middleware/python/README.md), certified adapters including LiteLLM and recovery constraints.
- [Caveman 3.0 README](https://github.com/JuliusBrussee/caveman/blob/v3.0.0/README.md), product architecture and installation surface.

## Final thoughts

Caveman 3.0 turns a joke about terse answers into a more serious context-efficiency layer. The Apache-2.0 relicensing, recoverable local compression and stable middleware make it worth testing where coding agents repeatedly ingest large tool outputs. The right rollout is not to trust a 98% block-compression screenshot. It is to run paired tasks, use provider-reported usage, keep the negative cases, verify exact recovery and reject any saving that comes with lower accepted-task quality.

## You may also like..

[**Smarter Token Usage with Your AI Coding Agent** A broader cost-control playbook covering caching, routing and context reduction.](/blog/smarter-token-usage-with-your-ai-coding-agent/) [**Context Language Models vs Compaction** A different approach where the model edits its own live context instead of relying on a proxy compressor.](/blog/context-language-models-vs-compaction/)

Models and infrastructure

## Continue through this cluster

Model selection, inference economics, local deployment, compression and serving architecture.

[Start with the cornerstone**Self-Hosting LLMs in the EU: When Open Weights Actually Pay Off**](/blog/self-hosting-llms-eu-cost/)

- [Cloudflare Clef vs Jev: Pricing, Benchmarks and Migration](/blog/cloudflare-clef-vs-jev/)
- [Context Language Models vs Compaction: What to Pilot](/blog/context-language-models-vs-compaction/)
- [LiteLLM Lens: Agent Trace Analysis with SQL and APIs](/blog/litellm-lens-agent-trace-analysis/)
- [SmythOS Studio Self-Hosting: Docker Setup, Costs and Limits](/blog/smythos-studio-self-hosting/)
- [DeerFlow 2.0: Docker Setup, Sandboxes and Memory](/blog/deerflow-2-docker-setup-sandbox-memory/)

[**Back**](/blog/overview/)

[![Kevin Riedl](/img/team/kevin.webp)](/team/kevin-riedl/)

[Kevin Riedl](/team/kevin-riedl/) https://linkedin.com/in/wsdt

15 min read · 2 Oct 2026 Last reviewed October 2, 2026

[**Next**](/blog/context-language-models-vs-compaction/)

## Structured Data

```json
{
  "@context": "https://schema.org",
  "@graph": [
    {
      "@id": "https://wavect.io/#organization",
      "@type": [
        "Organization",
        "ProfessionalService",
        "LocalBusiness"
      ],
      "employee": [
        {
          "@id": "https://wavect.io/team/kevin-riedl/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Kevin Riedl",
          "url": "https://wavect.io/team/kevin-riedl/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        },
        {
          "@id": "https://wavect.io/team/christof-jori/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Christof Jori",
          "url": "https://wavect.io/team/christof-jori/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        }
      ],
      "founder": [
        {
          "@id": "https://wavect.io/team/kevin-riedl/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Kevin Riedl",
          "url": "https://wavect.io/team/kevin-riedl/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        },
        {
          "@id": "https://wavect.io/team/christof-jori/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Christof Jori",
          "url": "https://wavect.io/team/christof-jori/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        }
      ],
      "legalRepresentative": [
        {
          "@id": "https://wavect.io/team/kevin-riedl/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Kevin Riedl",
          "url": "https://wavect.io/team/kevin-riedl/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        },
        {
          "@id": "https://wavect.io/team/christof-jori/#person",
          "@type": "Person",
          "jobTitle": "Managing Director",
          "name": "Christof Jori",
          "url": "https://wavect.io/team/christof-jori/",
          "worksFor": {
            "@id": "https://wavect.io/#organization",
            "@type": [
              "Organization",
              "ProfessionalService",
              "LocalBusiness"
            ]
          }
        }
      ],
      "name": "Wavect GmbH",
      "subjectOf": {
        "@id": "https://wavect.io/verified-claims.json#dataset",
        "@type": "Dataset",
        "creator": {
          "@id": "https://wavect.io/#organization",
          "@type": [
            "Organization",
            "ProfessionalService",
            "LocalBusiness"
          ]
        },
        "description": "A machine-readable registry of quantitative and qualitative claims published by Wavect, with review dates, localized page appearances and public third-party citations where available.",
        "inLanguage": "en",
        "isAccessibleForFree": true,
        "license": "https://creativecommons.org/licenses/by/4.0/",
        "name": "Wavect verified publication claims",
        "url": "https://wavect.io/verified-claims.json"
      },
      "url": "https://wavect.io/"
    },
    {
      "@id": "https://wavect.io/team/kevin-riedl/#person",
      "@type": "Person",
      "jobTitle": "Managing Director",
      "name": "Kevin Riedl",
      "sameAs": [
        "https://www.wikidata.org/wiki/Q139796365",
        "https://www.linkedin.com/in/wsdt",
        "https://github.com/wsdt"
      ],
      "url": "https://wavect.io/team/kevin-riedl/",
      "worksFor": {
        "@id": "https://wavect.io/#organization",
        "@type": [
          "Organization",
          "ProfessionalService",
          "LocalBusiness"
        ]
      }
    },
    {
      "@id": "https://wavect.io/team/christof-jori/#person",
      "@type": "Person",
      "jobTitle": "Managing Director",
      "name": "Christof Jori",
      "sameAs": [
        "https://www.wikidata.org/wiki/Q139796367",
        "https://www.linkedin.com/in/jocr77/",
        "https://github.com/jo-chris"
      ],
      "url": "https://wavect.io/team/christof-jori/",
      "worksFor": {
        "@id": "https://wavect.io/#organization",
        "@type": [
          "Organization",
          "ProfessionalService",
          "LocalBusiness"
        ]
      }
    },
    {
      "@id": "https://wavect.io/#website",
      "@type": "WebSite",
      "inLanguage": [
        "en",
        "de",
        "es",
        "zh"
      ],
      "name": "Wavect",
      "potentialAction": {
        "@type": "SearchAction",
        "query-input": "required name=search_term_string",
        "target": {
          "@type": "EntryPoint",
          "urlTemplate": "https://wavect.io/search/?q={search_term_string}"
        }
      },
      "publisher": {
        "@id": "https://wavect.io/#organization",
        "@type": [
          "Organization",
          "ProfessionalService",
          "LocalBusiness"
        ]
      },
      "url": "https://wavect.io/"
    },
    {
      "@id": "https://wavect.io/blog/caveman-3-claude-code-input-compression/#webpage",
      "@type": "WebPage",
      "dateModified": "2026-10-04",
      "inLanguage": "en",
      "isPartOf": {
        "@id": "https://wavect.io/#website",
        "@type": "WebSite"
      },
      "lastReviewed": "2026-10-04",
      "url": "https://wavect.io/blog/caveman-3-claude-code-input-compression/"
    }
  ]
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "BlogPosting",
  "abstract": "Caveman 3.0 is no longer just a terse-output skill. Its local proxy can compress logs, JSON, YAML, test output and other tool results before a coding agent sends them to the model, while keeping exact originals available for recovery. The project's pinned Claude Code benchmark reports 33.2% fewer provider-reported input tokens across 18 paired runs with 18/18 exact-answer checks, but one HTML workload used 9.9% more input and the raw benchmark artifacts are not published. Version 3.0 also relicenses the whole public repository under Apache-2.0 and stabilizes its TypeScript and Python middleware at 1.0.0. This guide explains what to pilot, what to measure, and which privacy and retention boundaries still matter.",
  "articleBody": " Blog overview/AI and agents/Models and infrastructure Caveman 3.0 for Claude Code: Local Input Compression, Recovery and Benchmarks TL;DR Caveman 3.0 is no longer just a terse-output skill. Its local proxy can compress logs, JSON, YAML, test output and other tool results before a coding agent sends them to the model, while keeping exact originals available for recovery. The project's pinned Claude Code benchmark reports 33.2% fewer provider-reported input tokens across 18 paired runs with 18/18 exact-answer checks, but one HTML workload used 9.9% more input and the raw benchmark artifacts are not published. Version 3.0 also relicenses the whole public repository under Apache-2.0 and stabilizes its TypeScript and Python middleware at 1.0.0. This guide explains what to pilot, what to measure, and which privacy and retention boundaries still matter. Reviewed 2 October 2026. Product baseline: Caveman 3.0.0, runtime bin-v2.0.0, SDK 1.2.0 and middleware 1.0.0. This is a source and benchmark-method review, not a Wavect production benchmark. Caveman 3.0.0 release Caveman 3.0 is interesting for a different reason than the meme that made it famous. The original idea was simple: make a coding agent answer with fewer words. The newer architecture targets the more expensive side of long agent sessions, what the model repeatedly reads. A local proxy can compress large tool results before the next model request while retaining the original bytes for recovery. Caveman 3.0 README That is a narrower search intent than our general coding-agent token optimization guide, which covers caching, routing and context reduction as a stack. It is also narrower than our tool-output compression guide, which explains the category. This article is specifically about Caveman 3.0, Claude Code input compression, recoverability, local deployment and the evidence behind its claims. What changed in Caveman 3.0? Caveman 3.0.0 was published on 30 September 2026. The release turns several previously separate product decisions into one clearer deployment story. Release notes. Caveman 3.0 changes that matter for engineering teams AreaCaveman 3.0Why it matters LicenseThe whole public repository is Apache-2.0 from 3.0.0 onward.The engine, proxy, CLI, SDKs and middleware can be forked, embedded and self-hosted under one permissive license. MiddlewareTypeScript and Python packages reach 1.0.0 and require SDK 1.2.0.Teams can put the same compression layer inside their own agent code instead of only wrapping a terminal agent. caveman learnBackground refresh, memory checks, trends and consent-gated fixes.The product can show where repeated context comes from before you decide what to compress. Runtimebin-v2.0.0 adds stronger middleware lifecycle, retention and deployment primitives.Recovery originals can be scoped and deleted with the session instead of becoming an unmanaged side store. The licensing change deserves precision. Before 3.0.0, parts of the repository used MIT while engine-linked runtime code used BSL-1.1. From 3.0.0 onward the repository is Apache-2.0, with no hosted-service restriction or change date for that code. Earlier releases keep the license terms they originally shipped with. Caveman Cloud is a separate hosted product and is not covered by the repository license. Caveman licensing notes The expensive part of an agent session is often what it rereads A coding agent does not only pay for its final explanation. Long sessions repeatedly carry forward instructions, previous messages and tool results. A test command that emits 80 KB once can influence multiple later requests if the framework keeps it in history. The same is true for JSON payloads, YAML configuration, diffs and search output. Caveman separates two ideas that are easy to confuse: The response skill asks the agent to write terse prose. It attacks output verbosity. The proxy and middleware transform eligible context before inference. They attack input volume. The second is more architecturally interesting. The model does not need every INFO line from a log to diagnose the one failing stack trace, but blindly truncating the log is dangerous because the missing line might be the answer. Caveman's design is therefore compress plus recover, not simply discard. Project architecture. How the local proxy works The high-level path is straightforward: Claude Code / Codex / another agent | | request + tool results v local Caveman proxy |-- classify eligible content |-- replace noisy blocks with shorter representations |-- store exact originals locally |-- expose recovery handles v model provider chosen by the agent The proxy still forwards inference traffic to the provider the agent was already using. Running Caveman locally therefore does not make Claude, GPT or another hosted model local. It gives you control over the transformation and recovery layer. If the upstream model remains a cloud API, the provider still receives the resulting model request. Caveman's security documentation",
  "articleSection": "AI Agents",
  "author": {
    "@id": "https://wavect.io/team/kevin-riedl/#person",
    "@type": "Person",
    "name": "Kevin Riedl",
    "sameAs": [
      "https://www.wikidata.org/wiki/Q139796365",
      "https://www.linkedin.com/in/wsdt",
      "https://github.com/wsdt"
    ],
    "url": "https://wavect.io/team/kevin-riedl/"
  },
  "citation": [
    {
      "@type": "WebPage",
      "name": "Caveman 3.0.0 release",
      "url": "https://github.com/JuliusBrussee/caveman/releases/tag/v3.0.0"
    },
    {
      "@type": "WebPage",
      "name": "Caveman 3.0 README",
      "url": "https://github.com/JuliusBrussee/caveman/blob/v3.0.0/README.md"
    },
    {
      "@type": "WebPage",
      "name": "Caveman licensing notes",
      "url": "https://github.com/JuliusBrussee/caveman/blob/v3.0.0/LICENSING.md"
    },
    {
      "@type": "WebPage",
      "name": "Security and privacy model",
      "url": "https://github.com/JuliusBrussee/caveman/blob/v3.0.0/SECURITY.md"
    },
    {
      "@type": "WebPage",
      "name": "TypeScript middleware 1.0",
      "url": "https://github.com/JuliusBrussee/caveman/blob/v3.0.0/packages/middleware/typescript/README.md"
    },
    {
      "@type": "WebPage",
      "name": "Python middleware 1.0",
      "url": "https://github.com/JuliusBrussee/caveman/blob/v3.0.0/packages/middleware/python/README.md"
    },
    {
      "@type": "WebPage",
      "name": "CaveBench wrap benchmark",
      "url": "https://github.com/JuliusBrussee/caveman/blob/v3.0.0/docs/WRAP-BENCHMARK.md"
    },
    {
      "@type": "WebPage",
      "name": "caveman learn technical documentation",
      "url": "https://github.com/JuliusBrussee/caveman/blob/v3.0.0/docs/technical/learn.md"
    }
  ],
  "dateModified": "2026-10-02",
  "datePublished": "2026-10-02",
  "description": "Caveman 3.0 is no longer just a terse-output skill. Its local proxy can compress logs, JSON, YAML, test output and other tool results before a coding agent sends them to the model, while keeping exact originals available for recovery. The project's pinned Claude Code benchmark reports 33.2% fewer provider-reported input tokens across 18 paired runs with 18/18 exact-answer checks, but one HTML workload used 9.9% more input and the raw benchmark artifacts are not published. Version 3.0 also relicenses the whole public repository under Apache-2.0 and stabilizes its TypeScript and Python middleware at 1.0.0. This guide explains what to pilot, what to measure, and which privacy and retention boundaries still matter.",
  "headline": "Caveman 3.0 for Claude Code: Local Input Compression, Recovery and Benchmarks",
  "image": "https://wavect.io/img/blog/headers/header_caveman-3-claude-code-input-compression.svg",
  "inLanguage": "en",
  "keywords": "AI Coding Agents, Context Compression, Claude Code",
  "mainEntityOfPage": {
    "@id": "https://wavect.io/blog/caveman-3-claude-code-input-compression/",
    "@type": "WebPage"
  },
  "publisher": {
    "@id": "https://wavect.io/#organization",
    "@type": [
      "Organization",
      "ProfessionalService",
      "LocalBusiness"
    ]
  },
  "url": "https://wavect.io/blog/caveman-3-claude-code-input-compression/",
  "wordCount": 3282
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "BreadcrumbList",
  "itemListElement": [
    {
      "@type": "ListItem",
      "item": "https://wavect.io/",
      "name": "Home",
      "position": 1
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/overview/",
      "name": "Blog overview",
      "position": 2
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/topics/ai-agents/",
      "name": "AI and agents",
      "position": 3
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/clusters/models-infrastructure/",
      "name": "Models and infrastructure",
      "position": 4
    },
    {
      "@type": "ListItem",
      "item": "https://wavect.io/blog/caveman-3-claude-code-input-compression/",
      "name": "Caveman 3.0 for Claude Code: Input Compression",
      "position": 5
    }
  ]
}
```
