Back
Kevin Riedl

10 min read · 7 Oct 2026
Last reviewed

Next
Made on your device, with no Instagram connection. We copy the post link for Instagram’s Link sticker.

Claude Model Router: What Switches the Model, and When?

You install a Claude model router, see a recommendation for a cheaper model, and the main conversation keeps using the expensive one. That can be expected behavior. The missing detail is which boundary the router controls: the active conversation, a future session, a delegated task or an outgoing API request.

Two similarly named community projects illustrate the distinction. tzachbon/claude-model-router-hook recommends models and routes subagent spawns; its opt-in main-session autoswitch affects new sessions. musistudio/claude-code-router, usually called CCR, is a local gateway that routes requests to configured models and providers. The repository owner is part of the identity: searching for “Claude Model Router” alone can lead to different software.

Reviewed on . This guide compares documented behavior and gives an evaluation workflow. The configuration examples are checked against those sources; they are not a live-provider benchmark or a claim that Wavect deployed either project for a client.

What does a Claude model router actually switch?

Ask what changes, when it changes, and how you can verify it. A routing recommendation, a saved setting and an actual model response are three different pieces of evidence.

Model-routing boundaries in Claude Code
BoundaryTypical controlEvidence to inspect
Current conversationNative /model selectionThe active model after the change
Next sessionA saved default, including the hook’s opt-in autoswitchThe model selected by a fresh launch
Delegated taskSubagent model configuration or a spawn hookThe model actually used by that subagent
Provider requestA gateway rule and upstream configurationThe resolved provider, model and response usage

First choose the boundary you need. A team wanting cheaper file-location tasks may only need explicit subagent models. A team needing multiple providers, credentials and request-level fallback has a different infrastructure requirement. Our LLM gateway and router comparison covers that broader purchase and architecture decision.

Can Claude Code switch models without a plugin?

Yes. Anthropic’s model configuration reference documents the launch flag, in-session selection and opusplan. That alias uses Opus for plan mode and Sonnet for execution. It is a specific phase-based policy, not a general classifier for every prompt.

For a native baseline, start a session with:

claude --model sonnet

Inside an interactive session, use /model opus to select Opus, or /model opusplan to select the planning/execution policy. Model availability and organization settings still apply. Use aliases for a quick experiment; record the resolved model version for comparisons you need to reproduce.

For a bounded delegated task, create .claude/agents/repo-locator.md. This example follows Anthropic’s subagent configuration reference and is our suggested starting role:

---
name: repo-locator
description: Locate files and symbols for a bounded repository question.
tools: Read, Glob, Grep
model: haiku
---
Find the requested files or symbols. Return paths and short evidence.
Do not edit files or infer a fix. Escalate ambiguity to the parent agent.

Ask Claude to use repo-locator for a narrow lookup. Then inspect the subagent’s actual model. A per-invocation model can override its frontmatter, and organization restrictions can cause substitution. The frontmatter expresses your intent; it is not a receipt proving that intent was executed.

Why is Claude Model Router Hook not switching my current session?

Its documented main-session behavior is warning or next-session configuration. The project changelog describes autoswitch writing the default for new sessions. Use native /model when you want to change the current conversation.

The same changelog records the August 2026 default routing update: mechanical tasks use Haiku, while implementation, debugging and broader reasoning use Opus with different effort levels. An old tutorial promising that ordinary implementation automatically goes to Sonnet may describe a previous policy.

A small project-level configuration in .claude/model-router.json can make the behavior explicit:

{
  "version": 2,
  "apply_mode": "warn",
  "subagent_enforcement": "on",
  "classifier": {
    "cli_fallback": false
  }
}

This is our evaluation configuration, not the upstream defaults. It keeps main-session warnings and subagent routing enabled while turning off the optional Claude CLI classifier fallback. apply_mode: warn does not mean that subagent routing is disabled. The project README documents these separate controls and the log at ~/.claude/hooks/model-router-hook.log.

Check the hook alongside other installed automation. Anthropic’s hooks reference documents how PreToolUse can rewrite tool inputs. Command hooks also execute with your user account’s permissions. Review the installed commands and existing hooks before introducing another component that rewrites agent spawns.

What changes when you use Claude Code Router as a proxy?

The request path changes. CCR receives the model request and forwards it according to its configuration. “Local gateway” describes where that component runs; the selected upstream determines where inference happens. A cloud provider still receives the material sent to its model.

The current CCR CLI guide separates management access from model-request access: the management token and CCR client keys are independent credentials. Treat access to the management interface and permission to call models as separate decisions.

Anthropic’s gateway documentation explains that billing follows the credential forwarded upstream. Setting ANTHROPIC_BASE_URL alone does not replace a saved subscription credential. A router therefore does not make another provider’s inference part of your Claude subscription.

Also test more than a successful greeting. The official gateway compatibility guide covers streaming, capability headers, request fields and cache markers. Broken forwarding can cause errors or remove expected behavior. Run a real tool round trip and inspect the resolved route and usage before trusting a provider combination.

Why can a cheaper Claude model increase task cost?

A model switch can replace cheap cache reads with a new cache write. Anthropic’s API cost optimization cookbook explains that caches are per model and that a subagent starts with a fresh prefix. Keeping the conversation text does not transfer the parent model’s cache.

Consider a long investigation followed by a small formatting task. Routing the complete investigation history to a cheaper model might cost more than formatting the result where the cache is already warm. Delegating a short, self-contained excerpt creates another option. Compare both against the same accepted output instead of assuming the lower price wins.

Effort needs the same care. The prompt-caching reference distinguishes top-level effort changes, which invalidate cached message blocks, from supported per-message effort changes that preserve the earlier prefix. “Same model” does not automatically mean “same cache.”

Our evaluation rule: count classification, generation, cache writes, retries and human correction across the whole task. Keep model routing at a useful boundary unless measured quality or total cost justifies another switch. A fixed savings percentage cannot tell you whether your own workload benefits.

Claude model router troubleshooting

Symptoms and the next evidence to inspect
SymptomFirst check
Recommendation changes, main model does notCheck whether the hook is warning or configuring the next session. Use /model for an active-session change.
Fresh sessions still start on the old modelInspect launch flags, environment variables, project settings and organization policy that can outrank a saved user default.
Subagent runs on an unexpected modelCompare the spawn’s explicit model, agent frontmatter, restrictions and resolved model.
Proxy chat works but tools or streams failCheck protocol forwarding and model capabilities with a complete tool round trip.
Lower model price, higher billCompare cache writes, extra attempts, classifier usage and duplicated context.

These checks follow the native configuration, subagent and gateway contracts above. Record what happened before changing several settings at once.

How should a team evaluate automatic model routing?

  1. Choose one workload. Separate file lookups, bounded edits and ambiguous debugging; give each an acceptance check.
  2. Record a native baseline. Keep the repository snapshot and task instructions fixed. Log actual models, completion time and correction work.
  3. Change one boundary. Compare explicit subagents with hook routing before adding provider switching to the same experiment.
  4. Inspect exceptions. Include an explicit model override, an ambiguous request, a long warm-cache session and a provider failure in a test environment.
  5. Retain evidence. Store route reasons, resolved models, usage and acceptance outcomes. Recheck the policy after client, plugin or model updates.

A useful rollout starts with one workflow whose mistakes are easy to spot. Expand when the routed version meets the quality target and improves total cost or completion time. For application-owned agent loops, our LiteAgents routing guide addresses the SDK boundary. Our OpenAI Decisions API guide covers classifier outputs and fallback policy when your application owns the decision.

Claude model routing: common questions

Is Claude Model Router an official Anthropic feature?

The two projects compared here, tzachbon/claude-model-router-hook and musistudio/claude-code-router, are community projects. Claude Code separately provides native model selection, subagent configuration and other controls. Check the repository owner before following installation instructions.

Can I switch Claude Code models without restarting?

Yes, native /model changes the current session’s selection. The hook project’s opt-in autoswitch instead writes a default for new sessions. Its subagent routing is a separate mechanism.

Does warn mode disable subagent routing?

No. In the example, apply_mode is warn while subagent_enforcement is on. The main session receives recommendations and delegated tasks can still be routed. Configure both controls deliberately.

Does model routing automatically save tokens?

No. Selecting a lower-priced model changes the price of its tokens, not necessarily the number required. Classification, cache rebuilds, retries and extra review can offset savings. Compare complete tasks that meet the same acceptance criteria.

Does a local Claude router keep all code on my computer?

Only if the complete configured inference path does. A local gateway can forward code to a cloud provider. Check the selected upstream, classifier calls, logs and fallback destinations before sending confidential material.

Production AI help

Building an AI product and worried about inference cost, architecture, or production readiness? Wavect helps founders turn AI prototypes into reliable production systems.

Explore the service path:

Inbox, without the noise

Follow the work that matters to you

Get a short email when we publish something new. Follow the whole blog or only the problems you care about.

What would you like to receive?
Choose your topics

Free, double opt-in, no tracking pixels.

Back
Kevin Riedl

10 min read · 7 Oct 2026
Last reviewed

Next

Get the next AI and agents field note

One concise email when we publish. No tracking pixels, and no inbox filler.

Free, double opt-in, no tracking pixels.