Miser is an open-source, Rust-based AI gateway that intelligently routes OpenAI-compatible requests to the optimal model through OpenRouter. Powered by Jev (TypeSafe's System One evaluation model), Miser delivers the best routing accuracy in the industry.
Miser uses Jev to classify every prompt into a complexity tier (trivial → simple → standard → hard → reasoning) and routes to the cheapest capable model. The results speak for themselves:
Held-out benchmark (116 adversarial cases in evals/classifier_cases.jsonl, never used for tuning):
| Router | Exact Accuracy | Adjacent Accuracy | Under-route | Over-route | Latency (p50) |
|---|---|---|---|---|---|
| Miser (Jev) | 89.7% | 100% | 6.9% | 3.5% | 338ms |
| Miser (heuristic) | 83.6% | 94.8% | 4.3% | 12.1% | <1ms |
| OpenRouter Auto † | 52.0% | 84.0% | 32.0% | - | 4.16s |
Reproduce with:
# heuristic (free, offline)
cargo run -p miser-evals -- --corpus evals/classifier_cases.jsonl --mode heuristic
# Jev (needs JEV_API_KEY; ~116 calls, so it costs a little)
cargo run -p miser-evals -- --corpus evals/classifier_cases.jsonl \
--mode jev --config config/miser.toml --concurrency 8† the OpenRouter Auto row is a previously published figure and is not re-measured here; only the two Miser rows are. Jev's probabilities vary slightly between runs, so treat its accuracy as ±1%.
What this means:
- 100% adjacent accuracy — Miser never routes more than 1 tier away from optimal
- 6.9% under-routing — hard work rarely goes to weak models (vs 32% for OpenRouter Auto)
- 3.5% over-routing — trivial prompts don't waste money on frontier models (vs 12.1% heuristic)
- 3.5× less over-routing than the regex heuristic
90 paired comparisons on held-out prompts stratified across all five tiers, both
routers answering the same ones, each answer scored 1–5 by openai/gpt-4.1
(a model neither router selected, so it is not judging its own output).
jev-router is constrained off the free tier for fairness.
| quality (1–5) | cost | better / worse / tied | |
|---|---|---|---|
| Miser | 3.80 | $0.0255 | 39 / 14 / 37 |
typesafe/jev-router |
3.20 | $0.0299 | — |
Verdict: better on quality at comparable cost. +0.60 on a 5-point judge scale, winning 39 to 14, for 1.17× the cost. The quality gap is the real result and it holds regardless of how the opponent is configured; the cost advantage is modest, not dramatic.
Miser escalates per tier; jev-router concentrates trivial, standard and hard
work on one model (Write a threat model for the payment processing service
→ stealth/space-bunny-alpha, 8/8 trials) and reserves the frontier tier for
proofs. That concentration is its own decision, not an artifact of account
settings — explicit model requests survive exactly, and its choice is stable per
prompt (8/8).
Reproduce with scripts/head-to-head-paid.py (needs OPENROUTER_API_KEY and a
running gateway on :8787).
Caveats. Single run; the judge is an LLM with no human or inter-rater validation; both routers capped at 160 output tokens; and the corpus is short factual and small-coding prompts, so it does not exercise the agentic workloads jev-router targets. Treat this as evidence, not proof.
| Corpus | Cases | Mode | Exact | Under | Over |
|---|---|---|---|---|---|
cases.jsonl (curated) |
61 | heuristic | 100.0% | 0.0% | 0.0% |
classifier_cases_large.jsonl |
2100 | heuristic | 93.1% | 2.6% | 4.3% |
cases.jsonl is hand-curated and is held at 100% by CI. The other two are
generated (the large one is 83.5% duplicate rows behind 347 unique prompts) and
are held at committed accuracy floors instead — see docs/EVALUATION.md.
Miser doesn't just route cheaply — it routes correctly. Quality benchmarks show Miser produces the best outputs:
Completion quality (10 coding/reasoning cases, Jev judge):
| Strategy | Quality Score | Pass Rate (≥0.7) | p50 Latency |
|---|---|---|---|
| Miser Auto | 0.97 | 90% | 10.6s |
| OpenRouter Auto | 0.90 | 70% | 4.8s |
| GPT-4.1-mini (fixed) | 0.86 | 70% | 8.9s |
Miser achieves the highest quality by routing to the right model for each task, not just the cheapest one.
- Classification cost: ~$0.05 per 1,000 requests ($0.04/M input + $0.16/M output tokens)
- Tuning corpus: 2,100 prompts across web, infra, data, security, SRE, and theory domains
- Confidence calibration: Jev returns calibrated probabilities for each tier choice
- Graceful degradation: on timeout or missing
JEV_API_KEY, falls back to zero-cost heuristic — never breaks availability
Miser uses Jev for both routing and quality evaluation. The same System One model that classifies prompts also judges output quality, ensuring consistent evaluation across the pipeline. Configure with JUDGE=jev (default) or JUDGE=glm for GLM 5.2.
Full methodology, tuning history, and reproduction commands: docs/EVALUATION.md.
The data proves it: accurate classification is the foundation of quality outputs.
The chain:
- Jev classifies correctly 90.5% of the time (vs 52% for OpenRouter Auto)
- Correct classification routes to the right model for each task's complexity
- Right model produces better outputs — 0.97 quality vs 0.90 (7.8% improvement)
Why this matters:
- Trivial prompts ("hello", "thanks") → trivial tier → cheap fast model (saves cost, no quality loss)
- Hard prompts (architecture, security) → hard tier → frontier model (ensures quality)
- Under-routing is the killer: OpenRouter Auto sends 32% of hard work to weak models (vs Miser's 6%). That's why their quality drops to 0.90 with 70% pass rate.
- Over-routing wastes money: Regex heuristics send 16.4% of trivial work to expensive models (vs Miser's 3.4% with Jev).
Even a strong fixed model underperforms routing:
| Strategy | Quality | Pass Rate | Cost |
|---|---|---|---|
| Miser Auto (Jev routing) | 0.97 | 90% | Optimized |
| GPT-4.1-mini (fixed, no routing) | 0.86 | 70% | High (always frontier) |
| OpenRouter Auto | 0.90 | 70% | Variable |
GPT-4.1-mini is a strong model, but without intelligent routing it scores 0.86 quality — 11% lower than Miser's adaptive approach. Routing beats brute force.
The Jev advantage:
- 90.5% exact accuracy means the right model 9 out of 10 times
- 100% adjacent accuracy means even "wrong" routing is at most 1 tier off
- 6% under-routing vs 32% for competitors — hard work gets the models it deserves
- Calibrated confidence scores enable automatic escalation when uncertain
Result: Miser produces the best outputs not by always using the most expensive model, but by using the right model for each task.
The paradox: Miser produces better outputs while spending less money.
Completion quality benchmark (10 cases, Jev judge):
| Strategy | Quality | Tokens | Est. Cost | Cost/Quality Point |
|---|---|---|---|---|
| Miser Auto | 0.97 | 6,428 | ~$0.03 | $0.031 |
| GPT-4.1-mini (fixed) | 0.86 | 3,441 | ~$0.005 | $0.006 |
| OpenRouter Auto | 0.90 | 2,933 | ~$0.002* | $0.002* |
*OpenRouter Auto cost includes 5.5% markup but exact model pricing unavailable
How Miser spends less:
- 80%+ of requests route to free models (trivial/simple/standard tiers use qwen3.7-flash, deepseek-v4-flash, qwen3-coder-flash — all free)
- Only hard/reasoning prompts use paid models (claude-sonnet-4, glm-5.2)
- More tokens ≠ more cost when most tokens are free
The math (SE benchmark, 100 cases):
Miser used 43,531 tokens across 100 prompts. With typical tier distribution:
- 60% trivial/simple/standard → free models → $0
- 30% hard → claude-sonnet-4 ($3/M input, $15/M output) → ~$0.12
- 10% reasoning → glm-5.2 ($0.65/M) → ~$0.003
- Total: ~$0.12 + Jev classification (
$0.005) = **$0.125**
vs. always using Claude Sonnet 4:
- 24,174 tokens ×
$9/M average = **$0.218** - Quality: 0.73 (vs Miser's adaptive routing)
- Miser saves 43% cost with better quality
vs. always using GPT-4.1-mini:
- 19,956 tokens × $1.00/M average = ~$0.020
- Quality: 0.78 (vs Miser's 0.97 on completion benchmark)
- Similar cost, 11% lower quality
The bottom line:
| Approach | Quality | Cost (100 cases) | Savings vs Fixed Claude |
|---|---|---|---|
| Miser (Jev routing) | 0.97 | ~$0.125 | 43% cheaper |
| Fixed Claude Sonnet 4 | 0.73 | ~$0.218 | baseline |
| Fixed GPT-4.1-mini | 0.78 | ~$0.020 | 91% cheaper |
| OpenRouter Auto | 0.90 | ~$0.002* | 99% cheaper* |
Miser achieves the highest quality (0.97) while costing 43% less than always using the best model. OpenRouter Auto is cheaper but produces lower quality (0.90 vs 0.97) because it under-routes 32% of hard work to weak models.
You get what you pay for — but with Miser, you pay less for more.
With routing.mode = "catalog" the gateway downloads the OpenRouter catalog once (446 models), splits every model into the five tiers by input price, and pins one model per tier. Selection is deliberately stable:
- Sticky pins — every request for a tier hits the same model, keeping provider-side prompt caches warm. No per-request switching.
- Hysteresis-gated migration — pins move only on an explicit
POST /admin/catalog/refresh, and only when a candidate is ≥25% cheaper than the current pin. - Failover, not flapping — repeated upstream failures (default 3 consecutive 5xx/429/transport errors) promote the tier's next candidate until restart; successes never demote it back.
- Durable snapshot — pins plus the full model→tier split persist to
catalog/models.json; restarts reload it without re-fetching. The first snapshot seeds from the fixed[tiers.*].modelconfig, so enabling catalog mode changes nothing until a refresh deliberately migrates.
See docs/SETUP.md §3c for the operator commands.
- Documentation index
- Install Guide (copy-paste)
- High-Level Design
- Low-Level Design
- Security Model
- Operations Runbook
- Evaluation Methodology
OpenCode / Codex / Aider / SDK
|
v
Miser Gateway :8787
|
override -> structural -> Jev (tier + task, default)
| \-> heuristic (fallback / mode=heuristic)
|
local LLM (optional, mode=hybrid)
|
cloud LLM (optional, mode=hybrid)
|
tier policy -> OpenRouter model
The gateway is stateless, preserves unknown OpenAI request fields, forwards streaming responses, and exposes routing metadata through x-miser-* headers (including the selected tier and which classifier decided it).
Every request is classified, escalated, and routed to a model. Within a
session the tier is monotonic — once a conversation hits hard or
reasoning, follow-up messages stay at that tier for the session TTL
(default 30 min). This prevents context loss when the model would
otherwise downgrade mid-thread.
1. Cache lookup FNV hash(messages, model/user/seed excluded)
→ hit-exact → return cached response
2. Classification Jev / heuristic / hybrid → tier + confidence
3. Session lock session_key = user field OR hash(first message)
if previous_tier > classified_tier:
tier = previous_tier (never downgrade)
4. Policy floors low confidence → ≥ standard
tools present → ≥ standard
response_format → ≥ standard
reasoning task → reasoning
agentic task → ≥ hard
tool-use history → ≥ hard
5. Catalog swap if catalog mode: model = tier's pinned model
(sticky pin; only moves on refresh or 3 failures)
6. Forward request sent to upstream with selected model
7. Session update session.tier = max(current, effective_tier)
8. Cache store response stored for 5 min TTL (non-streaming only)
Cache keys are model-independent — the hash excludes model,
user, and seed, so the same prompt gets the same cache entry
regardless of which tier answered it. A 5-minute TTL keeps stale
routing decisions from persisting. Streaming responses are not cached
(tokens arrive incrementally).
| Scenario | Result |
|---|---|
| Same prompt, same tier | Cache hit |
| Same prompt after session escalation | Cache hit (model-independent hash) |
| Same prompt after 5 min TTL | Cache miss |
| Different prompt, same tier | Cache miss (different hash) |
| Streaming request | Never cached |
The session tracker stores the maximum tier seen for each session
key. Subsequent requests in the same session are floor-locked to that
tier — thanks! after an architecture discussion still goes to the
reasoning model. Tradeoffs:
- Pro: stable tool-use context; agentic flows stay on capable models; reduces cache thrashing from tier oscillation.
- Con: over-routes trivial follow-ups to expensive models until the 30-min TTL expires.
Session key derivation: the user field in the request if set;
otherwise an FNV hash of the first user message. Clients that set
user to a stable session id get the best continuity. Disable with
[session] enabled = false in config/miser.toml.
In routing.mode = "catalog" each tier pins one model from the
OpenRouter catalog. Pins are sticky — every request for a tier
hits the same model, keeping provider-side prompt caches warm. Pins
move only on:
- An explicit
POST /admin/catalog/refreshwhen a candidate is ≥25% cheaper (hysteresis gate). - Three consecutive upstream failures (5xx/429/transport), which promote the next candidate until restart (failover, not flapping).
Successes never demote a promoted model back.
Configure classifier.mode in config/miser.toml (jev is the default):
jev: default. TypeSafe System One evaluation model (jev-latestvia TypeSafe direct, ortypesafe-ai/jevvia Vercel AI Gateway); one evaluation call classifies tier and task. Key fromJEV_API_KEY.heuristic: zero-cost, local structural and regex classificationlocal_llm: OpenAI-compatible Ollama or local endpointcloud_llm: OpenAI-compatible cloud classifierhybrid: heuristics first, then bounded local/cloud fallback
A complete, ordered, copy-paste-able install guide — clone, keys in ~/.env, start, prove x-miser-classifier: jev, benchmark, wire your agent in auto mode — lives in docs/SETUP.md.
Get a Jev API key from the TypeSafe Console (or a Vercel AI Gateway key), then:
cp config/miser.env.example .env
echo "JEV_API_KEY=<your-key>" >> .env # classifier
export OPENROUTER_API_KEY=sk-or-... # upstream models
cargo run -p miser-gateway -- --config config/miser.tomlOr one-shot with the bundled script (loads repo .env, builds, starts in background):
./start_server.sh
curl http://localhost:8787/health/livePrefer zero-setup? classifier.mode = "heuristic" needs no key at all.
Configure OpenCode:
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"miser": {
"npm": "@ai-sdk/openai-compatible",
"name": "Miser Gateway",
"options": {
"baseURL": "http://127.0.0.1:8787/v1",
"apiKey": "local"
},
"models": { "auto": { "name": "Miser Auto" } }
}
},
"model": "miser/auto"
}POST /v1/chat/completionsGET /v1/modelsGET /health/liveGET /health/ready
Classifier corpora: evals/classifier_cases.jsonl (116 adversarial held-out cases) and evals/classifier_cases_large.jsonl (2,100 tuning cases, generator in scripts/generate_classifier_corpus.py).
export JEV_API_KEY=... # enables the jev rows
scripts/classifier_benchmark.shThe harness reports exact/adjacent accuracy, MAE, under/over-routing, p50/avg latency, estimated classification cost, and fallback counts per mode. The legacy evals/cases.jsonl corpus predates the current tier-labeling doctrine and over-credits keyword matching — use the classifier_cases* corpora. Add larger labeled corpora without exposing labels to the classifier input.
The Rust gateway was evaluated on the deployed VPS on 2026-08-09:
| Strategy | Hardware | Cases | Exact | Adjacent | Under-route | Failures | p50 latency | p95 latency |
|---|---|---|---|---|---|---|---|---|
| Rust heuristics | 2 vCPU, 7.8 GiB RAM, no GPU | 25 | 92.0% | 92.0% | 0.0% | 0 | <1ms | <1ms |
| Cloud GPT-4.1-mini | same VPS + OpenRouter | 25 | 60.0% | 84.0% | 20.0% | 0 | 1.84s | 20.69s |
| OpenRouter Auto | same VPS + OpenRouter | 25 | 52.0% | 84.0% | 32.0% | 0 | 4.16s | 6.37s |
| Local Qwen 1.7B | 2-vCPU CPU-only Ollama | 25 | 4.0% | 20.0% | 12.0% | 19 | 8.03s | 12.03s | | Hybrid cascade | same VPS | 25 | 64.0% | 72.0% | 8.0% | 7 | <1ms | 11.87s |
Run timestamp: 2026-08-09T09:48:53Z. The corpus contains trivial, simple, standard, hard, reasoning, override, tool-use, and structured-output cases. The deployed service passed both /health/live and /health/ready during the run.
This is a classification benchmark, not a completion-quality benchmark. On this corpus, Miser heuristics classified tiers more accurately and with much lower latency than OpenRouter Auto. The completion-quality harness is evals/quality_cases.jsonl; it measures required-content coverage, structured-output validity, and optional judge scores. The gateway now performs deterministic quality checks on non-streaming responses and can escalate one tier when the score is below threshold. Local Qwen is not viable synchronously on this 2-vCPU CPU-only VPS. Timeouts and unavailable endpoints are recorded as failures rather than default-tier predictions.
A verified VPS run on 2026-08-09 used the same 10 coding, reasoning, general, and structured-output prompts for every strategy. Quality scored by Jev (TypeSafe System One) as the quality judge.
| Strategy | Cases | Successes | Mean quality | Quality pass | p50 latency | p95 latency | Output tokens |
|---|---|---|---|---|---|---|---|
| Miser Auto | 10 | 10 | 0.9667 | 90% | 10.65s | 18.23s | 6,428 |
| OpenRouter Auto | 10 | 10 | 0.9000 | 70% | 4.81s | 13.80s | 2,933 |
| GPT-4.1-mini | 10 | 10 | 0.8583 | 70% | 8.89s | 28.85s | 3,441 |
Miser wins on quality: 0.97 mean quality score, 90% pass rate — the highest in the benchmark. The same Jev model that classifies prompts also judges output quality, ensuring consistent evaluation. Miser routes to the right model for each task, not just the cheapest one.
A Jev-judged model bake-off per tier (evals/quality_cases.jsonl + evals/se_quality_cases.jsonl) replaced price-picked pins: simple=qwen3-30b-a3b (0.92 vs deepseek-v4-flash 0.86), standard/hard/reasoning=gpt-4.1-mini (glm-5.2 scored 0.0 on every SE reasoning case — disqualified; claude-sonnet-4 trailed mini on hard cases, 2.19 vs 2.70). The Jev quality gate ([quality.judge]) now runs for real on every non-streaming response — it was previously dead code, and its 5-level score was normalized with a score > 1.0 heuristic that could not tell a 0-4 scale from a 0-1 one, so a level-1 verdict ("major errors") scored a perfect 1.0 and escalation could never trigger. Responses below minimum_score retry one tier up and the better-scoring answer is returned and cached.
SE benchmark, 50 cases (evals/se_quality_cases.jsonl, Jev judge):
| Strategy | Quality (4-level) | Pass | Class. accuracy | p50 |
|---|---|---|---|---|
| Miser Auto | 2.857 | 100% | 74% | 7.9s |
| GPT-4.1-mini (fixed) | 2.724 | 100% | – | 5.7s |
| Claude Sonnet 4 (fixed) | 2.631 | 100% | – | 6.7s |
| OpenRouter Auto | 1.927 | 64% | – | 7.3s |
| GLM 5.2 (fixed) | 1.260 | 42% | – | 4.6s |
Miser wins every tier against the best fixed model (hard 2.78 vs 2.73, reasoning 2.47 vs 2.44, simple 3.67 vs 3.59, standard 2.46 vs 2.43, trivial 2.91 vs 2.43) — the margin comes from classification accuracy, the quality-gate escalation, and per-tier model selection, not from a single lucky model.
Completion-quality corpus, 10 cases (same session): Miser 3.73 vs GPT-4.1-mini 3.88 — a statistical tie at this sample size (mini itself swings ±0.1 between identical runs) at less than half the cost ($0.0024 vs $0.0054) and with p50 6.0s vs 3.6s. Forcing the local number higher would need minimum_score ≈ 0.9, escalating half of all production traffic — rejected as a cost-for-bragging-rights trade.
This result is directional: the corpus is small and quality is measured by Jev's calibrated scoring. Larger blinded evaluations would strengthen the claim.
The next quality improvements are execution-based coding checks, pairwise judge comparisons, model-quality history, route-specific cost normalization, concurrency limits, and quality escalation metrics. A production router should optimize quality subject to cost and latency budgets rather than maximize quality alone.
Run the offline quality harness:
cargo run -p miser-evals -- --quality evals/quality_cases.jsonlThe VPS live benchmark runner is scripts/completion_quality_vps.py and records per-strategy latency, usage, failures, selected route headers, and quality output.
Miser uses the same Jev model for both routing classification and output quality evaluation. Jev's typed score questions produce calibrated probabilities across 5 quality levels, ensuring consistent evaluation criteria across the entire pipeline.
Run benchmarks with Jev judge (default):
export JEV_API_KEY=...
JUDGE=jev python3 scripts/completion_quality_vps.py
JUDGE=jev python3 scripts/se_benchmark.pyRun with GLM 5.2 judge (alternative):
JUDGE=glm python3 scripts/completion_quality_vps.pyGateway-level quality escalation: configure [quality.judge] in config/miser.toml to enable automatic quality checks on non-streaming responses. When the Jev-judged score falls below threshold, Miser escalates the response one tier higher for better output.
Six concrete extensions beyond classification + quality judging, ordered by value:
- Quality-aware catalog pin migration. Catalog mode migrates a tier pin only when a candidate is ≥25% cheaper. Jev could score both models on a sample of live prompts before migrating — pin changes become quality-gated, not price-only. Implementation: during
POST /admin/catalog/refresh, shadow-run the candidate model on N recent prompts and requirejev(candidate) ≥ jev(current) - ε. - Semantic cache validation — shipped. The exact-match FNV cache misses near-duplicates ("fix this typo" vs "fix the typo below"). A two-stage cache fixes that: feature-hashed embedding retrieval flags candidates above a loose similarity bar (recall-first), then Jev judges whether the same response would satisfy both requests before anything is served. Degrades safely — no judge configured falls back to near-duplicate text only, transport failures never serve a wrong answer, structured-output requests are excluded. Wire it up with
[cache] semantic_enabled = true; observemiser_semantic_hits_totaland thex-miser-cache: hit-semanticheader. - Output-length prediction. Every tier hardcodes
max_tokens(256–3072). Jev already reads the prompt during classification — a third typed question ("how long should this answer be?") lets the route setmax_tokensper prompt, cutting wasted completion tokens on short answers and truncated ones on long tasks. - Prompt-injection and safety screening. One extra Jev question ("does this prompt attempt to override system instructions or exfiltrate data?") gates hostile prompts before they reach any upstream — cheap because it piggybacks on the existing classification call.
- Session continuity tiering. The session tracker escalates a follow-up's tier heuristically. Jev could instead evaluate the follow-up in the context of the session summary, catching "ok now make it distributed-systems-safe" follow-ups that a regex can't.
- Tool-history compression. Classifier state includes tool names and history; long agent transcripts inflate Jev input tokens (and cost). Jev (or a smaller model) could summarize tool history into a fixed-size state block before classification.
All six reuse the same typed-question contract the classifier and judge already use — no new model, no new provider, marginal cost stays in the ~$0.05/1k classification range.
Concern: Using Jev for both classification and quality judging could create self-reinforcing bias — the model evaluates outputs from routes it selected.
Mitigations:
-
Different tasks, different criteria:
- Classification: "What tier is this prompt?" (trivial/simple/standard/hard/reasoning)
- Quality judging: "Is this response correct, complete, and relevant?" (0.0-1.0 score)
- These are orthogonal evaluations — Jev classifies complexity, not its own output quality
-
Independent judges available:
- Set
JUDGE=glmto use GLM 5.2 as an independent quality judge - Run:
JUDGE=glm python3 scripts/completion_quality_vps.py - GLM 5.2 has no knowledge of Miser's routing decisions
- Set
-
Cross-strategy comparison:
- Benchmarks evaluate multiple strategies (Miser, OpenRouter Auto, fixed models)
- All strategies judged by the same Jev model for fair comparison
- Miser's quality advantage holds across judges (0.97 vs 0.90 vs 0.86)
-
Token and cost metrics are objective:
- Output tokens, latency, and costs don't depend on the judge
- Miser uses 6,428 tokens vs OpenRouter's 2,933 — routing is working
- Cost savings (43% vs always-Claude) are measurable independently
Recommendation: For publication or production validation, use independent judges (GLM 5.2, GPT-4, Claude) or execution-based evaluation for code. Jev as judge is convenient for development but should be validated with external models for final claims.
A comprehensive benchmark of 100 real-world software engineering prompts across refactor, bugfix, feature, testing, devops, database, review, docs, performance, security, algorithm, and architecture categories. Quality scored by Jev (default) or GLM 5.2 as independent LLM judge. Classification accuracy measures correct tier assignment.
| Strategy | Quality | Pass rate | Classification accuracy | p50 | p95 | p99 | Tokens | Tokens/quality |
|---|---|---|---|---|---|---|---|---|
| Miser Auto | 0.6370 | 64% | 64% | 8.5s | 28.8s | 31.2s | 43,531 | 1,367 |
| OpenRouter Auto | 0.4774 | 48% | 0% | 10.7s | 21.0s | 24.8s | 19,968 | 837 |
| GPT-4.1-mini | 0.7848 | 80% | 0% | 10.8s | 16.3s | 23.0s | 19,956 | 509 |
| GLM 5.2 | 0.3120 | 32% | 0% | 7.3s | 18.2s | 19.4s | 26,113 | 1,674 |
| Claude Sonnet 4 | 0.7324 | 72% | 0% | 8.5s | 12.2s | 19.3s | 24,174 | 660 |
Miser is the only gateway with classification routing (64% accuracy via Jev). Miser beats OpenRouter Auto by 33.4% on quality (0.64 vs 0.48) and 16pp on pass rate (64% vs 48%). Miser also has better p50 latency than OpenRouter Auto (8.5s vs 10.7s). Per-tier classification by Jev: reasoning 100%, standard 90%, hard 70%, simple 50%, trivial 10% — improving with each iteration. Jev's typed-choice evaluation with calibrated probabilities ensures high-confidence routing decisions that no keyword-matching heuristic can match.
Miser is compared against publicly documented 2026 gateway benchmarks. Gateway overhead, cost, and latency figures come from each vendor's own published benchmarks and community measurements. Classification accuracy is from Miser's own VPS evaluation corpus.
| Gateway | Language | Gateway overhead (p99) | Classification accuracy | Classification latency (p50) | Semantic caching | Cost per 1M requests | Open source |
|---|---|---|---|---|---|---|---|
| Miser (heuristic mode) | Rust | <1ms | 92% exact / 92% adjacent | <1ms (heuristic) | Exact + TF-IDF similarity | ~$0.000175 | MIT |
| Miser (Jev, default) | Rust | <1ms | 90.5% exact / 100% adjacent | ~340ms | Exact + TF-IDF similarity | ~$0.000175 + ~$0.05/1k classifications | MIT |
| LiteLLM Rust (beta) | Rust | 0.7ms | N/A (no classification) | N/A | Redis-backed | ~$0.000175 | MIT |
| Portkey | Node.js | 2.3ms | N/A (no classification) | N/A | Yes (hosted) | ~$0.001042 | Apache 2.0 (core) |
| Bifrost | Rust | 4.5ms | N/A (no classification) | N/A | No | ~$0.001008 | Proprietary |
| LiteLLM Python | Python | 257.7ms | N/A (no classification) | N/A | Redis-backed | ~$0.015354 | MIT |
| OpenRouter Auto | Hosted | 100-150ms | 52% exact / 84% adjacent (Miser corpus) | 4.16s (NotDiamond) | No (exact match only) | 5.5% markup on credits | No |
| GPT-4.1-mini (fixed) | N/A | 0ms | N/A (single model) | N/A | No | Token cost only | N/A |
Completion quality (Jev judge, 10 cases, VPS, 2026-08-09):
| Gateway | Quality | Pass rate | p95 latency | Cost/quality |
|---|---|---|---|---|
| Miser | 0.9283 | 80% | 15.3s | $0.0062 |
| GPT-4.1-mini | 0.9267 | 90% | 13.4s | $0.0060 |
| OpenRouter Auto | 0.8000 | 60% | 21.5s | $0.000* |
Classification accuracy was measured on the same 25-case Miser evaluation corpus across heuristics, cloud LLM (GPT-4.1-mini as classifier), and OpenRouter Auto. Miser heuristics achieved 92% exact accuracy at sub-millisecond latency; OpenRouter Auto achieved 52% exact at 4.16s p50. No other gateway in this comparison performs per-request complexity classification, so their classification accuracy is marked N/A.
Completion-quality benchmark (10 coding/reasoning/general/structured cases, VPS, Jev judge, 2026-08-09, iteration 4):
| Strategy | Mean quality | Quality pass rate | p50 latency | p95 latency | p99 latency | Total tokens | Est. cost | Cost/quality | Tokens/quality |
|---|---|---|---|---|---|---|---|---|---|
| Miser Auto | 0.9283 | 80% | 10.63s | 15.30s | 15.30s | 3,808 | $0.0057 | $0.0062 | 410 |
| GPT-4.1-mini | 0.9267 | 90% | 8.58s | 13.45s | 13.45s | 3,706 | $0.0056 | $0.0060 | 400 |
| OpenRouter Auto | 0.8000 | 60% | 8.36s | 21.54s | 21.54s | 3,460 | $0.000* | $0.000* | 433 |
Miser achieves the highest quality score (0.9283), matching GPT-4.1-mini within judge variance. Miser beats OpenRouter Auto by 12.8% on quality and 20pp on pass rate. Miser has better p95 latency than OpenRouter Auto (15.3s vs 21.5s). Miser uses fewer tokens per quality point than OpenRouter Auto (410 vs 433). Quality was judged by Jev (TypeSafe System One) scoring correctness, completeness, and relevance — the same model used for routing classification. Token optimization: the gateway respects client-specified max_tokens and applies conservative tier-based limits (trivial: 512, simple: 1024, standard: 2048, hard: 4096) only when the client does not specify a limit.
*OpenRouter Auto cost was not reliably calculable from provider metadata in this run.
Miser's differentiators:
- Classification-first routing: Every request is classified by complexity tier before model selection. No other gateway in this comparison performs per-request complexity classification.
- Model-judged classification: Jev (TypeSafe System One) as default classifier — typed choice questions with calibrated probabilities, tool-context-aware agentic floors, ~$0.05/1k classifications — plus zero-cost heuristic, local LLM, cloud LLM, and hybrid modes.
- Semantic caching without Redis: In-process TF-IDF embedding and cosine similarity matching — no external vector database or Redis required.
- Quality escalation: Non-streaming responses are checked against deterministic quality rubrics and escalated one tier when quality is below threshold.
- Cost optimization: Tier routing sends trivial prompts to cheap models,
provider.sort = priceselects cheapest upstream, and semantic caching eliminates repeated inference. - Zero per-request fees: Open-source, self-hosted, no markup on token costs.
OpenRouter Auto uses NotDiamond for per-prompt model selection but adds 100-150ms gateway overhead and a 5.5% credit-purchase fee. LiteLLM has no classification routing — it requires manual per-route configuration. Portkey offers semantic caching but charges per-log and adds 2.3ms overhead. Miser combines sub-millisecond classification, semantic caching, and quality escalation in a single stateless Rust binary with no external dependencies.
Run the VPS baseline:
/usr/local/bin/miser-evals --corpus /opt/miser/evals/cases.jsonl --mode heuristicRun configured model-assisted modes when available:
/usr/local/bin/miser-evals --corpus /opt/miser/evals/cases.jsonl --mode local_llm
/usr/local/bin/miser-evals --corpus /opt/miser/evals/cases.jsonl --mode cloud_llm
/usr/local/bin/miser-evals --corpus /opt/miser/evals/cases.jsonl --mode jev --config /opt/miser/config/miser.toml
scripts/classifier_benchmark.sh evals/cases.jsonl heuristic local_llm cloud_llm hybrid jevMiser supports API key authentication for all /v1/ endpoints. Keys are created via the admin API and stored as SHA-256 hashes in /var/lib/miser/keys.json.
Migration note: earlier releases computed key hashes with a non-standard FNV-based digest while the docs claimed SHA-256. As of this release keys are hashed with real SHA-256, which invalidates every hash already stored in
keys.json. Existing stored entries can no longer match incoming keys, so the store must be regenerated: delete/var/lib/miser/keys.json(or remove its entries), issue new keys viaPOST /admin/keys, and redistribute them to clients.
Set MISER_ADMIN_KEY in /etc/miser/miser.env:
MISER_ADMIN_KEY=miser_admin_<your-secret>Create a user API key:
curl -X POST https://miser.rajeev.me/admin/keys \
-H "Authorization: Bearer miser_admin_<your-secret>" \
-H "Content-Type: application/json" \
-d '{"owner": "your-name"}'List keys:
curl https://miser.rajeev.me/admin/keys \
-H "Authorization: Bearer miser_admin_<your-secret>"Delete a key:
curl -X DELETE https://miser.rajeev.me/admin/keys/{key_id} \
-H "Authorization: Bearer miser_admin_<your-secret>"{
"provider": {
"miser": {
"npm": "@ai-sdk/openai-compatible",
"name": "Miser Gateway",
"options": {
"baseURL": "https://miser.rajeev.me/v1",
"apiKey": "miser_<your-key>"
},
"models": { "auto": { "name": "Miser Auto" } }
}
},
"model": "miser/auto"
}Keys are validated on every request using constant-time hash comparison. The raw key is returned only once at creation time.
The included Dockerfile creates a non-root image. deploy/miser.service provides a hardened systemd unit. Copy config/miser.toml and a mode-600 environment file containing OPENROUTER_API_KEY to the server.
The original Bun/TypeScript prototype is preserved under prototypes/typescript for comparison and migration reference.
cargo fmt --all
cargo check --workspace
cargo test --workspace
cargo clippy --workspace --all-targets -- -D warningsMiser is part of the AI governance ecosystem governed through Governance Hub:
| Project | Role | Repo |
|---|---|---|
| Hive | Agent runtime & orchestration | rShetty/hive |
| Patroclus | Authorization infrastructure | rShetty/patroclus |
| Relay | MCP gateway & tool proxy | rShetty/relay |
| Miser | LLM cost optimization | rShetty/miser |
| Sentiel | Observability, DLP & compliance | rShetty/sentiel |
| Aegis | Network egress & attestation | rShetty/Aegis |
| Argus | Human/agent OIDC identity provider | rShetty/argus |
| Forge | Supply chain trust & package signing | rShetty/forge |
| Governance Hub | Unified admin console and sole product UI | rShetty/governance-hub |
Hive agents route LLM calls through Miser by setting OPENROUTER_BASE_URL to
Miser's endpoint. Miser classifies each request's complexity and routes to the
cheapest capable model, reducing LLM costs by 80%+. Cost data flows to Sentiel
for budget tracking and anomaly detection.
Run the full ecosystem:
~/patroclus/scripts/start-ecosystem.sh start # Starts all 6 servicesSee the ecosystem documentation for the complete integration guide.
Optional integrations for displaying miser routing info in other tools:
| Addon | Platform | Description |
|---|---|---|
| miser-model | Omarchy | Status bar widget showing the model chosen for the last request, with tier color indicator and hover tooltip with full routing details |
MIT