Skip to main content
Model Routing

Not every question needs your most expensive model.

You have one API key. Right now, every query hits the same model at the same price. Routing classifies each question and sends it to the right model. Within a single provider, or across providers.

Simple questions get the fast, cheap model. Complex ones get the heavyweight. Same key. Same provider. Less waste.

Without Routing
100 API calls/day, all to Claude Sonnet
What's the capital of France?
Sonnet · $3/Mtokoverkill
Review this auth middleware
Sonnet · $3/Mtok
Analyze HIPAA implications of...
Sonnet · $3/Mtokunderkill
~$47/mo provider spend
With Routing
100 API calls/day, intelligently routed
What's the capital of France?
Haiku 4.5 · $1/Mtok3x cheaper
Review this auth middleware
Sonnet 4.6 · $3/Mtok
Analyze HIPAA implications of...
Opus 4.8 · $5/Mtok
~$22/mo provider spend
Same quality. Right model for each question.

Routing works through the ASURIQ API or MCP server. You send queries to us with your provider key. We classify, route, and return the response.

Add Routing · $13/moTypically saves $20+/mo · works with any provider · cancel anytime
8 behavioral dimensions · multi-model panel accuracy · independent error detection · learns from every query
Start here

One key. Three models. Immediate savings.

Start with the key you already have. Routing classifies every query by complexity and routes within your provider's model family. Simple questions go to the fast model. Complex ones get the heavyweight. Your API key stays the same.

Claude
Haiku 4.5 · $1/Mtok
Sonnet 4.6 · $3/Mtok
Opus 4.8 · $5/Mtok
One Anthropic key
OpenAI
GPT-5.6 Luna · $1/Mtok
GPT-5.6 Terra · $2.50/Mtok
GPT-5.5 · $5/Mtok
One OpenAI key
Gemini
Flash 2.5 · $0.15/Mtok
Pro 3.1 · $2/Mtok
Ultra · $7/Mtok
One Google key

The cheapest model in each family is 3x to 40x less expensive than the most capable. Routing knows which one to use for each question. You save money without losing quality.

The crown jewel

One model has opinions. Multiple models have knowledge.

Single-provider routing saves money. Multi-model routing saves money AND dramatically improves accuracy. Here's why: models trained on different data fail on different questions. When Claude and DeepSeek disagree on an answer, that disagreement is the most valuable signal in the entire system.

Why more models means more accuracy

In 1785, the Marquis de Condorcet proved a mathematical law: if each voter in a jury is right more than 50% of the time, and their errors are independent, the group accuracy converges toward 100% as you add jurors.

LLMs are jurors. Claude, GPT, Gemini, DeepSeek, Kimi: each is right more than 50% of the time. And because they're trained on different data, their errors are genuinely independent. The conditions for the theorem are met.

One model at 85% accuracy. Three independent models at 85% accuracy: 96.6% panel accuracy. Five models: 99.3%.

Why error independence matters

Claude was trained by Anthropic. GPT by OpenAI. DeepSeek by a Chinese lab. Gemini by Google. They consumed different training corpora, applied different RLHF, and developed different failure modes.

When Claude confidently says X and DeepSeek confidently says not-X, something interesting is happening. One of them is wrong, and the disagreement itself tells you which questions need more scrutiny.

If they all agreed on everything, you wouldn't need more than one. The value is in the disagreement surface.

The panel recommendation
1Your query arrives at Routing
2Routing classifies: domain, complexity, stakes
3Low stakes? Single best model. Fast, cheap, done.
4High stakes? Panel of 3-5 independent models
5Models answer independently. No consensus pressure.
6Routing compares answers. Flags disagreements.
7Returns the best answer with a confidence assessment
What multi-model panels catch that single models miss
Hallucination detection
When 4 of 5 models agree and one is an outlier, the outlier is hallucinating. Single-model usage gives you no way to detect this.
Domain blind spots
Every model has topics where its training data was thin. Claude is weaker on some niche medical subfields. GPT struggles with certain legal frameworks. The panel covers each other's gaps.
Sycophancy resistance
Some models cave when pushed. They agree with you even when you're wrong. A panel of independent models can't all be sycophantic simultaneously. Dissent from any model is a signal worth examining.
Temporal accuracy
Models have different knowledge cutoffs and different rates of knowledge decay. A question about recent events might get the wrong answer from a model with a stale training set. The panel cross-checks.
Reasoning diversity
Chain-of-thought, few-shot, direct reasoning: different models approach the same problem differently. When they reach the same answer via different reasoning paths, your confidence should be high.
Cost optimization
Not every question needs a panel. Routing only convenes a multi-model panel when the stakes justify the cost. Simple questions still route to the cheapest model. You pay for the panel only when it matters.
The economics work

"But doesn't calling 3 models cost 3x more?" Not with Routing.

60% of your queries are simple. They route to one cheap model. $1/Mtok. 30% are moderate. One model, mid-tier. $3/Mtok. Only 10% are high-stakes enough to justify a multi-model panel. And even then, the panel uses the cheapest effective combination.

Blended cost is 30-40% below running everything on Sonnet, with higher accuracy on the questions that actually matter.

Pair with Cognitive Stack for the Intelligence Layer bundle: $19/mo instead of $22 separate.

The science

Your models have personalities. We profile them.

Static benchmarks are snapshots from months ago. Routing builds live behavioral profiles from your actual queries across 8 dimensions. These profiles drive both single-provider routing and multi-model panel selection.

Accuracy
Running performance per domain. Not a benchmark score from 6 months ago. Live accuracy from your actual queries.
Calibration
When the model says 90% confident, is it right 90% of the time? Tracks predicted vs actual correctness.
Blind spots
Specific subdomains where performance drops vs the panel average. Every model has them. Routing knows where.
Complementarity
How independent is this model's error pattern from others? Cross-provider models fail differently. That's the signal.
Cost efficiency
Accuracy per dollar, by task type. The metric that actually matters for routing decisions.
Dissent quality
When this model disagrees with the panel, how often is it the one that's right? High dissent quality means independent thinker.
Sycophancy
Does the model cave when challenged, or hold its ground? A model that always agrees with the crowd adds no signal.
Decay rate
How quickly does its knowledge go stale on time-sensitive topics? Some models age faster than others.

Day 1: routes based on published benchmarks and cost.
Day 30: routes based on how each model actually performs on your questions.
Day 90: knows your domain well enough to predict which model, or which panel, will nail each query.

Go deeper

Every key you add deepens the savings and sharpens the accuracy.

Start single-provider. Add more keys as you see the value. Each new provider gives Routing more models to choose from and more independent perspectives for panel decisions. Keys are stored server-side in an encrypted vault — add them once in settings and every connector uses them immediately, with no re-entry per session.

DeepSeek
Simple classification routes to DeepSeek V4 Flash at $0.14/Mtok. Genuinely different training data gives your panels real independence.
21x cheaper than Sonnet for simple queries
Gemini
Large-context questions route to Gemini Flash at $0.15/Mtok. Google's training corpus covers areas where Anthropic and OpenAI are thinner.
20x cheaper for long-context work
Kimi K2
Code-heavy tasks route to Kimi at $0.55/Mtok, which benchmarks above Sonnet on SWE-bench. Chinese lab training data provides genuinely independent error patterns.
5x cheaper, stronger at code
Integration

Switch in one line. Switch back in one line.

Before (direct provider call)
POST https://api.anthropic.com/v1/messages
{ "model": "claude-sonnet-4-6", "messages": [...] }
Single provider (routed)
POST https://consensus-production-6eeb.up.railway.app/api/v1/routing/route
{ "query": "Your question", "api_key": "sk-..." }
Multi-provider panel
POST https://consensus-production-6eeb.up.railway.app/api/v1/routing/route
{ "query": "High-stakes question", "mode": "panel",
  "api_keys": { "anthropic": "sk-...", "openai": "sk-...",
                "deepseek": "sk-..." } }
Example: this month's routing stats
Queries routed3,247
Routed to Haiku (simple)1,948 (60%)
Routed to Sonnet (moderate)1,104 (34%)
Routed to panel (high-stakes)195 (6%)
Cross-provider panels47 (1.4%)
Saved vs all-Sonnet$28.40
More than cost savings

Routing is ASURIQ's quality guarantee.

Every ASURIQ tool has a minimum model requirement. WHETSTONE needs a Sonnet-class model to generate meaningful adversarial challenges. Confidence scoring needs mid-tier reasoning to produce calibrated axes. Without Routing, tools that need a stronger model simply won't fire. You'll see an explanation and an offer to add Routing.

With Routing, every tool call silently upgrades to the cheapest sufficient model. Your conversation stays on Haiku. Your WHETSTONE challenge fires on Sonnet. Invisible quality guarantee, zero wasted spend.

BYOK · $13/mo

Bring your own API keys. ASURIQ stores them server-side in an encrypted key vault, so every connector — extension, MCP, ChatGPT, Gemini, API — routes through the same keys without you re-entering them per session. You pay your providers directly; ASURIQ handles the intelligence: classification, model selection, behavioral fingerprinting, quality guarantees.

Your keys: Anthropic, OpenAI, DeepSeek, Gemini, Kimi
Full routing intelligence
Cost tracking and savings dashboard
Quality gate for all ASURIQ tools
No per-query charges from ASURIQ
Managed · from 3 credits

No API keys needed. ASURIQ provides the models. You pay credits per routed query. Same routing intelligence, we handle the infrastructure.

No API keys to manage
Same routing intelligence as BYOK
Single route: 3 credits
Panel route (3 models): 10 credits
Works from API, extension, or any connected surface
What this replaces
Claude Pro subscription$20/mo
ChatGPT Plus subscription$20/mo
Gemini Advanced$20/mo
$40–60/mo → one $13/mo subscription routes to the right model
See the verification engine
ASURIQ verification — click to watch
Your data stays yours
Prompts stay with your provider. We see analysis metadata only.
Keys never stored
One-way hash for authentication. Your credentials pass through. Never persist.
~2 second responses
Single-model cognitive tools return in about 2 seconds.
37 of 100 flagged
We ran 100 ChatGPT answers through verification. See the study →

Currently in friends-and-family beta. Built on peer-reviewed cognitive architecture.

Baars — Global Workspace TheoryACT-R — Memory Decay ModelWang et al. 2025 — Silent AgreementLi et al. EMNLP 2024 — Sparse Debate

Stop paying premium prices for every question.

$13/mo · Typically saves $20+/mo · Works with any provider · Cancel anytime