Respan Router BetaYour agent shouldn’t need the same model for every task.
Try Respan Router
Use sample data only.
Result
No response yet.

A support agent can look up an order, interpret a refund policy, and work through a disputed charge in the same conversation. Those requests do not all ask the same thing of a model. Sending everything to your strongest model is simple. Paying for that decision on every turn is expensive.
Respan Router puts model selection inside the gateway. Powered by Span-01, it combines task context, model capabilities, and cache-aware cost estimates behind one model ID: span-router.
You pay the serving model’s list price, including applicable cache discounts. No markup.
Route the task, not just the prompt
“Fix this” tells you very little. “Fix this failing migration in a production database; preserve existing records and use the available database tools” tells you considerably more.
Your application already knows much of that context. Give it to the router through the system or developer message: what the agent does, the environment it operates in, and the constraints it must respect. You do not need to make every end user explain your application again.
When classifying a task, Respan Router uses that context alongside the latest exchange. It also checks whether a model can handle the request’s context length, images, tools, and structured-output requirements.

Keep continuity. Reconsider when it matters.
A conversation is not a series of unrelated prompts. Keeping the same model can preserve a warm cache and avoid unnecessary switching.
Pass a consistent thread_identifier across turns. Respan Router prefers to keep the current model while it remains usable, checking capabilities and context limits as the conversation evolves. When the cost-aware switching policy is enabled, it can also reconsider the route using cost estimates.
This is not a promise to switch models on every turn. The useful decision is whether switching helps. A new turn alone is not a reason to switch.

Cache belongs in the routing decision
The cheaper model is not always the cheaper next call. A model with a warm cache can cost less than moving a long conversation to a lower-priced model and processing the input from scratch.
Respan Router’s cost-aware policy considers cached reads, applicable cache writes, fresh input, and expected output. The model’s list price is one input to that decision, not the whole decision.
To make caching effective, keep system instructions and tool definitions stable, and send the full conversation history. Cache reuse depends on the provider and matching prefixes; it is not guaranteed.
What the benchmarks show
On 557 held-out requests, our routing evaluation achieved 94.1% pass rate at 37% lower cost than the most accurate fixed model, which scored 94.3%.

View exact request results
| Approach | Pass rate | Total cost |
|---|---|---|
| Respan Router | 94.1% | $0.46 |
| Most accurate fixed model | 94.3% | $0.73 |
| Cheapest fixed model | 91.0% | $0.09 |
| Hindsight oracle (upper bound) | 97.1% | $0.18 |
The held-out set spans coding, math, knowledge, and function calling, drawn from HumanEval, MBPP, GSM8K, MATH-500, MMLU-Pro, BFCL, LiveCodeBench v6, and an internal set. These 557 items were not used to tune routing. Each item was evaluated on every candidate model.
A separate live run through the router scored 93.0%. The cost figures above belong to the routing evaluation, not that separate live run.
Agent tasks, head to head
We also tested complete agent tasks under matched settings for each router. These go beyond isolated requests.

View exact results
| Benchmark | Respan Router | Jev Router | OpenRouter Auto |
|---|---|---|---|
| τ-bench airline | 88.7% | 83.3% | 46.0% |
| τ-bench retail | 93.3% | 88.6% | 62.3% |
| τ³-bench banking | 31.3% | 30.6% | 4.1% |
Respan Router also completed 82.0% of airline tasks in all three trials, versus 72.0% for Jev Router. On retail, that consistency measure was 86.8% versus 78.9%. Average airline task latency was 28.6 seconds versus 32.9 seconds.
Respan Router and Jev Router scores are averaged over three trials; OpenRouter Auto was run once. The task sets contain 50 airline, 114 retail, and 97 banking tasks. The τ benchmarks use the same user simulator, GPT-4.1 at temperature 0. Results were measured September 26–27, 2026.
The strongest gains here are in airline and retail workflows. Banking is close.
One model name. Your existing application.
Respan Router works through the Respan Chat Completions endpoint. Change the model name, include the context your agent already has, and keep a stable thread identifier.
request.model = "span-router"
request.thread_identifier = conversation_id
request.messages = [system_context, ...conversation]
send_via_respan_gateway(request)Streaming, tool calls, and structured outputs are supported. On subsequent turns, send the full message history with the same thread_identifier. Do not combine span-router with fallback or load-balancing options.
Use span-router directly through the Respan gateway. Router calls use Respan credits; bring-your-own provider keys are not supported.
One model name.
More possibilities.
Create a Respan account and start routing with span-router through the gateway.