A router that picks the model for you is only as useful as your view into what it picked. Six alternatives compared on pricing and on whether the routed answer gets scored.
Dylan Cable · October 6, 2026 · 17 min read

OrcaRouter launched in May 2026 as an OpenAI-compatible router that grades each prompt and picks a model to match, with an open-source edition you can run on your own servers. Most of the attention it has picked up comes from two things: adaptive routing that learns from your traffic, and a pricing model with no token markup.
Handing model choice to a router changes where quality problems come from. When the router swaps a cheaper model into a request, the answer your user sees depends on that decision, so you need a way to tell whether it held up once real traffic arrives.
OrcaRouter handles routing and adds request logs, a session timeline, and evals you run against your own test cases. For teams running agents in production, the alternatives worth a look are the ones that trace every step and score live traffic in the same place they route it.
Here are the top 6 OrcaRouter alternatives and competitors in 2026.
All seven tools route requests across providers with fallbacks, so what separates them is how much each one records about a request and whether it scores output on live traffic.
| Tool | Routing | Tracing | Evals on live traffic | Pricing model |
|---|---|---|---|---|
| Respan | Fallbacks, task-aware router (beta) | Nested spans per agent step | Yes, sampled production spans | Free tier, Team $199/mo |
| OrcaRouter | Adaptive routing, YAML/CEL rules | Session timeline of calls | Offline, hand-written cases | Free, no token markup |
| OpenRouter | Fallbacks, Auto Router | Exported via Broadcast | Via exported traces | 5.5% fee on credits |
| Vercel AI Gateway | Provider and model fallbacks | Request logs | Separate tool needed | No markup, metered add-ons |
| LiteLLM | Load balancing, fallbacks | Callbacks to Langfuse, OTel | Separate tool needed | Free OSS, you host |
| Bifrost | Fallbacks, weighted load balancing | OpenTelemetry, Prometheus | Via Maxim plugin | Free OSS, Enterprise custom |
| Portkey | Fallbacks, load balancing | Built-in logs and traces | Guardrail checks as feedback | Free tier, Production $49/mo |
If scoring production traffic is a requirement, Respan runs it inside the router. Portkey's guardrail checks cover part of that, and the other tools get there through an export or a separate product.
OrcaRouter is an OpenAI-compatible LLM router from Continuum AI that sends each request to one of 200+ models across 12+ providers. Its adaptive routing grades every prompt and picks a model that meets your quality and cost thresholds, either automatically through orcarouter/auto or through routing rules you write in YAML and CEL.
Around the router, OrcaRouter adds model fallbacks, session affinity, a prompt registry, guardrails, and an agent firewall that grades tool and MCP calls before they run. Request logs show the model, latency, cost, and each retry for every call, while Sessions lays out an agent run as a timeline of model calls with the tools each one requested. Evals score a model or a router against test cases you write yourself.
OrcaRouter Lite is the MIT-licensed edition, and it runs on your own infrastructure with SQLite by default. On compliance, OrcaRouter lists frameworks like HIPAA and SOC 2 as policy packs it enforces, and its trust center notes that the certifications it actually holds are shared under NDA.
OrcaRouter's pricing is built around zero token markup: you pay each provider's published rate, and OrcaRouter charges for team and governance features instead. The free Hacker plan includes all 200+ models, auto-failover, a basic dashboard, prompt versioning, and 10 API keys.
Team and Enterprise are both custom-priced. Team adds up to 10 seats, compliance enforcement and reports, and unlimited API keys, while Enterprise moves to unlimited seats, dedicated infrastructure, and a 99.99% uptime SLA.
The zero-markup promise has some edges worth reading before you commit. Enterprise plans may include a flat per-million-token routing fee depending on volume, and BYOK traffic may carry a platform fee. Eval runs are also real requests billed as such, so a large test suite costs tokens every time you run it.
Every router in this space handles the basics: an OpenAI-compatible endpoint, a large model catalog, and fallbacks when a provider errors. Those are worth checking, but they rarely decide which router a production team keeps, because they look nearly identical from one vendor to the next.
What really separates them is what happens after the route. An LLM gateway sits in the path of every request, so it already sees the prompt, the model that answered, and the cost of each call. The question is how much of that it turns into something you can act on:
Routing logic still matters, especially for teams comparing LLM routers on cost, since a smarter route only saves money if quality holds after the switch. OrcaRouter covers the routing side and adds a session timeline and offline evals, so the gap to check for in any alternative is on live traffic.

Respan is the AI router with built-in observability and automated evals. Every model call goes through one API to 1,000+ models, and every request it routes is already traced, priced, and ready to score, so there's no second or third tool to wire on after the gateway. Respan Router (beta) picks the model per task from the system prompt and the latest exchange, and it weighs cache cost before switching so a cheaper model doesn't end up costing more.
For teams leaving OrcaRouter, the difference shows up after the route. Instead of reading logs after the fact, use Respan to run observability in production, know when production shifts, and act before it spreads.
Features:
Pros:
Cons:
Pricing: Free tier with 100k logs and unlimited seats; Team is $199/month billed yearly with 5 members included, 30-day retention, and a 99.9% uptime SLA; Enterprise is custom.
Best for: Engineering teams running agents in production who want routing, tracing, and evals in one platform.

OpenRouter is a hosted model marketplace and gateway where one API key and one bill cover models from many providers. Its routing options include model fallbacks, provider selection, an Auto Router that picks a model per prompt, and a Switchyard router that sends each request to the cheapest model in your shortlist that can handle it. Ownership is also changing, since Stripe announced in August 2026 that it had agreed to acquire OpenRouter.
Logs and an activity dashboard cover spend, tokens, and individual generations, and Broadcast sends traces out to Langfuse, LangSmith, Braintrust, Datadog, or any OpenTelemetry backend. Scoring output happens in whichever of those tools you send traces to, so a quality check on production traffic means running a second product next to the router.
OpenRouter charges a 5.5% fee on credit purchases, and bring-your-own-key traffic is free up to $25,000 of list-price inference per month before a 5% fee applies. Because that fee scales with every dollar of credits, it's worth weighing against other OpenRouter alternatives once monthly spend grows.
Pros:
Cons:
Pricing: Pay-as-you-go with a 5.5% fee on credit purchases; BYOK is free up to $25,000 a month in list-price inference, then 5%.
Best for: Teams that want one bill across many providers and already run a separate observability or eval tool.

If your stack already uses the AI SDK, Vercel AI Gateway is the managed gateway built to sit behind it, though your app doesn't have to run on Vercel to use it. It routes each model across healthy providers, supports ordered provider and model fallbacks, and covers image, video, speech, and embeddings alongside text.
Request logs record each call's provider attempts, latency, tokens, status, and cost, and budgets can be set per team, project, API key, or team member. Those budgets work as soft caps on system-credential spend, though, and BYOK spend doesn't count toward them, so a team on its own provider keys needs another way to cap spend.
Vercel adds no markup to provider token prices, including with BYOK, and each account gets $5 of gateway credit a month. Custom reporting, provider allowlisting, and zero data retention are metered add-ons, which is where costs tend to show up in any comparison of Vercel AI Gateway alternatives, since the token price itself stays flat.
Pros:
Cons:
Pricing: $5 monthly credit, then pay-as-you-go at provider rates with no markup; reporting, allowlisting, and zero data retention are metered add-ons.
Best for: Teams on the AI SDK or Vercel who want a managed gateway with request logs and budgets.

For teams that want to run the gateway themselves, LiteLLM is an open-source proxy that puts 100+ LLMs behind one OpenAI-compatible interface. It handles load balancing, fallbacks, traffic mirroring for A/B tests, caching, guardrails, and per-key and per-team budgets through virtual keys, with an admin UI for logs and spend.
What you take on with tools like LiteLLM is the infrastructure around them. LiteLLM's own production reference uses PostgreSQL for keys and spend data, Redis for rate limits and caching across instances, and a secrets manager for credentials, and your team owns upgrades and uptime for all of it. Observability works through callbacks to tools like Langfuse or an OpenTelemetry backend, so tracing depth and any scoring of output depend on what you connect.
The open-source proxy has no license fee, so the real cost is infrastructure and engineering time. LiteLLM's Enterprise pricing comes through a sales conversation, so it's a bit more difficult to find.
Pros:
Cons:
Pricing: Free open source; Enterprise is custom.
Best for: Platform teams with the capacity to run and maintain their own gateway.

Bifrost is Maxim's open-source AI gateway, written in Go under the Apache 2.0 license, and it starts with a single npx command or a Docker container. It covers 20+ providers with automatic fallbacks, weighted load balancing across API keys, semantic caching, an MCP gateway, and budgets and rate limits at the virtual key, team, and customer level.
Its observability is built for teams that already run their own monitoring stack, with native Prometheus metrics and OTLP export for distributed tracing in tools like Grafana or Honeycomb. Evals come from Maxim's separate platform through a plugin rather than from the gateway itself.
The open-source build is free, and Enterprise is custom-priced. Several features a production team may need sit in the Enterprise tier, including content safety guardrails, clustering for high availability, RBAC, and audit logs, so the free build and the version you'd run at scale can differ more than the license suggests.
Pros:
Cons:
Pricing: Free open source; Enterprise is custom.
Best for: Teams that want a self-hosted Go gateway feeding an existing Prometheus or OpenTelemetry setup.

Portkey pairs an AI gateway with observability, guardrails, and prompt management, and its gateway is open source. Palo Alto Networks completed its acquisition of Portkey in May 2026, and the product now carries the Prisma AIRS AI Gateway name.
Beyond routing, much of Portkey's depth is in guardrails. They check inputs and outputs, either synchronously before a request reaches the model or the response reaches the user, or asynchronously with no added latency, and they mix 20+ deterministic checks with LLM-based ones like prompt injection scanning. Results can be logged as feedback on each request, so live scoring is tied to whichever checks you configure.
Portkey's Developer tier is free, Production costs $49 a month for 100k logs with overages at $9 per additional 100k requests, and Enterprise adds VPC hosting and SSO. Because logs are the metered unit, the bill tracks traffic, which is worth modeling before comparing Portkey alternatives on headline price.
Pros:
Cons:
Pricing: Free Developer tier; Production is $49/month for 100k logs plus $9 per additional 100k requests; Enterprise is custom.
Best for: Security-focused teams that want guardrails and governance on every request.
For teams running agents and LLM apps in production, Respan is the OrcaRouter alternative to choose, because it routes every model call through one API and then traces and scores each routed request in the same platform. OrcaRouter's evals run against test cases you write, with no scheduling yet, while Respan runs the same evaluators on sampled production spans, builds test sets from real requests, and links every score to the trace behind it.
The AI gateway with built-in observability and evals also handles fallbacks and hard spend limits on the same path, so there's no second integration to maintain as traffic grows.
The rest of the field fits narrower needs. OpenRouter suits teams that want one bill across many providers, and Vercel AI Gateway fits apps built on the AI SDK. LiteLLM and Bifrost suit teams that need to self-host the gateway, while Portkey leans toward security teams that want guardrails on every request.
Use OpenRouter if you want one bill across a large provider marketplace and plan to send traces to a separate observability tool. Use OrcaRouter if you want adaptive routing with no token markup, a self-hosted edition, and offline evals in the same product. The pricing gap grows with spend, since OpenRouter takes 5.5% of credit purchases while OrcaRouter charges for features instead.
Neither documents scoring output on live production traffic inside the router. If that's the gap you're trying to close, Respan routes every call through one API and runs evals on sampled production spans with the trace attached.
Yes, the Hacker plan is free and adds no markup to token prices, so you pay each provider's published rate for every call. It covers all 200+ models, auto-failover, a basic dashboard, prompt versioning, and up to 10 API keys.
Team and Enterprise are custom-priced, and Enterprise may add a per-million-token routing fee depending on volume. Eval runs are also billed as real requests, so testing a router against a large suite costs tokens even on the free plan.
Yes. Respan has a free tier with 100k logs, 7-day retention, and unlimited seats, and it includes the gateway, tracing, and evaluators on the same plan, so you can route traffic and score its output before paying anything.
LiteLLM and Bifrost are free and open source, though you run and maintain the infrastructure yourself. Vercel AI Gateway includes $5 of monthly credit with no token markup, which covers light usage before pay-as-you-go pricing starts.
OrcaRouter supports 200+ models across 12+ providers. Respan reaches 1,000+ models through one API, Portkey's homepage lists 1,600+ LLMs, OpenRouter lists 500+, and LiteLLM's proxy supports 100+. Bifrost publishes provider coverage rather than a model count, at 20+ providers, and Vercel AI Gateway lists its catalog model by model without a published total.
Once your app's models are covered, catalog size usually stops being the deciding factor. It's worth confirming the specific models and providers you call, then comparing routers on what each one does after the request.