Portkey is the third LLM gateway or observability vendor to change hands in five months. ClickHouse acquired Langfuse on January 16, Mintlify acquired Helicone on March 3, and on May 29 Palo Alto Networks acquired Portkey. Six weeks later the gateway shipped as Prisma AIRS AI Gateway, generally available and positioned as the AI control plane for the enterprise.
For the engineering teams who picked Portkey because it was the gateway with the best developer experience, that is a real shift. The roadmap is no longer about prompt iteration speed or trace depth. It is about agent identity, least-privilege execution, and inspecting every AI transaction at runtime (that is Palo Alto's framing, not ours).
Nothing breaks on a deadline here. Free signup still works, the open-source gateway is still published, and existing keys keep routing. The question is narrower than whether Portkey works: it is whether the routing, logs, prompt templates and budgets you consolidated into one integration are still being built for you, or for a security team standardizing AI traffic across an enterprise.
This guide covers 15 Portkey alternatives, what each one costs, how much work the cutover is, and how much of Portkey's surface each one actually replaces.
What Is Portkey AI?
Portkey is an AI gateway and control plane. Your application calls Portkey instead of calling OpenAI or Anthropic directly, and Portkey handles the routing, the retries, and the record of what happened.
The routing layer covers a universal OpenAI-compatible API, automatic fallbacks, load balancing, request timeouts, and both simple and semantic caching. Virtual keys wrap provider credentials with per-key budgets so a runaway agent hits a ceiling instead of an invoice. On top of that sits observability with logs, traces, custom metadata, filters and alerts, plus prompt management with versioned templates, guardrails, and an MCP gateway.
The core gateway is open source and self-hostable with no request limit, which is how a meaningful share of its users run it. The managed platform is where the observability, prompt templates, and governance live.
For teams that adopted it, the appeal was that one integration covered four jobs. That bundling is also why replacing it is harder than swapping a router, and it is the thing to hold onto when comparing options below.
Portkey Pricing
Portkey's tiers are sized by monthly volume, and the free tier draws a hard line: the pricing page says outright that Developer is not suitable for production workloads. Three-day log retention makes that concrete. An incident you investigate the following week has no data behind it.
- Open Source - Self-hosted with no request limit and no license fee. Routing, fallbacks, load balancing, retries, and basic guardrails are all in the build.
- Developer - Free forever, 10,000 recorded logs a month, three-day log retention, three prompt templates, and simple caching. No overage allowed.
- Production - $49 a month for 100,000 recorded logs with paid overage above that, 30-day log retention, unlimited prompt templates, RBAC, and semantic caching.
- Enterprise - Custom pricing for 10M or more logs, adding VPC hosting, private tenancy, SSO, configurable retention, data export to data lakes, and custom BAAs.
The line that matters for most teams sits between Production and Enterprise. Compliance certification, SOC 2 Type 2, ISO 27001, GDPR and HIPAA, appears only in the Enterprise column, as does SSO, granular budgets, and any deployment inside your own network. A regulated workload was never a $49 decision, and post-acquisition that conversation now starts with Palo Alto's sales team.
What the Palo Alto Networks Acquisition Changes
Palo Alto announced its intent to acquire Portkey on April 30, 2026 and closed the deal on May 29. Portkey now serves as the AI Gateway for Prisma AIRS, inspecting AI traffic and enforcing security and governance policy at runtime.
The integration moved fast. Prisma AIRS AI Gateway reached general availability on July 16, six weeks after close, with Palo Alto reporting 68 trillion tokens processed in the prior month and sub-millisecond routing latency. The capabilities it leads with are agent identity, least-privilege execution, runtime inspection against prompt injection, and shadow AI discovery.
Read that feature list next to the one that made engineers adopt Portkey and the divergence is clear. Nothing in the GA announcement is about prompt iteration speed, trace depth, or eval workflows. The buyer is a CISO consolidating AI traffic under existing policy, not an engineer trying to work out why an agent returned the wrong answer at 2am.
For teams running Portkey today, there is no failure date and no forced migration. What there is instead is a roadmap being written for a different customer, packaging moving toward enterprise security SKUs, and a self-serve upgrade path that now goes through a sales team. Those are the conditions under which it is worth knowing your options before you need them.
15 Best Portkey Alternatives in 2026
Two things determine how hard leaving Portkey actually is, and they pull in opposite directions.
The first is cutover cost. Most tools here speak OpenAI format, so pointing traffic at them is a base URL change. Others need infrastructure running before you can send a request, and a few need SDK instrumentation inside your application rather than a proxy in front of it.
The second is replacement scope. Portkey bundled routing, observability, prompt management and governance into one integration. A tool that takes an afternoon to adopt and replaces a quarter of that surface leaves you shopping for three more products. The fastest migration on this list is not the cheapest one.
| Tool | Cutover | Replaces | Pricing |
|---|---|---|---|
| Respan | Base URL swap | Gateway, obs, evals, prompts | Free, $199/mo Team |
| LiteLLM | Self-host, then base URL | Routing and key management | Free OSS, Enterprise custom |
| OpenRouter | Base URL and one key | Routing and model access | 5.5% fee on credits |
| Bifrost | Self-host, then base URL | Routing, caching, governance | Free OSS, Enterprise custom |
| Vercel AI Gateway | Base URL swap | Routing and spend visibility | No token markup |
| Cloudflare AI Gateway | Base URL swap | Routing, caching, logs | Free on Cloudflare plans |
| TrueFoundry | Deploy, then base URL | Routing plus model hosting | Free tier, $499/mo Pro |
| Kong AI Gateway | Plugin config, self-hosted | Routing and traffic governance | $500/mo per control plane |
| LLM Gateway | Base URL swap | Routing and cost analytics | Free, 5% fee on credits |
| Eden AI | New SDK integration | Routing across model types | 5.5% platform fee |
| Langfuse | SDK instrumentation | Observability and prompts | Free OSS, paid cloud tiers |
| LangSmith | SDK instrumentation | Observability and evals | Free tier, per-seat paid |
| Helicone | Base URL swap | Observability only | Free tier, paid tiers above |
| Braintrust | SDK instrumentation | Evals and datasets | Free tier, $249/mo Pro |
| Arize Phoenix | SDK instrumentation | Observability only | Free self-hosted |
1. Respan

Respan runs the gateway, tracing, evaluations, and prompt management on one data plane. Route, observe, and evaluate every LLM call: one endpoint reaches 1,000+ models with ordered fallback, per-key spend caps, and around 10ms added P95, and every request comes back as a trace tree showing the model attempted, whether a fallback fired, the cache result, cost, and any eval scores attached.
Plus, Respan works at scale. For example, Mem0 used Respan to build a 99.99% reliable memory layer for AI agents and run hundreds of millions of LLM calls a day across multiple providers.
Migration effort: Point your OpenAI-compatible client at a new base URL and traffic moves. It is the same one-line change Portkey users already made once, and it covers the whole surface rather than the routing slice: virtual keys become per-key spend limits, Configs become ordered fallback lists, prompt templates land in versioned prompt management, and logs land in tracing that goes deeper than Portkey's did. Nothing gets left behind for a second vendor to pick up.
Pricing: Free covers the full platform including the gateway, evals, and prompts. Team is $199 a month billed yearly, raising throughput and retention, with self-hosted deployment, SAML SSO, and a HIPAA BAA on Enterprise. SOC 2, ISO 27001, GDPR, and HIPAA are all covered.
Best fit for: Teams who used Portkey as a full developer platform and want one independent vendor covering everything it bundled.
Know which step broke, not just which request failed
Respan returns the whole agent run as a trace tree, with cost and eval scores attached to the request the gateway served. Same base URL change you made for Portkey, free to test on real traffic.
2. LiteLLM

LiteLLM is an open-source proxy that normalizes calls to more than 100 providers into OpenAI format. Routing, load balancing, per-key budgets, and spend tracking are all in the open-source build, and because you operate it, keys and traffic never leave your infrastructure. Governance features including SSO, audit logs, and JWT authorization sit in the Enterprise tier.
Migration effort: The proxy has to exist before any traffic moves, so this is infrastructure work first and a base URL change second. It needs a database, a cache layer, monitoring, and someone on call once it sits in the request path. What ports well is the routing and key management half of Portkey: fallback chains, load balancing, and virtual keys with budgets all have direct equivalents. What does not port is everything else, since observability stops at request logging and there is no prompt management or eval layer.
Pricing: The open-source gateway is free and the real cost is the infrastructure and the on-call rotation. Enterprise is quoted rather than published.
Best fit for: Teams with platform engineers who want to own the gateway outright and already have somewhere to send telemetry.
3. OpenRouter

OpenRouter exposes 400 or more models from 70 or more providers behind one OpenAI-compatible API, with auto-routing, budgets, prompt caching, and activity logs. Data policy routing lets you constrain which providers can see a request based on their logging terms. Requests transit OpenRouter's infrastructure and there is no self-hosted build.
Migration effort: This is the shortest cutover on the list: a base URL change and a single key, with no infrastructure to stand up first. It is also the narrowest replacement. You get model access and routing, and you lose prompt templates, guardrails, trace-level detail, and the governance surface, so teams who used Portkey for more than the router end up sourcing three more tools behind it.
Pricing: Tokens pass through at provider list price with a 5.5% platform fee on credit purchases, and bring-your-own-key is free up to a generous monthly ceiling before a smaller percentage applies.
Best fit for: Teams who used Portkey mainly to reach a lot of models and want the fastest possible cutover.
4. Bifrost

Bifrost is built to run on infrastructure you control. The open-source build covers routing, automatic fallback across providers and models, weighted load balancing, semantic caching, virtual keys with per-consumer budgets and rate limits, MCP support, and native Prometheus and OpenTelemetry export. Nothing in it is gated behind a paid tier, governance included.
Migration effort: Self-hosting comes first, though the deployment is lighter than most and adoption after that is a base URL change against an existing OpenAI, Anthropic, or Google client. It covers Portkey's routing, caching and governance closely, including the virtual key model. Observability is exported rather than resident, so traces and metrics land in whatever stack you already run, and evaluation lives in Maxim's separate platform rather than in the gateway.
Pricing: Free to self-host with no feature gating. Enterprise is custom and buys clustering, SAML SSO, private networking, and in-VPC or air-gapped deployment.
Best fit for: Teams who want the gateway inside their own perimeter with low overhead and already run a metrics stack.
5. Vercel AI Gateway

Vercel AI Gateway gives you one endpoint across hundreds of models with automatic retries, managed fallback, load balancing, and spend monitoring in the dashboard a Vercel team already uses. It is tightly integrated with the Vercel AI SDK. There is no self-hosted build and no VPC option, so prompts and completions transit Vercel's infrastructure.
Migration effort: A base URL change, and close to zero work if your app already deploys to Vercel. The replacement scope is narrower than with other similar tools in a specific way: you keep routing and spend visibility but the record you get back is cost and tokens per request rather than a trace of a multi-step run, and there is no prompt management or guardrail layer. Teams that adopted Portkey partly to avoid platform coupling are also trading one first-party gateway for another.
Pricing: No markup on tokens including bring-your-own-key, with optional per-request surcharges for custom reporting, provider allowlisting, and zero data retention.
Best fit for: Next.js teams whose stack already lives on Vercel and whose model calls are mostly single completions.
6. Cloudflare AI Gateway

Cloudflare AI Gateway proxies model traffic across Cloudflare's edge network, so the control plane sits close to your users rather than in one region. It exposes a universal endpoint alongside OpenAI-compatible and Anthropic-compatible paths, and a default gateway is created on your first request. Caching, rate limiting and analytics are included on every plan.
Migration effort: A base URL change with essentially no setup, which makes it one of the easier moves here. What you take on is a narrower record: observability is request-level rather than trace-level, so an agent run arrives as a sequence of unrelated calls. Persistent log storage is also capped by plan, and once you hit the ceiling new logs stop being written until you delete old ones, which puts a manual step between you and the data.
Pricing: Core gateway features cost nothing beyond your existing Cloudflare plan, with log storage the metered part and Logpush billing separately.
Best fit for: Teams already on Cloudflare who want edge routing and caching and do not need trace depth.
7. TrueFoundry

TrueFoundry pairs an LLM gateway with a model deployment platform, so the same product routes to hosted provider APIs and to models you run on your own GPUs. Routing covers failover, retries and load balancing, with virtual keys, budgets, SSO, RBAC and audit logs for governance, plus an MCP gateway. It runs as SaaS, in your cloud account, or on-prem.
Migration effort: Deployment into your environment comes before the base URL change, and the platform is considerably broader than a gateway, so onboarding includes a model registry, fine-tuning orchestration and GPU scheduling that a routing-only team will not use. On replacement scope it covers Portkey's routing and governance well and adds self-hosted inference on top, though the governance controls that matter most sit behind the paid tier rather than the free one.
Pricing: A free Developer tier is enough to evaluate rather than run. Pro is $499 a month and is where RBAC, per-team budgets, and rate limiting unlock, with self-hosted model compute billing through your own cloud on top.
Best fit for: Teams who need one vendor for both API routing and self-hosted model inference inside their own VPC.
8. Kong AI Gateway

Kong added AI routing to its existing API gateway rather than shipping a separate product, so model traffic is governed by the same plugin architecture, policies and control plane as the rest of your APIs. The governance surface is broad: PII sanitization, semantic prompt guards, per-user and per-model token quotas, and chargeback across internal teams.
Migration effort: This is the heaviest cutover on the list. There is no OpenAI-compatible base URL to point at, so adoption means configuring AI Proxy plugins inside a full API gateway you also have to run. What you get in return is traffic governance that goes well past what Portkey offered, and what you lose is the LLM-shaped half: Kong emits telemetry rather than analyzing it, visibility is request-shaped, and there is no eval layer or prompt management at all.
Pricing: $500 a month per control plane, with separate environments or regions each paying again, and AI capabilities metering separately per unique model routed.
Best fit for: Organizations already standardized on Kong who want model traffic under the same governance as everything else.
9. LLM Gateway

LLM Gateway is an open-source router reaching 200 or more models across 40 or more providers through an OpenAI-compatible endpoint, with cost and latency analytics, secure key management, budgets, prompt caching, and guardrails covering prompt injection and PII detection. It is SOC 2 Type II certified and can be self-hosted under AGPLv3 or used as a hosted platform.
Migration effort: A base URL change on the hosted platform, or a self-host first if you want the gateway inside your own network, which is a genuine choice rather than an Enterprise upsell. It replaces Portkey's routing, key management and cost analytics closely, and the per-request cost breakdown is more detailed than most routers at this price. Prompt templates and eval workflows have nowhere to land, so those need a second home.
Pricing: Free forever with a flat 5% fee on credit top-ups, and no fee at all when you bring your own provider keys. Self-hosting is free, and Enterprise adds SSO, audit logs, an uptime SLA, and volume discounts.
Best fit for: Teams who want OpenRouter's shape with a self-host option and a lower fee on credits.
10. Eden AI

Eden AI routes across a much wider surface than LLMs alone, covering OCR, speech-to-text, text-to-speech, translation, vision and document parsing through one unified API alongside the model providers. Smart routing and fallbacks select by cost, latency or execution region, and its European infrastructure and EU endpoint make data residency a configuration rather than an Enterprise negotiation.
Migration effort: Not a drop-in swap. Eden AI has its own API surface rather than an OpenAI-compatible endpoint, so this is an integration rather than a base URL change, and the fallback and budget logic gets rebuilt from scratch. The replacement scope is unusual: broader than Portkey across model types, narrower within LLM workflows, with no tracing, prompt versioning or evaluation to inherit what you had.
Pricing: Pay as you go with no markup on provider pricing and a 5.5% platform fee at checkout, unlimited seats included. An Advanced tier adds private deployments, bulk discounts and an SLA.
Best fit for: Teams whose AI stack spans OCR, speech and vision alongside LLMs, particularly with EU data residency requirements.
11. Langfuse

Langfuse is a tracing backbone with prompt management attached, and the MIT-licensed core is genuinely usable self-hosted rather than a limited demo build. OpenTelemetry support is solid, the trace UI handles multi-step runs well, and callback handlers cover the OpenAI SDK, LangChain, LlamaIndex, LiteLLM and others. ClickHouse acquired the company in January 2026.
Migration effort: SDK instrumentation inside your application rather than a proxy in front of it, which is more work than a base URL change and buys correspondingly more depth on agent runs. It replaces Portkey's observability and prompt management and none of its routing, so you need a gateway underneath it. Langfuse also logs traces without scoring them, so quality monitoring means writing your own LLM-as-judge logic.
Pricing: The self-hosted core is free to license and costs whatever Postgres plus ClickHouse plus containers costs to run. Managed cloud has a free tier with paid plans above it.
Best fit for: Teams with data residency requirements and the appetite to build their own evaluation layer.
12. LangSmith

LangSmith comes from the LangChain team, and traces render the full execution tree including tool selections, retrieved documents, and model parameters. Annotation queues route specific traces to domain experts for labelling, and that output feeds evaluation datasets.
Migration effort: SDK instrumentation, and the depth you get depends heavily on your framework, since trace fidelity outside LangChain and LangGraph is noticeably thinner. On scope it covers observability and evals and leaves routing entirely unaddressed, so the gateway half of Portkey needs replacing separately. Self-hosting is Enterprise only, which rules it out for teams who left over deployment control.
Pricing: A free developer tier, then per-seat pricing with trace volume metering separately, and custom Enterprise above that.
Best fit for: Teams whose stack is LangChain today and will still be LangChain in two years.
13. Helicone

Helicone instruments by proxy, so changing a base URL brings request logs, caching and cost analytics without touching application code. Mintlify acquired it on March 3, 2026 and the platform is now in maintenance mode, defined by the team as security updates, bug fixes, and new model support with feature development ended.
Migration effort: As easy a cutover as exists, and the same proxy shape Portkey used, so the mechanics are familiar. The scope is narrow, covering observability and basic gateway behavior without prompt management or evals, and proxy instrumentation is request-shaped, so agent chains are largely invisible. There is also a harder question than effort here: moving off one acquired product onto another whose roadmap has already stopped is a short-lived migration.
Pricing: A free tier covers prototyping, with paid tiers above it billed on volume and compliance coverage arriving at the higher tier. The open-source build stays free to self-host.
Best fit for: Existing Helicone installs. Teams choosing fresh today have better options above, and Respan maintains a fuller breakdown of Helicone alternatives.
14. Braintrust

Braintrust treats evaluation as the product. Datasets, scoring functions, comparison reports and regression testing are all first-class, and experiment comparison across prompt versions is more rigorous than most platforms attempt.
Migration effort: SDK instrumentation, and a narrow slice of the problem. It replaces the eval workflow Portkey barely had and none of the routing, caching, or governance that Portkey actually did, so this is an addition to a stack rather than a replacement for one. Tracing exists to serve the eval workflow, which makes it thinner than tools built for production debugging, and self-hosting is Enterprise only.
Pricing: A free developer tier, Pro at $249 a month, and undisclosed Enterprise pricing that tends to arrive sooner than teams plan for, since eval volume grows with the test suite rather than with traffic.
Best fit for: Teams where eval discipline is the constraint and routing is already solved elsewhere.
15. Arize Phoenix

Phoenix is the open-source half of Arize, running in a notebook, locally, or via Docker with no external dependencies, which makes it usable during development rather than only after deploy. Instrumentation uses OpenInference, built on OpenTelemetry, covering LlamaIndex, LangChain, Haystack, DSPy and others.
Migration effort: OpenTelemetry instrumentation, which is the most portable option here and the cheapest to reverse if you change your mind again. Scope is observability only, so routing, caching, prompt templates and governance all need homes elsewhere, and the LLM evaluation layer is the shallow part of a product built for model monitoring first. Custom evaluators are supported; research-backed metrics for faithfulness and hallucination are not.
Pricing: Phoenix is free and self-hosted. The managed Arize AX platform has a capped free tier with paid plans above it.
Best fit for: Teams already running OpenTelemetry who want tracing they control and will build evaluation separately.
How to Choose a Portkey Alternative
Start by writing down what you actually used Portkey for, because the answer usually surprises people. Teams who describe it as "our gateway" often turn out to have prompt templates in production, budgets on virtual keys, and dashboards someone checks weekly. Every one of those is a separate product on this list unless you pick one that bundles them.
Then work through what disqualifies rather than what appeals:
- Deployment - If model traffic cannot leave your network, the managed-only options are gone in one pass and the decision is a different one.
- Cutover budget - A base URL change fits in an afternoon. Standing up a proxy with a database, a cache and an on-call rotation is a project, and Kong is a larger project still.
- What you need back per request - Cost and tokens per call is enough for single completions. An agent making a dozen calls per user action needs the whole run as a tree, and a tool that only records the request boundary cannot reconstruct one after the fact.
- Whether quality gets measured - Portkey ships eval templates, but evaluation was never the product and it does not score live traffic. If a wrong answer that returned HTTP 200 is your real failure mode, scoring production requests is the requirement, not a test suite you run by hand.
- Independence - Portkey is the third gateway or observability vendor acquired in five months, alongside Langfuse and Helicone. Ask who funds a vendor and what happens to your data if they are next.
Most of these tools solve one of those cleanly. The ones worth shortlisting are the ones that solve the two or three that apply to you at once, because the alternative is correlating request IDs across three dashboards during an incident.
How to Migrate Off Portkey
The mechanics are easier than the inventory. Almost everything here speaks OpenAI format, so the code change is small. What takes time is the configuration Portkey held that does not port anywhere.
- Export your logs before retention lapses. Log retention runs three days on Developer and 30 on Production, so anything you want for baselining needs pulling now rather than after the cutover.
- Inventory your Configs. Fallback chains, load balancing rules, retry policy and cache settings live in Portkey Configs and have no export format. Write them down explicitly, including the order.
- Map your virtual keys. Each one carries a budget, a rate limit and a set of permissions. These become per-key spend caps or budget rules in the new tool, with different names and different scoping.
- Account for prompt templates. If you deployed prompts by ID rather than shipping them in code, only a few destinations on this list can receive them. Check this before you pick, not after.
- Run both in parallel for a day. Send real traffic through the new path and reconcile token counts and cost against your provider invoices before tearing anything out.
- Verify you kept what you moved for. If the reason was visibility, confirm you can now answer which model served a request, whether a fallback fired, and whether output quality shifted. If you cannot, the migration solved the ownership question and left the harder one in place.
Frequently Asked Questions
What is the Portkey LLM gateway?
The Portkey LLM gateway is the routing layer that sits between your application and model providers, exposing one OpenAI-compatible endpoint across 1,600 or more models. It handles automatic fallbacks, load balancing, retries, request timeouts, and simple and semantic caching, with virtual keys wrapping provider credentials in per-key budgets. The core is open source and self-hostable with no request limit. Since the Palo Alto acquisition it also ships as the AI Gateway inside Prisma AIRS. For a broader look at the category, see how LLM gateways differ from plain API proxies.
Portkey vs LiteLLM: which one should you use?
They answer different questions. Portkey is a managed platform where observability, prompt management and governance come with the routing, and someone else runs it. LiteLLM is an open-source proxy you operate, which gives you full control of keys and traffic and hands you a database, a cache layer, monitoring and an on-call rotation along with it. Pick LiteLLM if data control is the requirement and you have platform engineers to spend. Pick a managed platform if the gateway needs to answer questions rather than just forward requests. Respan covers the middle case, running the gateway with tracing, evals and prompts on the same data plane and self-hosted deployment available on Enterprise. There is a fuller breakdown in our guide to LiteLLM alternatives.
Who acquired Portkey?
Palo Alto Networks, in a deal announced April 30, 2026 and closed May 29, 2026. Portkey now serves as the AI Gateway for Prisma AIRS, Palo Alto's AI security platform.
Is Portkey free?
There are two free paths. The open-source gateway is free to self-host with no request limit and covers routing, fallbacks, load balancing and retries. The hosted Developer tier is free forever with 10,000 recorded logs a month and three-day retention, and Portkey's own pricing page says it is not suitable for production. Paid plans start at $49 a month.
What is the best Portkey alternative?
For teams who used Portkey as a full developer platform, Respan is the closest replacement, because it covers the gateway, observability, evaluations and prompt management that Portkey bundled rather than one slice of it, and the cutover is the same base URL change. If you only ever used the router, LiteLLM and Bifrost are strong open-source options and OpenRouter is the fastest cutover on the list. If you need model traffic governed alongside the rest of your APIs, Kong is the fit.
Is there a free Portkey alternative?
Several, in two shapes. Respan's free tier covers the full platform including the gateway, tracing, evals and prompt management, which is the closest free equivalent to what Portkey bundled. LiteLLM, Bifrost, Langfuse and Phoenix are all free to self-host, where the cost is infrastructure and engineering time rather than a subscription. LLM Gateway is free with your own provider keys, and Cloudflare AI Gateway's core features are included with any Cloudflare plan.



