AI applications rarely stay tied to one model provider for long. A production application may use one model for reasoning, another for fast responses, and a third for lower-cost background tasks. As those models, prices, rate limits and providers change, managing every integration directly in application code becomes difficult.
AI gateways and LLM routers add a control layer between applications and model providers. Depending on the platform, that layer can manage model routing, fallbacks, load balancing, caching, observability, security, budgets and governance.
The key difference is where that control lives. Some platforms are managed services, some can run inside your own infrastructure, while others are built into larger cloud or API platforms.

Best AI Gateway and LLM Router Platforms at a Glance
| Platform | Best for | Deployment | Price | Key capabilities | Tradeoff |
|---|---|---|---|---|---|
| Portkey | Enterprise AI governance | Managed / self-hosted | From $49/mo | Routing, fallbacks, guardrails, observability | More enterprise-focused |
| OpenRouter | Broad model access | Managed | 5.5% Standard fee | 500+ models, 80+ providers, routing | No self-hosting |
| LiteLLM | Self-hosted AI infrastructure | Self-hosted / cloud | Free OSS | Routing, load balancing, fallbacks | Your team operates the infrastructure |
| Vercel AI Gateway | Vercel and AI SDK apps | Managed | $5 monthly free credits | Provider switching, failover, unified API | Best fit within Vercel ecosystem |
| Cloudflare AI Gateway | Edge-native applications | Managed | Core features free | Routing, caching, rate limiting | Cross-provider failover needs configuration |
| Helicone | AI gateway + observability | Managed / self-hosted | Free; Pro from $79/mo | Gateway, caching, fallbacks, monitoring | Advanced usage adds usage-based costs |
| Kong AI Gateway | Enterprise API platforms | Cloud / self-hosted | From $25/mo + usage | AI/API governance, routing, observability | More infrastructure than a simple proxy |
| Requesty | Managed routing and cost control | Managed | 5% markup | Routing, caching, budgets, fallbacks | Percentage fee scales with usage |
| Merge Gateway | Production AI control plane | Managed | Free; Pro adds 5% | Routing, budgets, failover, cost controls | Broader than a basic gateway |
| Future AGI | AI agents and evaluations | Managed | Free; usage-based | Routing, guardrails, caching, evaluations | Best value comes from its broader platform |
Pricing and feature information was checked against public vendor information in October 2026. Plans, limits and model availability can change.
What Is an AI Gateway or LLM Router?
An AI gateway sits between an application and its AI providers. It gives developers a common interface while centralizing controls around requests, models and providers.
An LLM router has a narrower primary role: deciding which model or provider should handle a request. That decision can consider cost, latency, capability, availability or predefined routing rules.
The two categories increasingly overlap. A modern AI gateway may include routing as well as authentication, caching, observability, rate limits and governance. Vercel's current AI gateway comparison also emphasizes deployment model and failover behavior as important factors when evaluating gateways.
How AI Gateways Route Requests
In a production application, a request can pass through the gateway before reaching the selected model. The gateway evaluates the request against routing rules, chooses an eligible model or provider, forwards the request and can move to a fallback if the first option fails.
Routing policies can consider model capability, cost, latency, provider availability, user or team metadata, budgets and geographic requirements. LiteLLM, for example, supports load balancing, cooldowns, retries and fallbacks across multiple deployments and providers.
Cloudflare's Dynamic Routing takes a more configurable approach, allowing conditions, traffic splits, budgets, rate limits, model selection and fallback policies.
For platform engineers, this matters because routing logic can stay in the infrastructure layer instead of being duplicated across individual applications.
How We Evaluated These AI Gateway Platforms
We looked at the parts that affect a production AI stack rather than simply counting features. The comparison considers provider and model coverage, routing flexibility, fallback behavior, caching, observability, guardrails, deployment options, pricing structure and self-hosting.
We also considered where each gateway fits architecturally. This matters because a self-hosted gateway, a model marketplace and an enterprise API gateway solve related problems from different positions in the stack.
1. Portkey
Portkey is an AI gateway focused on production control and governance. It combines routing with fallbacks, load balancing, retries, observability, guardrails, prompt management and caching.
Its Production plan currently starts at $49 per month and includes 100,000 recorded logs, while enterprise pricing is custom.
Portkey is particularly relevant for teams that need centralized AI policies across multiple applications and providers. Palo Alto Networks completed its acquisition of Portkey in May 2026, integrating the technology into its Prisma AIRS strategy for governing AI agents.
2. OpenRouter
OpenRouter is built around broad access to AI models and providers through a unified API. Its current Standard plan lists more than 500 models and 80+ providers, with automatic routing and preferred vendor selection. The Standard platform fee is 5.5%.
For engineering teams evaluating different models without building individual integrations for every provider, this broad catalog is a major advantage.
Its main limitation is deployment flexibility. OpenRouter is a managed service, so teams that need to operate the gateway inside their own infrastructure will need another approach.
3. LiteLLM
LiteLLM is an open-source option for teams that want direct control over their AI gateway infrastructure.
Its router supports load balancing, retries, cooldowns, timeouts and fallbacks across deployments and providers. Production deployments can also use Redis for usage and cooldown tracking.
LiteLLM is particularly useful when self-hosting, network control or custom infrastructure requirements matter. The tradeoff is operational responsibility: your team manages deployment, scaling, upgrades and reliability.
4. Vercel AI Gateway
Vercel AI Gateway is designed for applications already using Vercel and the AI SDK. It provides a unified interface across available models and handles provider access, usage tracking and failover.
Vercel currently offers $5 in monthly AI Gateway credits on its free tier and says its paid tier uses zero markup, including when customers bring their own provider key.
This makes it attractive when Vercel is already part of the application architecture. Teams looking for a fully self-hosted gateway should consider another option.
5. Cloudflare AI Gateway
Cloudflare AI Gateway is designed around Cloudflare's edge infrastructure.
Its core features, including analytics, caching and rate limiting, are currently available for free. Cloudflare also provides Dynamic Routing for conditional routing, traffic splits, budgets and fallbacks.
This is a strong choice for applications already using Cloudflare services. One detail matters for production planning: automatic retries and cross-provider failover are different capabilities, with Dynamic Routing required for provider-level failover scenarios.
6. Helicone
Helicone combines an AI gateway with observability. Its gateway supports model switching, failovers, rate limiting and caching, while the broader platform focuses on monitoring AI traffic.
Its current pricing starts with a free Hobby plan offering 10,000 requests, while Pro starts at $79 per month plus usage-based pricing. Enterprise plans include on-premise deployment.
Helicone makes sense for teams that want gateway functionality closely connected to AI observability rather than maintaining separate systems.
7. Kong AI Gateway
Kong AI Gateway extends API gateway concepts into LLM, MCP and agent traffic. It provides routing, authentication, governance, quotas and observability alongside broader API management capabilities.
Kong's AI Management Plus plan starts at $25 per month plus usage. Its platform also supports dedicated cloud, hybrid and fully self-hosted deployment models, with enterprise pricing available for larger environments.
This makes Kong particularly relevant for organizations that already use API management as part of their infrastructure strategy.
8. Requesty
Requesty is a managed AI gateway focused on routing, caching and cost control.
Its current pay-as-you-go pricing charges 5% on application spend and provides access to more than 600 models. Routing policies, caching, fallbacks, spend limits, budget caps, EU data residency and observability are included in the paid offering.
For teams that want managed infrastructure with explicit spending controls, Requesty offers a relatively straightforward model. The percentage-based fee should be included when calculating costs at higher volumes.
9. Merge Gateway
Merge Gateway positions itself as a production AI control plane. It provides intelligent routing, automatic failover and budget enforcement while tracking request-level cost and routing information.
Merge's current Pro pricing charges the provider cost plus a 5% Merge fee, while its Free plan has no platform fee.
This approach suits teams that need routing tied closely to production cost and reliability controls rather than a simple provider abstraction layer.
10. Future AGI
Future AGI combines an AI gateway with evaluation, guardrails and observability. Its Agent Command Center provides routing, caching, fallbacks, cost tracking and guardrails across multiple providers.
The gateway currently includes 100,000 requests and 100,000 cache hits per month in its free tier. Usage-based pricing starts after those allowances.
For teams building AI agents, this broader connection between routing and evaluation can be useful because the same control plane can manage both runtime behavior and quality checks.
How to Choose the Right AI Gateway
If your team wants the least infrastructure to manage
A managed gateway is usually the practical starting point. Vercel, OpenRouter, Requesty and Merge remove much of the operational work involved in running the gateway yourself. Your decision then comes down to model coverage, pricing, routing depth and the ecosystem your application already uses.
If you need the gateway inside your own infrastructure
Self-hosting becomes more important when your organization has strict network boundaries, data residency requirements or an existing platform engineering team. LiteLLM is the clearest fit in this group, while Portkey, Helicone and Kong also provide self-hosted options.
If your application already runs on a major cloud platform
Look at the gateway that fits the infrastructure you already operate. Vercel AI Gateway makes sense for Vercel and AI SDK applications, while Cloudflare AI Gateway is particularly relevant when Cloudflare's edge network is already part of the architecture.
If governance matters as much as routing
For enterprise environments, model selection is only one part of the problem. You may also need access controls, guardrails, auditability, budgets and centralized observability. Portkey and Kong are strong candidates here, while Future AGI becomes interesting when evaluations and agent workflows are part of the same platform.
If you are optimizing for model choice and cost
OpenRouter is useful when broad provider and model access is the priority. Requesty and Merge are worth considering when routing decisions also need to incorporate budgets and cost controls.
For a deeper comparison of Portkey, LiteLLM, Kong and Cloudflare, see our AI gateway comparison. If monitoring is your main concern, our guide to LLM observability platforms covers that layer separately.
AI Gateway vs LLM Router: Which One Do You Need?
An LLM router mainly decides where a request should go. An AI gateway usually provides a wider control layer around that request.
If your main requirement is choosing between models based on cost, latency or capability, an LLM router may cover what you need. If several applications and teams share models, the gateway becomes more useful because authentication, routing, budgets, caching, observability and governance can be managed centrally.
The distinction is increasingly blurred because modern platforms combine both functions.
Final Verdict
The best AI gateway depends on the architecture around it.
LiteLLM is a strong choice for teams that want self-hosted control. OpenRouter stands out for broad model and provider access. Vercel AI Gateway fits Vercel-based applications, while Cloudflare AI Gateway is well suited to edge-native infrastructure. Kong is a natural option for organizations already using enterprise API management.
Portkey is particularly relevant for enterprise AI governance, while Requesty and Merge focus heavily on managed routing and cost controls. Helicone connects gateway functionality with observability, and Future AGI brings routing, guardrails and evaluation into the same platform.
The better buying question is not simply which gateway has the most features. It is where you want control over your AI traffic to live, how much infrastructure your team wants to operate, and how much routing and governance your production workloads actually require.
FAQs
What is an AI gateway?
An AI gateway is a control layer between an application and AI model providers. It can centralize routing, authentication, fallbacks, caching, observability, budgets and governance.
What is an LLM router?
An LLM router selects the model or provider that should handle a request based on factors such as cost, latency, capability or availability.
Is an AI gateway the same as an LLM router?
They overlap, but an AI gateway generally covers a broader set of controls. Routing is one of the main functions that an AI gateway can provide.
Which AI gateway is best for self-hosting?
LiteLLM is one of the clearest choices for self-hosted AI routing. Portkey, Helicone and Kong also provide self-hosted deployment options, depending on the requirements of the environment.
Do AI gateways add a markup to model costs?
It depends on the platform. Vercel currently advertises zero markup, OpenRouter's Standard plan charges a 5.5% platform fee, Requesty charges 5%, and Merge's Pro plan adds a 5% fee to provider costs.



