Skip to content

What is an LLM Gateway?

Arrow down icon
What is an LLM Gateway?
What is an LLM Gateway?

TL;DR

  • An LLM gateway is a middleware layer between your applications and your LLM providers that centralizes routing, failover, rate limiting, caching, and cost tracking through a single API.

  • Direct API integration works fine for one team calling one model. It breaks when multiple teams, models, and applications are involved.

  • A gateway handles the traffic layer only. It does not manage retrieval, workflow orchestration, or data integration.

  • Gateway tools range from lightweight open-source proxies to full platforms where routing is one layer in a broader architecture.

In January 2026, OpenAI logged 11 incidents in 28 days. Anthropic had multiple incidents per week during the same period, with one resolution cycle for a Claude model stretching to 30 hours. 

The Nordic APIs reliability report, which tracked 215+ services across 25 categories from October 2025 through February 2026, ranked AI and ML APIs last for reliability across every category it measured, and not by a small margin. 

For every application calling those APIs directly, there was no automated fallback. When a provider went down, the application went down with it, and engineers absorbed the cost of managing the recovery manually.

That is the core problem an LLM gateway solves. Not provider outages themselves, because those will keep happening. The gap it closes is on the application side: a single layer that handles routing, failover, and recovery automatically, so a provider having a bad week does not become your problem every time. 

What is an LLM Gateway?

An LLM gateway is a middleware layer that sits between your applications and your LLM providers. Instead of each application connecting directly to OpenAI, Anthropic, Google, or a self-hosted model, all traffic routes through the gateway. 

The gateway handles authentication, translates requests into provider-specific formats, routes traffic based on rules you define, and returns standardized responses back to your application.

Core functions include:

  • Unified API: One interface across all providers. Your application code stays the same whether the request goes to GPT, Claude, or a self-hosted Llama instance.
  • Request routing: The gateway directs each request to the right provider based on cost, latency, availability, or rules you configure.
  • Failover: When a provider goes down or hits rate limits, the gateway automatically reroutes to a backup without your application registering the problem.
  • Rate limiting: Controls how many requests each team, user, or application can make in a given time window.
  • Response caching: Stores responses to repeated or semantically similar queries so you do not pay for the same inference twice.
  • Cost tracking: The gateway logs token usage and spend per request, attributable to the team, project, or application that generated it.

AI gateway” and LLM gateway are often used interchangeably. An AI gateway is the broader term, covering language models plus vision, audio, and other modalities. An LLM gateway is specific to language model traffic. For most organizations working primarily with text-based models, the distinction does not matter in practice.

How Does an LLM Gateway Work?

When an application sends a request, the gateway intercepts it before it reaches any provider. It validates the API key, checks permissions, and applies rate limits, rejecting anything that fails those checks before a single token gets consumed.

The gateway then routes the request based on rules you configure: cost, latency, availability, or custom logic. It translates your standardized request into the provider’s expected format, makes the API call, and translates the response back into a consistent structure your application always expects. Once the response returns, the gateway logs which model handled it, how many tokens it consumed, how long it took, and what it cost.

The entire round-trip adds single-digit milliseconds of overhead, which is imperceptible compared to model inference that typically takes hundreds of milliseconds to several seconds.

When Do You Need an LLM Gateway?

Not every organization needs a gateway. Direct API integration works fine in certain situations, and adding infrastructure you do not need creates complexity without payoff.

One team calling one model from one application does not need a gateway. The integration is straightforward, the API key lives in one place, and you can track costs by checking your provider dashboard.

The inflection points that signal you have outgrown direct integration tend to follow a pattern:

  • Credentials start scattering: API keys live in environment variables across multiple applications managed by different teams, and nobody has a complete picture of which keys exist, where they are stored, or who has access.
  • Cost becomes invisible: Multiple teams use the same provider account, the monthly invoice shows total spend, but you cannot attribute costs to the team, project, or application that generated them.
  • Provider outages cascade: One provider goes down and every application that calls it directly goes down with it. There is no fallback because each integration was built independently, and building retry-and-reroute logic into every application separately is exactly the work a gateway does once.
  • Switching providers means rewriting code: You want to try a new model, but every application has hardcoded the current provider’s request format and response parsing, so a change that should take an afternoon becomes a multi-sprint project.
  • Teams spin up their own accounts: A new team needs model access, and rather than wait for procurement, they create their own provider account with a corporate card, putting model usage outside any centralized tracking or policy.

If two or more of those sound familiar, a gateway is worth the investment. The operational cost of managing that fragmentation exceeds the cost of the gateway.

Where a Gateway Fits in the AI Stack

A gateway handles one specific layer: the traffic between your applications and your models. That layer matters more as model usage scales across teams, but it helps to be precise about what a gateway does not handle, because the boundaries define where you will need other tools.

For organizations where routing is the only AI infrastructure problem, a standalone gateway is the right tool. For organizations that need routing alongside data integration, retrieval, orchestration, and delivery, the gateway function typically lives inside a broader platform. Understanding that distinction before you evaluate tools saves significant time.

Top LLM Gateway Tools in 2026

The gateway spectrum runs from lightweight proxy to full platform. The further along the spectrum, the more the tool handles beyond traffic management.

Open-Source Proxy

LiteLLM

LiteLLM is the most widely adopted open-source gateway. It provides a Python-based proxy server that unifies access to 100+ LLM APIs in OpenAI-compatible format. You configure it with a YAML file, point your applications at it, and it handles translation, routing, and basic logging. Budget and rate limit management per user or team is built in, and the open-source community is active.

While the software license costs nothing, self-hosting in production requires a server or container, a database for logging and configuration, monitoring, and ongoing engineering time to maintain and update it. 

The limitations: no formal SLAs or dedicated escalation path on the open-source tier. Users report regressions between versions and instability at higher scale. Enterprise tier pricing is sized to annual gateway request capacity and deployment architecture, not published publicly.

Deployment: Self-hosted (open-source); Enterprise Standard and SCALE add air-gapped support. 

Best for: Development teams and startups that want provider abstraction without vendor commitment and have the DevOps capacity to run their own infrastructure.

Managed Gateway-as-a-Service

Portkey

Portkey positions itself as an LLMOps platform rather than a pure gateway, adding prompt management, guardrails integration, and observability alongside routing and failover. It provides access to 1,600+ models with SOC2, ISO27001, HIPAA, and GDPR compliance and a 99.99% uptime SLA. 

The limitations: enterprise pricing runs $2,000-$10,000+/month depending on volume, retention, deployment model, and support level. Key features like budget limits and air-gapped deployment are restricted to enterprise tier. Model deployment is not natively supported.

Deployment: Managed SaaS, hybrid, air-gapped (Enterprise). 

Best for: Teams that want gateway capabilities plus prompt management and guardrails in one vendor, and are operating at enterprise scale.

Cloudflare AI Gateway

Cloudflare extends its edge network to LLM traffic. Core features (caching, rate limiting, analytics) are free on all Cloudflare plans with no per-call gateway fee and no token markup. Log storage limits vary by plan, and free tier users get 100,000 AI Gateway logs per month across all gateways before hitting limits. If you already use Cloudflare, adding it takes a single line of code.

The limitations: limited routing intelligence compared to dedicated gateway tools, no advanced budget management by team, and tightly coupled to the Cloudflare ecosystem. Logpush requires the Workers Paid plan.

Deployment: Cloudflare edge (managed SaaS). 

Best for: Teams already on Cloudflare that want basic gateway capabilities at low or no cost without adopting a new platform.

Orq

Orq provides 500+ models across 30+ providers behind a single OpenAI-compatible API, alongside prompt management, evaluation tooling, knowledge base, agent runtime, MCP gateway, and AI governance including EU AI Act compliance tooling. Its pay-as-you-go tier includes 1M BYOK requests per month free, with a 4% fee after that, and a 4.5% fee on Orq-managed model credits. 

The limitation: it is more platform than teams need if they only want a routing proxy. Teams evaluating Orq purely as a gateway are underusing what they are paying for.

Deployment: EU SaaS (pay-as-you-go), VPC and on-prem (Enterprise). 

Best for: Teams that want gateway routing connected to evaluations, prompt management, and AI governance in one platform, particularly those with EU data residency requirements.

Cloud-Native API Extension

Kong AI Gateway

Kong extends its established API gateway platform to handle AI model traffic through a plugin architecture. For organizations already running Kong for API infrastructure, the AI Gateway plugins add LLM-specific routing, rate limiting, and analytics without introducing a new tool.

Pricing is complex. The open-source version (Kong OSS) is free but lacks GUI, RBAC, analytics, and the AI Gateway plugins themselves, which require paid tiers. Kong Konnect Plus bills per gateway per month plus per-request overages. 

The limitation: adopting Kong only for LLM routing introduces more infrastructure and licensing complexity than most AI teams need. It earns its cost when you are already running Kong across a broader API estate.

Deployment: Self-hosted (OSS), managed SaaS (Konnect), hybrid. 

Best for: Organizations already running Kong for API management that want to extend it to AI traffic without adopting a separate tool.

Full Platforms with Routing Built In

TrueFoundry

TrueFoundry provides a complete AI infrastructure platform where the gateway is one component alongside model deployment, training, fine-tuning, and MCP gateway capabilities. Sub-10ms gateway overhead, granular cost attribution by user, team, or geography.

The limitation: the comprehensive feature set may exceed what teams need if they only want a routing layer.

Deployment: SaaS, VPC, on-premises, air-gapped. 

Best for: Organizations that need AI infrastructure beyond routing: model deployment, fine-tuning, and operational tooling in one platform.

UNIFI (AISquared)

UNIFI treats model routing as one layer in a seven-layer enterprise AI architecture called the AI Controls Model. Its routing applies policy-driven selection at runtime based on workspace, role, data sensitivity, task type, latency targets, and cost guardrails, with every routing decision logged alongside the selected model and governing policy.

That routing layer sits alongside data connectivity to 50+ enterprise systems, retrieval-augmented generation, workflow orchestration, and embedded delivery. UNIFI originated in Department of Defense environments where deterministic execution and air-gapped deployment are non-negotiable, and those same design constraints carry into its commercial deployments in financial services, healthcare, and defense.

The limitation: UNIFI is a platform, not a standalone gateway product. Organizations that only need a routing proxy will find it overscoped.

Deployment: SaaS, customer-managed cloud (VPC), air-gapped on-premises. 

Best for: Enterprises that need model routing as part of a governed AI deployment stack, particularly in regulated industries and defense environments.

Gateway Spectrum at a Glance

CategoryToolsBest when you need
Open-source proxyLiteLLMProvider abstraction, fast setup, DevOps capacity to self-host
Managed servicePortkey, Cloudflare AI Gateway, OrqProduction routing with compliance, prompt management, or edge performance
Cloud-native extensionKong AI GatewayAI routing layered onto an existing Kong API management estate
Full platformUNIFI, TrueFoundryModel routing alongside data integration, orchestration, or full AI operations

How to Choose an LLM Gateway

Three questions tend to narrow the field quickly: 

  1. What are your deployment constraints? 

If your data cannot leave your network, SaaS-only tools drop off the list immediately. If you need air-gapped operation, the list gets very short. Deployment model is the single biggest disqualifier for defense, financial services, and healthcare buyers, and it is the question most gateway evaluations address last instead of first.

  1. Build or buy? 

An open-source proxy gives you flexibility and control, but your team owns the operational burden: infrastructure, updates, scaling, monitoring, and incident response. A managed service trades some flexibility for operational simplicity. The right answer depends on whether your platform team has bandwidth for another piece of infrastructure to maintain.

  1. What else are you solving? 

If model routing is the only problem on the table, a standalone gateway is the right tool. If you are also solving data integration, retrieval, orchestration, or delivery, look at whether a broader platform handles routing as part of its architecture. Consolidating reduces integration work, and over-consolidating locks you into a single vendor’s roadmap for problems you have not encountered yet.

Conclusion

AI and ML APIs are the least reliable category of infrastructure being tracked right now, and that is not going to change while providers are shipping this fast. The answer is not to find a provider that never goes down. It is to build your stack so that when one does, you do not feel it. An LLM gateway is the layer that makes that possible.

Book a demo to see how UNIFI handles model routing across your stack.

Frequently Asked Questions

Is an LLM gateway the same as an AI gateway?

Mostly. An LLM gateway handles language model traffic specifically. An AI gateway is the broader term that also covers vision, audio, and other modalities. Most organizations use the terms interchangeably.

Can I build my own LLM gateway or should I buy one?

You can build a basic proxy in a few days, but routing, failover, caching, rate limiting, and cost attribution at production quality take significantly longer. The build-vs-buy decision comes down to how much operational burden your platform team can absorb.

Does an LLM gateway add latency?

A well-engineered gateway adds single-digit milliseconds per request, which is imperceptible compared to model inference that typically takes hundreds of milliseconds to several seconds. The failover and caching a gateway provides often reduce average response time more than the gateway hop adds.