How LiteLLM's AI Gateway Works

LiteLLM is an open-source AI gateway that sits between your applications and your model providers, exposing one OpenAI-compatible API and one login for 140+ providers and 1,800+ models. Requests flow through the gateway, which handles key management, routing, spend tracking, and governance before forwarding to the right backend. It's aimed at platform teams that need to give an entire organization access to models, agents, and MCP servers without becoming a bottleneck — and it can be self-hosted, including in air-gapped environments.

The core idea: one key in front of everything

Instead of every developer holding separate credentials for OpenAI, Anthropic, Bedrock, or an internal fine-tuned model, the gateway holds the provider keys and issues virtual keys to teams, projects, or apps. Your app code points at the gateway using the OpenAI SDK format, so swapping the underlying model becomes a configuration change rather than a code change.

The site describes this as putting "your full AI stack behind one key" — and the stack is broader than LLMs. Agents and MCP servers are reachable through the same gateway, not just chat completions.

What sits behind the single endpoint

Layer What the gateway fronts
Providers 140+ providers, 1,800+ models
Model types Hosted APIs, plus your own internal, fine-tuned, and self-hosted models
Beyond LLMs Agents and MCP servers
Access control One login with SSO, scoped by team, project, or app

How a request actually flows

  1. Client sends a request to the gateway using the OpenAI-compatible API format, authenticated with a virtual key rather than a raw provider key.
  2. The gateway resolves identity and policy — which team, project, or app the key belongs to, and what budget or scope applies.
  3. Routing decides the destination. The gateway sends each request to the model that should handle it, which may mean a specific provider, a self-hosted model, or a fallback.
  4. The provider call is made using the real credentials the gateway manages, drawn from the secret manager you already run.
  5. The response returns to the client, and the request is recorded for observability and spend tracking.

Because the interface is OpenAI-compatible, the client-side change is typically just a base URL and key swap. The site notes day-zero support for new models, so a newly released model can be exposed to developers the day it ships.

Observability and governance

The gateway's dashboard shows virtual keys with their team, last-active time, and spend against budget. In the example shown on the site, keys like acme-prod-gateway (platform-eng) and wayne-data-science (data-science) display running totals such as $4,182.55 / $10,000 and $6,740.15.

That structure answers two questions platform teams care about:

  • Who is driving spend? Spend is attributed per key, which maps to a team, project, or app.
  • How do you cap it before it runs? Budgets are attached to keys, so limits are enforced at the gateway rather than discovered on a provider invoice.

The site frames the ownership model as "your keys, your infra, your audit trail" — relevant if you need requests to stay within your own infrastructure and logging.

Routing and optimization

Routing is the mechanism that makes the single endpoint useful rather than just convenient. The gateway can send each request to the model that should handle it, which supports:

  • Model swaps without app changes — as Okta's Dennis Henry describes it, switching the backend model is "a simple configuration update in the gateway; no code changes, procurement cycles, or repetitive security reviews required."
  • Cost-driven routing — AT&T's Mark Austin reports that after adopting LiteLLM as a router provider, costs for some advanced AI tasks such as coding fell by as much as 56%.
  • Access at scale — NVIDIA's Ajay Dogra describes LiteLLM as giving engineers "a single, consistent way to access more than 100 AI model endpoints."

These are vendor-published testimonials, so treat the specific percentages as reported outcomes from those teams rather than guaranteed results. The mechanism they point to — centralizing routing so the backend is a config decision — is the part you can evaluate for your own setup.

Performance and deployment

Two claims matter for whether this fits your stack:

  • Overhead: the site states sub-millisecond overhead and links to a benchmark. If your workload is latency-sensitive, read that benchmark rather than assuming the number applies to your traffic pattern.
  • Deployment time: the site says you can "go live in your stack in an afternoon" and "self-host in minutes." Self-hosting is the default posture, and air-gapped deployment is explicitly supported.

The site offers a free start with no credit card required, alongside a sales contact path. Pricing details for paid tiers are not specified in the material available here, so check the pricing page directly before assuming what a production deployment costs.

When this architecture fits

The gateway model is a good match when:

  • Multiple teams or apps need model access and you don't want to distribute provider keys.
  • You expect to change models or providers and want that to be a config change, not a code migration.
  • You need per-team spend visibility and hard budget caps.
  • You want agents and MCP servers behind the same access layer as your LLMs.
  • You need to self-host, including in an air-gapped environment.

It's a weaker fit if a single app talks to a single provider and you have no governance or cost-attribution problem to solve — the gateway adds a hop and a component to operate for benefits you wouldn't use.

Getting started

  1. Confirm your interface. If your app already uses an OpenAI-compatible client, the integration is largely a base URL and key change.
  2. Decide hosting. Self-host in your own infrastructure, or start with the free option to evaluate before committing.
  3. Connect providers and secret management. The gateway works with the secret manager you already run, so plan how provider credentials will be sourced.
  4. Create virtual keys scoped by team, project, or app, and attach budgets.
  5. Verify with a single request that routing, spend attribution, and logging all register as expected before opening access broadly.

The main things to check before rolling out: the benchmark numbers against your own latency budget, and the current pricing terms for the tier you'd run in production.

litellm.ai
LiteLLM is the open-source AI gateway that puts your full AI stack behind one OpenAI-compatible key. Track and cap LLM spend, route to the right mode…