What is LiteLLM and what does it do?
LiteLLM is an open-source AI gateway and LLM proxy that puts your full AI stack — models, agents, and MCP servers — behind one OpenAI-compatible API and one login. It is built for platform teams that need to give an entire organization access to many models while keeping visibility and control over every request. You would use it when you want to swap or add models without changing application code, track and cap LLM spend, and self-host the gateway in your own infrastructure, including air-gapped environments.
The core idea: one key, many models
Instead of wiring each application to each provider separately, you point your apps at LiteLLM and let it handle the connections. The site describes this as putting your full AI stack behind one key, with one OpenAI-compatible API reaching 140+ providers and 1,800+ models (the source description cites 1,892 models). Because the interface is OpenAI-compatible, existing code that already talks to OpenAI-style endpoints can generally be redirected to the gateway rather than rewritten.
A practical consequence: if you decide to switch the backend model, it becomes a configuration update in the gateway rather than a code change. Okta's Dennis Henry is quoted making exactly this point — no code changes, procurement cycles, or repetitive security reviews required.
What it actually does
The site groups the gateway's value into four areas, which map to concrete jobs:
- Access — Give the whole company every model, agent, and MCP behind one API and one login. LiteLLM handles key management and works with the secret manager you already run, so platform teams can open access without becoming a bottleneck. Access can be scoped by team, project, or app, and it supports SSO.
- Routing — Send each request to the model that should handle it. New models get day-zero support, so developers can use the latest model the day it ships.
- Observability — See who is driving spend and what is happening across requests.
- Governance and spend control — Cap spend before it runs, and keep your keys, your infrastructure, and your audit trail.
The site's own dashboard example shows this in practice: virtual keys listed with their team, last-active time, and spend against budget — for example, a platform-eng key at $4,182.55 / $10,000 and a data-science key at $6,740.15. That is the spend-tracking and budgeting layer in action.
Beyond LLMs: agents and MCP
LiteLLM is not limited to chat models. The same gateway reaches agents and MCP servers, not just LLMs, and it can also put your own internal, fine-tuned, and self-hosted models behind the same key. If your stack mixes hosted providers with in-house models, this keeps a single access point rather than one integration per source.
Deployment and overhead
The site positions deployment as fast — "go live in your stack in an afternoon" — and claims sub-millisecond overhead with a published benchmark. It supports self-hosting in minutes and can run air-gapped, which matters for teams that cannot send traffic to an external service. The framing "your keys, your infra, your audit trail" reflects that the gateway runs in your environment rather than as a mandatory third party.
Who it is for
The primary audience is platform teams at organizations shipping AI at scale. The site cites NVIDIA (a single consistent way to access more than 100 AI model endpoints), Netflix (providing the latest models to users, usually within a day of release, saving months of work), Lemonade, Okta, and AT&T. AT&T's Mark Austin is quoted saying that after adopting LiteLLM as a router provider, costs for some advanced AI tasks such as coding fell by as much as 56% — a routing-and-cost outcome rather than a model-quality claim.
How to decide whether to try it
LiteLLM fits if you recognize these conditions:
- You have multiple apps and multiple model providers and want one integration point.
- You need per-team or per-project visibility and budget caps on LLM spend.
- You want to change models through configuration, not code.
- You must self-host, or even run air-gapped.
- You want agents and MCP servers reachable through the same gateway as your LLMs.
It is less obviously necessary if you use a single provider with a single app and have no governance or spend-visibility requirements.
To start, the site offers a free start path with no credit card and a "talk to sales" option for larger rollouts. Note that the pricing page is linked as "Get Started," and the source does not state specific plan prices or whether paid tiers are required for particular features — check the pricing page directly before assuming any capability is free.