Every team that adopts LLMs hits the same wall within months: a dozen provider SDKs, a dozen auth styles, a dozen error taxonomies, and no idea what any of it costs. LiteLLM, at version 1.106.0 and backed by Y Combinator, attacks that wall from both sides. The Python SDK exposes one OpenAI-shaped completion function across more than one hundred providers, and the AI Gateway, a FastAPI proxy the README positions as self-hosted and enterprise-ready, wraps that SDK with the operational layer real organizations need: virtual API keys, per-user and per-team budgets, spend tracking, rate limits, guardrails, load balancing, fallbacks, and an admin dashboard. The README quotes 8ms P95 overhead at 1k requests per second and lists adopters from Stripe to Netflix, and the repository backs the claim with an unusual artifact: an ARCHITECTURE.md file that maps the full request flow, the Redis and PostgreSQL division of labor, and even the background jobs, written for the contributors who maintain this codebase daily.

The layout separates the two products cleanly. The litellm package holds the SDK core and the gateway: litellm_core_utils for logging and cost plumbing, router.py for the load balancer, caching for the response cache stack, and llms, a directory of per-provider packages where each vendor gets its own transformation code. The proxy directory is the gateway itself, from proxy_server.py through auth, hooks, management endpoints, and the Prisma schema it runs on. Around them live litellm-rust for the Rust side, litellm-proxy-extras, the React dashboard under ui, the enterprise directory with features under a separate commercial license, plus helm charts, terraform modules, docker files, and migrations. Licensing is explicit in the LICENSE file: everything outside enterprise is MIT, and the content under enterprise follows its own commercial terms. As always in this series, what follows is an educational tour of published source code.

LiteLLM overview architecture diagram

LiteLLM at a glance: the FastAPI gateway mounts OpenAI and Anthropic-shaped routes, user_api_key_auth verifies keys against the DualCache and enforces hook-level rate limits, the Router picks a deployment with its latency strategies, the SDK dispatches through the shared HTTP transport into per-provider transformations, and cost lands in batched spend logs.

Reading the overview from left to right:

Why You Need This

The first reason is that the gateway is a real product, not a thin reverse proxy. The auth gate at litellm/proxy/auth/user_api_key_auth.py resolves each virtual key against the InternalUsageCache, checks team and user budgets, and stamps the request with metadata that follows it through the whole call. The limiter at litellm/proxy/hooks/parallel_request_limiter_v3.py tracks TPM and RPM counters per key, user, and team, and the management endpoints under litellm/proxy/management_endpoints expose key generation, team membership, and model access as a proper admin API. Behind all of it sits PostgreSQL through the Prisma schema at litellm/proxy/schema.prisma, with token, team, user, and spend-log tables. This is the layer you would otherwise spend a quarter building.

The second reason is the routing engine. The Router at litellm/router.py manages a list of model deployments with retries, timeouts, cooldowns for failing providers, and fallback chains, and its strategy modules under litellm/router_strategy include latency-aware selection in litellm/router_strategy/lowest_latency.py that scores deployments from observed response times. Router state, cooldowns and TPM tracking included, lives in its own DualCache so multiple gateway instances share it through Redis. On top of that, the response cache stack, the caching handler at litellm/caching/caching_handler.py over the Cache class in litellm/caching/caching.py, lets identical prompts return instantly when you enable it. Few open-source projects make the multi-provider, multi-replica problem this concrete.

The third reason is the provider abstraction, which is the quiet engineering marvel here. Providers are not hand-written clients; they are configuration classes that declare how to transform OpenAI-shaped requests into vendor payloads and back, executed by the shared BaseLLMHTTPHandler at litellm/llms/custom_httpx/llm_http_handler.py. A representative example is the Anthropic chat transformation at litellm/llms/anthropic/chat/transformation.py, and AWS Bedrock gets a whole package under litellm/llms/bedrock with its signature-based auth. Pricing for cost attribution comes from the model price map at model_prices_and_context_window.json, a curated JSON of per-token prices and context windows that the cost calculator consults. Adding a provider means adding a transformation, not a client library.

LiteLLM detail architecture diagram

The detail view: the gateway surface with OpenAI, Anthropic, passthrough, A2A, realtime, MCP, and management routes on the left; auth, pre-call metadata, hooks, and the rate limiter beside them; the Router with strategies, the SDK entry, batching, response caching, and logging in the middle; provider packages, transformations, and shared types on the right; and the DualCache, Redis, entity models, repositories, spend writer, Prisma schema, price map, and dashboard below.

Walking the detail diagram, the surface area is the first surprise. Beyond the OpenAI-compatible families, the gateway mounts provider passthrough routes under litellm/passthrough, an A2A agent protocol module at litellm/a2a_protocol, a realtime websocket API under litellm/realtime_api, and an experimental MCP client under litellm/experimental_mcp_client, so agents, voice clients, and raw vendor calls all share one auth and billing story. The Anthropic Messages adapter converts formats through the interface layer at litellm/anthropic_interface, meaning an Anthropic SDK client can point at the gateway unmodified.

In the middle, the flow from auth to provider is worth tracing once. After the auth gate stamps user and team context through litellm/proxy/litellm_pre_call_utils.py, the route helper at litellm/proxy/route_llm_request.py resolves wildcard model patterns to a Router deployment, and the SDK entry at litellm/main.py attaches the logging object that fans out callbacks from litellm/litellm_core_utils/litellm_logging.py. Cost is computed from the response token counts, returned to the client in a response header, and handed to the proxy cost callback at litellm/proxy/hooks/proxy_track_cost_callback.py, which queues increments for the batched writer. Persistence is layered: typed entities in litellm/models flow through the repository layer at litellm/repositories into Prisma, while Redis carries the hot state, and the React dashboard under ui drives the management API.

From Install to Gateway

The README starts with the SDK: uv add litellm gives you the completion function over every provider, and a newer litellm-core distribution packages the same import surface with no optional extras, CLI, or bundled dashboard, built from this checkout with the script in scripts/build_core_distribution.py. For the gateway, the repository ships docker-compose files, a Dockerfile, helm charts, and terraform, because a stateful proxy with Redis and PostgreSQL is an infrastructure decision, not a pip install. Configuration is a YAML file declaring model list, database URL, and master key, and docker-compose.yml wires the proxy to its database for a local run. Once up, the dashboard under ui is served alongside the API, and the periodic jobs, from the sixty-second spend flush to budget resets and deployment syncs described in ARCHITECTURE.md, keep budgets honest without operator attention.

Honest limits: the feature surface is wide, and that width has a support bill, because the proxy pulls in database, Redis, and dashboard components that each need operational care; the enterprise directory holds a meaningful set of advanced features behind a commercial license, so audit exactly which capabilities your deployment needs before standardizing on the MIT core; and provider coverage, while enormous, moves with vendor API churn, which means version upgrades can carry behavioral changes. The ARCHITECTURE.md file mitigates all three honestly, and that document plus the MIT SDK make this one of the most deployable gateways in the ecosystem.

Watch PyShine on YouTube

Contents