Call over 100 LLMs through a single OpenAI-compatible interface
Imagine moving away from managing API keys and request formats for multiple AI services separately, to calling all models through a single address. LiteLLM is an open-source AI gateway that unifies various LLM providers into an OpenAI-compatible format. You can write code directly using the Python SDK or deploy it as a proxy server shared by your entire team.
Combining a Rust core with a Python SDK
This repository provides a core written in Rust alongside a Python SDK. The Rust core handles performance, while the Python SDK offers a familiar interface for developers. The litellm package includes all features such as the proxy server, CLI, and dashboard, whereas litellm-core is a standalone distribution for environments that only need the SDK. Because the two packages share some Python files, they should not be installed in the same environment. You must have the Rust build tool uv and Git installed to build the core distribution directly.
Distinguishing between gateway and SDK roles
It is divided into two modes depending on the use case. The Python SDK is used when developers integrate directly into their codebase to call multiple LLMs. In this case, the router handles retry and fallback logic and returns OpenAI-compatible errors. The AI gateway (proxy server) is suitable when an organization or team builds a centralized service. It comes with access control via virtual keys, per-project cost tracking, guardrails, load balancing, and an admin dashboard by default. Companies like Netflix have adopted this open-source project.
Support for MCP and agent protocols
LiteLLM supports the MCP (Model Context Protocol) bridge and the A2A (Agent-to-Agent) protocol. Once you register an MCP server with the gateway, you can call MCP tools in OpenAI format through the /chat/completions endpoint. You can also connect directly to MCP servers from tools like Cursor IDE. The agent harness feature allows you to configure all model calls to pass through the AI gateway when running agents such as Claude Code, Codex, OpenCode, and DeepAgents. This ensures that agent activity is logged and tracked by the gateway.
Cloud deployment and security verification
Terraform modules for deploying to AWS and GCP are available in the public registry. AWS uses ECS Fargate, Aurora, ElastiCache, and ALB, while GCP uses Cloud Run, Cloud SQL, Memorystore, and HTTPS LB. Both stacks separate the gateway, backend, and UI into independent services and include managed Postgres, Redis, and object storage. Docker images are signed with cosign, and you can verify the signing key using a specific commit hash or release tag. It is recommended to use the -stable tagged Docker images, which have undergone stability testing, for production.
Performance figures and pre-adoption checks
These are benchmark results, and actual performance may vary depending on network conditions and model response times. Before adopting it, you should verify that your intended LLM provider is supported. While the support list is extensive, support for specific endpoints (e.g., /audio, /images) varies by provider. You must choose to install either litellm or litellm-core, not both. If Rust build tools are required, ensure that uv and the Rust toolchain are ready. The enterprise license includes SSO, professional support, and custom SLAs, and may differ in features from the open-source version.