LiteLLM alternatives fall into three groups: open-source gateways you run yourself, such as Portkey, Bifrost, agentgateway, Agent Router, Apache APISIX and Kong; hosted routers such as OpenRouter and Cloudflare AI Gateway; and projects to avoid for new work because they are archived or in maintenance mode. The right choice depends on runtime, licence tiers, budgets and where prompts may travel.
Why teams look beyond LiteLLM
LiteLLM is where many teams start. It is a Python SDK and proxy that exposes more than 100 LLM APIs in the OpenAI format, with virtual keys, spend tracking, budgets, rate limits, fallbacks, guardrails and an MCP gateway, and its repository had over 60,000 GitHub stars in October 2026. Code outside its enterprise/ directory is MIT licensed. The reasons teams review it come mostly from LiteLLM’s own documentation:
- Enterprise features sit under a commercial licence. LiteLLM aims its Enterprise tier at teams with 100 or more users or ten or more production use cases. Its docs list SSO beyond five users, JWT authentication, audit logs with retention, IP-based access lists, secret-manager integrations, guardrails per key or team, and tag-based budgets among the paid features.
- The operational footprint grows with scale. Virtual keys need a PostgreSQL database. LiteLLM’s production guidance asks for Redis 7.0 or newer as soon as more than one proxy instance runs, since each instance otherwise enforces limits on its own, and sizes each pod at 1 vCPU and 4 GiB of memory as a floor, with one Uvicorn worker per pod.
- The March 2026 supply-chain incident. On 24 March 2026, malicious versions 1.82.7 and 1.82.8 were published to PyPI after a compromise that LiteLLM traced to the Trivy dependency in its CI/CD security scanning. LiteLLM says the packages were live for about 40 minutes, that unpinned pip installs between 10:39 and 16:00 UTC that day were affected, and that the official proxy Docker image was not. It released a clean v1.83.0 through a rebuilt pipeline and signs Docker images with cosign from v1.83.0-nightly.
- Python overhead, at least for now. LiteLLM’s own benchmark of the forwarding path measured about 7.5 ms of overhead in Python against about 0.05 ms on its new Rust path. It is migrating route by route and targets a full Rust server by 1 December 2026, with config.yaml, the database schema and the client API unchanged. Performance on its own may not stay a reason to switch for long.
What to compare before switching
A replacement should be judged on the work LiteLLM is doing for you today, not on a feature grid. Check these points against your own config.yaml and database:
- Runtime and footprint. Language, memory, the databases and caches it needs, and how it scales horizontally.
- Data path. Whether it runs on your infrastructure, in your cloud account, or only as a hosted service.
- Routing. Retries, fallbacks, load-balancing strategies and per-request overrides.
- Keys and budgets. Virtual keys, budgets per team or user, token and cost limits, and whether those sit in the free edition.
- Guardrails. Built-in checks, partner integrations, and whether they block a request or only report it.
- Identity and audit. Single sign-on, role-based access and audit logs, and their tier.
- Project health. Release cadence, maintainers and recent commits, which matter more than usual this year.
For a wider view of the gateway market beyond the LiteLLM switch, see our broader AI gateway comparison.
LiteLLM alternatives at a glance
Rows come from each project’s own documentation or repository as of October 2026, and dashes mark items those pages did not document.
| Gateway | Runtime | Self-host | Routing and fallbacks | Budgets and limits | Guardrails | Licence and paid tier |
|---|---|---|---|---|---|---|
| LiteLLM, for reference | Python, moving to Rust | Yes | Retries, fallbacks, load balancing | Virtual keys, budgets, rate limits | Yes; per key or team in Enterprise | MIT; enterprise/ directory commercial |
| Portkey AI Gateway | TypeScript on Node.js | Yes; also Cloudflare Workers | Retries, fallbacks, load balancing, conditional routing | Hosted and enterprise versions | 40+ in the open-source build | MIT; hosted and enterprise editions |
| Bifrost | Go | Yes | Failover and load balancing; adaptive balancing in Enterprise | Virtual keys with budgets and rate limits | Enterprise | Apache 2.0; Enterprise edition |
| agentgateway | Rust | Yes | Load balancing and failover | Budget and spend controls | Guardrails on MCP traffic | Apache 2.0, Linux Foundation |
| Agent Router | Go, on Envoy | Yes; or hosted by Tetrate | Provider fallback, model name virtualisation | Token limits per team, app or model | – | Apache 2.0 |
| Apache APISIX | Lua | Yes | Weighted, hashed or semantic balancing; fallback on 429, 5xx or quota | Token limits per instance or consumer | Prompt guard, Lakera plugin | Apache 2.0 |
| Kong AI Gateway | Lua | Yes; or Kong Konnect | Load balancing and failover in a paid plugin | Token and cost limits in a paid plugin | Prompt guards, PII sanitiser, cloud guardrails | Apache 2.0 core; AI Gateway Enterprise |
| OpenRouter | Hosted | No | Price- and uptime-weighted routing with fallbacks | Budgets per key and per user | Regex prompt-injection detection, PII redaction | Hosted; fee on credit purchases |
| Cloudflare AI Gateway | Hosted | No | Dynamic routing with fallback, Auto Router | Spend limits, rate limiting | Guardrails, data loss prevention | Hosted; core features free |
| Helicone AI Gateway | Rust | Yes | Latency, cost and weighted balancing; fallbacks | Rate limits per user, team or globally | – | GPL-3.0; repository idle since November 2025 |
| TensorZero | Rust | Yes | Routing, retries, fallbacks | Rate limits and budgets | – | Apache 2.0; archived in June 2026 |
Open-source gateways you run yourself
Portkey AI Gateway
Portkey’s gateway is TypeScript under the MIT licence and starts with npx @portkey-ai/gateway, Docker or a Node.js server, and it can also run on Cloudflare Workers. The open-source build includes retries of up to five attempts with exponential backoff, fallbacks, load balancing, conditional routing, simple caching and more than 40 guardrails. Budgets, rate limits, semantic caching, logging and tracing are listed for the hosted and enterprise versions. The README announces Gateway 2.0, a pre-release that merges Portkey’s core enterprise gateway into the open-source project, while the repository’s most recent push was in May 2026, so check release activity before committing. Coming from LiteLLM, the trade is Python for Node.js, with spend controls on the paid side for now.
Bifrost
Bifrost, from Maxim AI, is a Go gateway under Apache 2.0 with one OpenAI-compatible API across more than 20 providers. It starts with npx -y @maximhq/bifrost, Docker or as a Go SDK. The open-source edition includes failover, load balancing, semantic caching, virtual keys that carry budgets, rate limits and routing per consumer, an MCP gateway, Prometheus metrics, OpenTelemetry and plugins in Go or WASM. Clustering, adaptive load balancing, content-safety guardrails, Okta and Entra sign-in with role-based access, audit logs and in-VPC deployment belong to the Enterprise edition. The repository advertises “50x faster than LiteLLM” and the docs cite 11 µs of overhead at 5,000 requests per second. Those are the vendor’s numbers, measured against LiteLLM’s Python path, so rerun them on your own traffic.
agentgateway
agentgateway is a Rust proxy that Solo.io donated to the Linux Foundation, licensed Apache 2.0. A single binary, or a Kubernetes deployment with its own controller, handles model, MCP and agent-to-agent traffic. For models it offers an OpenAI-compatible API to the major providers with budget and spend controls, prompt enrichment, load balancing and failover, next to JWT, API key and OAuth authentication, CEL-based authorisation, local and global rate limits and OpenTelemetry output. Of the Rust gateways checked, it is the one still under active development, with a repository push on 2 October 2026, and it suits teams that want one proxy for both model and tool traffic.
Agent Router, formerly Envoy AI Gateway
Envoy AI Gateway is now Agent Router, an Agentic AI Foundation project with the same code and the same maintainers. Written in Go on Envoy Gateway under Apache 2.0, it offers one OpenAI-compatible API across 17 providers and self-hosted vLLM, with provider fallback, model name virtualisation, token limits per team, app or model, centralised provider credentials, OpenTelemetry GenAI conventions and an MCP tool catalogue filtered by caller. It runs on a laptop, on Kubernetes, or as Tetrate’s hosted service, and it is the obvious candidate for teams that already operate Envoy.
Apache APISIX AI plugins
APISIX is an Apache 2.0 API gateway written in Lua that handles model traffic through plugins. ai-proxy-multi spreads requests across model instances by weighted round robin, consistent hashing or semantic similarity, honours priorities, retries, runs active health checks and falls back when an instance is rate limited or returns 429 or 5xx errors. Its providers include OpenAI, Anthropic, Azure OpenAI, Gemini, Vertex AI, Amazon Bedrock, DeepSeek, OpenRouter and any OpenAI-compatible endpoint. ai-rate-limiting caps total, prompt or completion tokens per instance or per consumer and can push overflow to an instance with quota left, and ai-prompt-guard, ai-lakera-guard, ai-cache and ai-rag complete the set. All of it is open source; you assemble and operate it yourself.
Kong AI Gateway
Kong Gateway’s core is Apache 2.0, and its AI Proxy plugin, available from Gateway 3.6, reaches commercial providers as well as self-hosted Ollama, vLLM and Hugging Face models. The features LiteLLM users lean on most are paid. AI Proxy Advanced, which balances load by latency, usage, semantic match or priority and fails over between models, and AI Rate Limiting Advanced, which limits prompt, completion or total tokens or cost per consumer, model or provider, both belong to AI Gateway Enterprise. Kong makes most sense where it already fronts your APIs.
Hosted routers and managed gateways
OpenRouter
OpenRouter is a hosted API in front of many models and providers. It passes provider prices through without markup and charges 5.5% on card credit purchases with a $0.80 minimum, or 5% for crypto; bring-your-own-key usage is free up to $25,000 a month of list-price inference on pay-as-you-go plans and $200,000 on Enterprise, then costs 5% (list, verified October 2026). By default it prefers providers without recent outages and weights the choice toward the lowest price, keeping the rest as fallbacks; requests can instead sort by price, throughput or latency, exclude providers, cap the price, or require zero-data-retention endpoints. Organisation guardrails add budgets per user and per key that reset daily, weekly or monthly, model and provider allowlists, regex prompt-injection detection and PII redaction. Prompts and completions are not logged unless you opt in. It replaces LiteLLM’s provider reach with nothing to run, at the price of a third party in the data path.
Cloudflare AI Gateway
Cloudflare’s gateway runs on its own network. Its documentation, updated 30 September 2026, lists caching, rate limiting, spend limits, dynamic routing with automatic fallback, an Auto Router, guardrails, data loss prevention, authentication, bring-your-own keys, logging and custom costs. Analytics, caching and rate limiting are free, Unified Billing adds a 5% fee on purchased credits, and log storage limits depend on the plan. It suits teams already on Cloudflare who accept a managed data path.
Projects to avoid for new deployments
Two names still appear in older “LiteLLM alternatives” lists but need care in October 2026:
- TensorZero. The Rust gateway and LLMOps platform is archived on GitHub, and its website says the project remains available but is no longer maintained.
- Helicone AI Gateway. Mintlify acquired Helicone on 3 March 2026 and runs it in maintenance mode, with security updates, bug fixes and new models still shipping and help for customers who migrate elsewhere. The Rust gateway repository, licensed GPL-3.0, was last pushed in November 2025.
That leaves agentgateway as the active Rust option today, alongside LiteLLM’s own Rust path once it ships.
Planning the move off LiteLLM
- Export what LiteLLM holds. That means the model list, fallbacks and router settings in
config.yaml, the virtual keys and budgets in PostgreSQL, guardrail settings and logging callbacks. - Map each item to the target. Note which pieces need a paid tier there, and which have no equivalent.
- Check client compatibility. Bifrost, agentgateway and Agent Router document OpenAI-compatible APIs; test the endpoints, streaming and tool-calling behaviour your applications actually use.
- Run both side by side. Send a share of traffic through the new gateway and compare latency, error rates, token counts and cost per request. Our explainer on how LLM routing works helps when routing rules differ.
- Be careful with caching. If the new gateway turns on semantic caching, review the semantic caching risks for regulated workloads first.
- Cut over one workload at a time. Rotate provider keys as you go, and keep LiteLLM pinned to a verified version and image digest until the last workload has moved.
- Recheck the cost case. Model what routing and limits will save on real traffic with the AI savings calculator before signing an enterprise tier.
How VDF AI fits
VDF AI’s gateway is VDF AI Router, a Docker-packaged Python service that runs on your VMs, Kubernetes or bare metal and exposes a REST API and a Python SDK. For model access it uses Ollama locally and OpenRouter for cloud models. Policy is the first layer of every decision: regulated domains are limited to approved models, workloads can be pinned to a named model, organisation-wide allow and deny lists apply before any call, and air-gap mode disables external APIs entirely.
Each decision returns the chosen model, a readable reason, up to five ordered failover candidates and per-model scores, and the SEEMR engine learns from observed quality, latency, failures and energy. Rate limits and budgets apply in the request path. The product pages do not list semantic caching, PII redaction or prompt-injection filtering, and the Router documents a REST API and Python SDK rather than an OpenAI-compatible endpoint, so plan for client changes. Tool traffic is governed separately by the MCP gateway, and the on-premises AI gateway overview explains how the two fit together.
Sources
- LiteLLM repository
- LiteLLM licence
- LiteLLM Enterprise
- LiteLLM virtual keys
- LiteLLM production best practices
- LiteLLM security update, March 2026
- LiteLLM migration to Rust
- Portkey AI Gateway repository
- Bifrost repository
- Bifrost documentation
- agentgateway repository
- agentgateway project site
- Agent Router
- Agent Router repository
- Apache APISIX AI gateway
- APISIX ai-proxy-multi plugin
- APISIX ai-rate-limiting plugin
- Kong AI Gateway
- Kong AI Proxy plugin
- Kong AI Proxy Advanced plugin
- Kong AI Rate Limiting Advanced plugin
- Kong Gateway repository
- OpenRouter FAQ
- OpenRouter provider routing
- OpenRouter guardrails
- Cloudflare AI Gateway features
- Cloudflare AI Gateway pricing
- Helicone AI Gateway repository
- Mintlify acquires Helicone
- TensorZero repository
- TensorZero website