AI Infrastructure

LiteLLM Alternatives in 2026: Open-Source, Self-Hosted and Enterprise LLM Gateways

For teams already running LiteLLM: what its own documentation says about enterprise licensing, operations, performance and the March 2026 PyPI incident, and how Portkey, Bifrost, agentgateway, Agent Router, Apache APISIX, Kong, OpenRouter and Cloudflare compare. Verified October 2026.

LiteLLM alternatives fall into three groups: open-source gateways you run yourself, such as Portkey, Bifrost, agentgateway, Agent Router, Apache APISIX and Kong; hosted routers such as OpenRouter and Cloudflare AI Gateway; and projects to avoid for new work because they are archived or in maintenance mode. The right choice depends on runtime, licence tiers, budgets and where prompts may travel.

Why teams look beyond LiteLLM

LiteLLM is where many teams start. It is a Python SDK and proxy that exposes more than 100 LLM APIs in the OpenAI format, with virtual keys, spend tracking, budgets, rate limits, fallbacks, guardrails and an MCP gateway, and its repository had over 60,000 GitHub stars in October 2026. Code outside its enterprise/ directory is MIT licensed. The reasons teams review it come mostly from LiteLLM’s own documentation:

  1. Enterprise features sit under a commercial licence. LiteLLM aims its Enterprise tier at teams with 100 or more users or ten or more production use cases. Its docs list SSO beyond five users, JWT authentication, audit logs with retention, IP-based access lists, secret-manager integrations, guardrails per key or team, and tag-based budgets among the paid features.
  2. The operational footprint grows with scale. Virtual keys need a PostgreSQL database. LiteLLM’s production guidance asks for Redis 7.0 or newer as soon as more than one proxy instance runs, since each instance otherwise enforces limits on its own, and sizes each pod at 1 vCPU and 4 GiB of memory as a floor, with one Uvicorn worker per pod.
  3. The March 2026 supply-chain incident. On 24 March 2026, malicious versions 1.82.7 and 1.82.8 were published to PyPI after a compromise that LiteLLM traced to the Trivy dependency in its CI/CD security scanning. LiteLLM says the packages were live for about 40 minutes, that unpinned pip installs between 10:39 and 16:00 UTC that day were affected, and that the official proxy Docker image was not. It released a clean v1.83.0 through a rebuilt pipeline and signs Docker images with cosign from v1.83.0-nightly.
  4. Python overhead, at least for now. LiteLLM’s own benchmark of the forwarding path measured about 7.5 ms of overhead in Python against about 0.05 ms on its new Rust path. It is migrating route by route and targets a full Rust server by 1 December 2026, with config.yaml, the database schema and the client API unchanged. Performance on its own may not stay a reason to switch for long.

What to compare before switching

A replacement should be judged on the work LiteLLM is doing for you today, not on a feature grid. Check these points against your own config.yaml and database:

  • Runtime and footprint. Language, memory, the databases and caches it needs, and how it scales horizontally.
  • Data path. Whether it runs on your infrastructure, in your cloud account, or only as a hosted service.
  • Routing. Retries, fallbacks, load-balancing strategies and per-request overrides.
  • Keys and budgets. Virtual keys, budgets per team or user, token and cost limits, and whether those sit in the free edition.
  • Guardrails. Built-in checks, partner integrations, and whether they block a request or only report it.
  • Identity and audit. Single sign-on, role-based access and audit logs, and their tier.
  • Project health. Release cadence, maintainers and recent commits, which matter more than usual this year.

For a wider view of the gateway market beyond the LiteLLM switch, see our broader AI gateway comparison.

LiteLLM alternatives at a glance

Rows come from each project’s own documentation or repository as of October 2026, and dashes mark items those pages did not document.

GatewayRuntimeSelf-hostRouting and fallbacksBudgets and limitsGuardrailsLicence and paid tier
LiteLLM, for referencePython, moving to RustYesRetries, fallbacks, load balancingVirtual keys, budgets, rate limitsYes; per key or team in EnterpriseMIT; enterprise/ directory commercial
Portkey AI GatewayTypeScript on Node.jsYes; also Cloudflare WorkersRetries, fallbacks, load balancing, conditional routingHosted and enterprise versions40+ in the open-source buildMIT; hosted and enterprise editions
BifrostGoYesFailover and load balancing; adaptive balancing in EnterpriseVirtual keys with budgets and rate limitsEnterpriseApache 2.0; Enterprise edition
agentgatewayRustYesLoad balancing and failoverBudget and spend controlsGuardrails on MCP trafficApache 2.0, Linux Foundation
Agent RouterGo, on EnvoyYes; or hosted by TetrateProvider fallback, model name virtualisationToken limits per team, app or model–Apache 2.0
Apache APISIXLuaYesWeighted, hashed or semantic balancing; fallback on 429, 5xx or quotaToken limits per instance or consumerPrompt guard, Lakera pluginApache 2.0
Kong AI GatewayLuaYes; or Kong KonnectLoad balancing and failover in a paid pluginToken and cost limits in a paid pluginPrompt guards, PII sanitiser, cloud guardrailsApache 2.0 core; AI Gateway Enterprise
OpenRouterHostedNoPrice- and uptime-weighted routing with fallbacksBudgets per key and per userRegex prompt-injection detection, PII redactionHosted; fee on credit purchases
Cloudflare AI GatewayHostedNoDynamic routing with fallback, Auto RouterSpend limits, rate limitingGuardrails, data loss preventionHosted; core features free
Helicone AI GatewayRustYesLatency, cost and weighted balancing; fallbacksRate limits per user, team or globally–GPL-3.0; repository idle since November 2025
TensorZeroRustYesRouting, retries, fallbacksRate limits and budgets–Apache 2.0; archived in June 2026

Open-source gateways you run yourself

Portkey AI Gateway

Portkey’s gateway is TypeScript under the MIT licence and starts with npx @portkey-ai/gateway, Docker or a Node.js server, and it can also run on Cloudflare Workers. The open-source build includes retries of up to five attempts with exponential backoff, fallbacks, load balancing, conditional routing, simple caching and more than 40 guardrails. Budgets, rate limits, semantic caching, logging and tracing are listed for the hosted and enterprise versions. The README announces Gateway 2.0, a pre-release that merges Portkey’s core enterprise gateway into the open-source project, while the repository’s most recent push was in May 2026, so check release activity before committing. Coming from LiteLLM, the trade is Python for Node.js, with spend controls on the paid side for now.

Bifrost

Bifrost, from Maxim AI, is a Go gateway under Apache 2.0 with one OpenAI-compatible API across more than 20 providers. It starts with npx -y @maximhq/bifrost, Docker or as a Go SDK. The open-source edition includes failover, load balancing, semantic caching, virtual keys that carry budgets, rate limits and routing per consumer, an MCP gateway, Prometheus metrics, OpenTelemetry and plugins in Go or WASM. Clustering, adaptive load balancing, content-safety guardrails, Okta and Entra sign-in with role-based access, audit logs and in-VPC deployment belong to the Enterprise edition. The repository advertises “50x faster than LiteLLM” and the docs cite 11 µs of overhead at 5,000 requests per second. Those are the vendor’s numbers, measured against LiteLLM’s Python path, so rerun them on your own traffic.

agentgateway

agentgateway is a Rust proxy that Solo.io donated to the Linux Foundation, licensed Apache 2.0. A single binary, or a Kubernetes deployment with its own controller, handles model, MCP and agent-to-agent traffic. For models it offers an OpenAI-compatible API to the major providers with budget and spend controls, prompt enrichment, load balancing and failover, next to JWT, API key and OAuth authentication, CEL-based authorisation, local and global rate limits and OpenTelemetry output. Of the Rust gateways checked, it is the one still under active development, with a repository push on 2 October 2026, and it suits teams that want one proxy for both model and tool traffic.

Agent Router, formerly Envoy AI Gateway

Envoy AI Gateway is now Agent Router, an Agentic AI Foundation project with the same code and the same maintainers. Written in Go on Envoy Gateway under Apache 2.0, it offers one OpenAI-compatible API across 17 providers and self-hosted vLLM, with provider fallback, model name virtualisation, token limits per team, app or model, centralised provider credentials, OpenTelemetry GenAI conventions and an MCP tool catalogue filtered by caller. It runs on a laptop, on Kubernetes, or as Tetrate’s hosted service, and it is the obvious candidate for teams that already operate Envoy.

Apache APISIX AI plugins

APISIX is an Apache 2.0 API gateway written in Lua that handles model traffic through plugins. ai-proxy-multi spreads requests across model instances by weighted round robin, consistent hashing or semantic similarity, honours priorities, retries, runs active health checks and falls back when an instance is rate limited or returns 429 or 5xx errors. Its providers include OpenAI, Anthropic, Azure OpenAI, Gemini, Vertex AI, Amazon Bedrock, DeepSeek, OpenRouter and any OpenAI-compatible endpoint. ai-rate-limiting caps total, prompt or completion tokens per instance or per consumer and can push overflow to an instance with quota left, and ai-prompt-guard, ai-lakera-guard, ai-cache and ai-rag complete the set. All of it is open source; you assemble and operate it yourself.

Kong AI Gateway

Kong Gateway’s core is Apache 2.0, and its AI Proxy plugin, available from Gateway 3.6, reaches commercial providers as well as self-hosted Ollama, vLLM and Hugging Face models. The features LiteLLM users lean on most are paid. AI Proxy Advanced, which balances load by latency, usage, semantic match or priority and fails over between models, and AI Rate Limiting Advanced, which limits prompt, completion or total tokens or cost per consumer, model or provider, both belong to AI Gateway Enterprise. Kong makes most sense where it already fronts your APIs.

Hosted routers and managed gateways

OpenRouter

OpenRouter is a hosted API in front of many models and providers. It passes provider prices through without markup and charges 5.5% on card credit purchases with a $0.80 minimum, or 5% for crypto; bring-your-own-key usage is free up to $25,000 a month of list-price inference on pay-as-you-go plans and $200,000 on Enterprise, then costs 5% (list, verified October 2026). By default it prefers providers without recent outages and weights the choice toward the lowest price, keeping the rest as fallbacks; requests can instead sort by price, throughput or latency, exclude providers, cap the price, or require zero-data-retention endpoints. Organisation guardrails add budgets per user and per key that reset daily, weekly or monthly, model and provider allowlists, regex prompt-injection detection and PII redaction. Prompts and completions are not logged unless you opt in. It replaces LiteLLM’s provider reach with nothing to run, at the price of a third party in the data path.

Cloudflare AI Gateway

Cloudflare’s gateway runs on its own network. Its documentation, updated 30 September 2026, lists caching, rate limiting, spend limits, dynamic routing with automatic fallback, an Auto Router, guardrails, data loss prevention, authentication, bring-your-own keys, logging and custom costs. Analytics, caching and rate limiting are free, Unified Billing adds a 5% fee on purchased credits, and log storage limits depend on the plan. It suits teams already on Cloudflare who accept a managed data path.

Projects to avoid for new deployments

Two names still appear in older “LiteLLM alternatives” lists but need care in October 2026:

  • TensorZero. The Rust gateway and LLMOps platform is archived on GitHub, and its website says the project remains available but is no longer maintained.
  • Helicone AI Gateway. Mintlify acquired Helicone on 3 March 2026 and runs it in maintenance mode, with security updates, bug fixes and new models still shipping and help for customers who migrate elsewhere. The Rust gateway repository, licensed GPL-3.0, was last pushed in November 2025.

That leaves agentgateway as the active Rust option today, alongside LiteLLM’s own Rust path once it ships.

Planning the move off LiteLLM

  1. Export what LiteLLM holds. That means the model list, fallbacks and router settings in config.yaml, the virtual keys and budgets in PostgreSQL, guardrail settings and logging callbacks.
  2. Map each item to the target. Note which pieces need a paid tier there, and which have no equivalent.
  3. Check client compatibility. Bifrost, agentgateway and Agent Router document OpenAI-compatible APIs; test the endpoints, streaming and tool-calling behaviour your applications actually use.
  4. Run both side by side. Send a share of traffic through the new gateway and compare latency, error rates, token counts and cost per request. Our explainer on how LLM routing works helps when routing rules differ.
  5. Be careful with caching. If the new gateway turns on semantic caching, review the semantic caching risks for regulated workloads first.
  6. Cut over one workload at a time. Rotate provider keys as you go, and keep LiteLLM pinned to a verified version and image digest until the last workload has moved.
  7. Recheck the cost case. Model what routing and limits will save on real traffic with the AI savings calculator before signing an enterprise tier.

How VDF AI fits

VDF AI’s gateway is VDF AI Router, a Docker-packaged Python service that runs on your VMs, Kubernetes or bare metal and exposes a REST API and a Python SDK. For model access it uses Ollama locally and OpenRouter for cloud models. Policy is the first layer of every decision: regulated domains are limited to approved models, workloads can be pinned to a named model, organisation-wide allow and deny lists apply before any call, and air-gap mode disables external APIs entirely.

Each decision returns the chosen model, a readable reason, up to five ordered failover candidates and per-model scores, and the SEEMR engine learns from observed quality, latency, failures and energy. Rate limits and budgets apply in the request path. The product pages do not list semantic caching, PII redaction or prompt-injection filtering, and the Router documents a REST API and Python SDK rather than an OpenAI-compatible endpoint, so plan for client changes. Tool traffic is governed separately by the MCP gateway, and the on-premises AI gateway overview explains how the two fit together.

Sources

Frequently asked questions

What are the best open-source alternatives to LiteLLM?

It depends on the runtime and the tiers you can accept. Portkey's gateway is MIT-licensed TypeScript with fallbacks, retries and guardrails, but keeps budgets and rate limits in its hosted and enterprise versions. Bifrost is Apache 2.0 Go with virtual keys, budgets and semantic caching in the open-source edition. agentgateway is Apache 2.0 Rust under the Linux Foundation. Agent Router, formerly Envoy AI Gateway, suits teams that run Envoy, and Apache APISIX keeps its AI plugins open source. Kong's core is open source, but its AI load balancing and token limits are paid.

Is there a LiteLLM alternative written in Rust or Go?

In Go, Bifrost from Maxim AI and Agent Router, formerly Envoy AI Gateway, are both Apache 2.0 and under active development. In Rust, agentgateway is the active option. TensorZero's repository was archived in June 2026, and Helicone's Rust gateway has had no repository pushes since November 2025 while Helicone runs in maintenance mode. LiteLLM itself is moving its hot paths to Rust, targeting a full Rust server for 1 December 2026 without changes to config.yaml or the client API.

Is LiteLLM still safe to use after the March 2026 PyPI compromise?

LiteLLM says malicious versions 1.82.7 and 1.82.8 reached PyPI on 24 March 2026 after a compromise that started in the Trivy dependency of its CI/CD workflow, and were quarantined after about 40 minutes. Unpinned pip installs during that day's window were exposed; the official proxy Docker image was not. LiteLLM rotated credentials, rebuilt its pipeline, released v1.83.0 and now signs Docker images with cosign. Whichever gateway you run, pin versions, verify image signatures and keep provider keys out of build environments.

What does LiteLLM Enterprise add over the open-source version?

LiteLLM's documentation aims Enterprise at teams with 100 or more users or ten or more production AI use cases. The paid features include SSO beyond five users, JWT authentication, audit logs with retention policies, role-based access for organisations and teams, IP-based access lists, key rotation, secret-manager integrations, tag-based and model-specific budgets, per-team logging controls and guardrails per key or team. Pricing depends on deployment size and comes from LiteLLM's sales team.

Should we use OpenRouter instead of a self-hosted gateway?

OpenRouter removes provider integration work with nothing to operate: one API, provider fallbacks, budgets per key and per user, and an option to route only to zero-data-retention endpoints. It passes model prices through without markup but charges a fee on credit purchases, and on bring-your-own-key usage beyond a monthly allowance. The trade-off is the data path, because every prompt passes through OpenRouter and the chosen provider, so it cannot serve air-gapped or strictly on-premises workloads. A self-hosted gateway can still use it as one upstream for approved cloud models.

How do we migrate from LiteLLM to another gateway?

Export the LiteLLM config.yaml, virtual keys, budgets, guardrails and logging callbacks, map each to the new gateway, and note which pieces need a paid tier there. Confirm the new gateway speaks the API format your clients use, run it beside LiteLLM on a slice of traffic, and compare latency, errors, token counts and cost per request. Cut over one workload at a time, rotate keys as you go, and keep LiteLLM pinned to a verified version until the last workload has moved.

Filed under
AI gatewayLLM gatewaymodel routingAI infrastructureopen-source AIon-premises AI
AI Agent Platform

Evaluating AI agent platforms?

See what a governed agent platform looks like in production — build, orchestration, retrieval, routing and audit in one product — then score it against the platforms on your shortlist.

Or start free — no credit card →

Keep reading