· via dev.to (home feed)
LiteLLM unifies local and cloud LLM endpoints behind one self-hosted gateway
A dev.to walkthrough shows LiteLLM unifying Ollama, llama.cpp and cloud APIs behind one OpenAI-compatible proxy, with waterfall fallbacks, load balancing and cost tracking.

The problem: endpoint sprawl
According to a walkthrough on dev.to, homelab AI setups tend to accumulate incompatible inference services: Ollama on port 11434 serving models such as qwen2.5-coder:14b and deepseek-r1:14b, a llama.cpp server on port 8080 for specialist GGUF models, Text Generation WebUI on 7860 for experiments, and cloud APIs from OpenAI, Anthropic and Groq. Each one brings its own endpoint, credentials and SDK.
The author lists the consequences: every application ends up with conditional logic for each provider, credentials and endpoints multiply, a dead Ollama instance crashes apps because there is no fallback, load cannot be spread across instances, usage has to be logged by hand, and nothing protects the upstream APIs from abuse.
One proxy in front of everything
LiteLLM is presented as the fix: a single proxy on port 4000 exposing an OpenAI-compatible API. Applications point their existing OpenAI SDK at one URL, and the configuration decides what actually serves each model name — a request that looks like gpt-4o-mini can be routed to local hardware or a paid API, depending on what the operator chose.
The post says the gateway adds token counting and cost estimation, per-model or global rate limits, retries with exponential backoff and jitter, load balancing across identical instances, and automatic fallback when a provider fails.
Providers and routing strategies
The write-up splits supported providers into local ones — ollama, llama_cpp, vllm and Hugging Face's Text Generation Inference among them — and remote ones, including OpenAI, Anthropic, Groq, Cohere, Mistral, Azure OpenAI and Amazon Bedrock.
Three routing patterns are described:
- Simple list: models are tried in order until one responds. The example attempts Ollama's qwen2.5-coder:14b, then deepseek-r1:14b, then OpenAI's gpt-4o-mini.
- Load balancing: identical deployments registered under one name share incoming requests, each with a requests-per-minute cap. The example spreads load across three Ollama instances limited to 30 rpm each.
- Fallback chains, which the author calls waterfall routing: tiers escalate from a fast local model to a larger local model, then a cheap remote option such as gpt-3.5-turbo, and finally a premium pick like gpt-4o. Each tier carries a tokens-per-minute budget, which keeps the expensive tiers rationed.
Deployment options
Three ways to run the proxy are covered. Docker is the recommended path, either as a docker run command or a compose file that mounts the config directory and passes API keys as environment variables. A direct pip install suits development and debugging. For Kubernetes hosts such as TrueNAS SCALE, a Helm chart installs LiteLLM into a dedicated namespace with the service exposed on port 4000.
Why it matters
The pattern generalises well beyond the homelab. Any team touching more than one LLM endpoint — a local inference box plus a couple of vendors — hits the same sprawl of keys, SDKs and failure modes. LiteLLM bundles the standard fixes into one self-hosted process: a stable OpenAI-shaped contract for clients, provider failover, load spreading and metering. It also softens lock-in, since swapping the model behind a virtual name becomes a configuration change rather than an application rewrite, and the built-in rate and token budgets become more valuable as usage grows.
- #litellm
- #llm-gateway
- #self-hosting
- #openai-compatible-api
- #ai-infrastructure