LiteLLM Proxy Multi Model Routing in Practice
What is LiteLLM?
LiteLLM is an LLM API proxy that provides a standard OpenAI-compatible interface and can route to 100+ LLM providers behind the scenes.
Client → LiteLLM Proxy (:4000) → DeepSeek / ModelStudio / Ollama / ...
Installation
pip install litellm[proxy]
Configuration
litellm_config.yaml:
model_list:
- model_name: deepseek-flash
litellm_params:
model: deepseek/deepseek-chat
api_key: sk-xxx
- model_name: deepseek-pro
litellm_params:
model: deepseek/deepseek-reasoner
api_key: sk-xxx
- model_name: qwen-max
litellm_params:
model: dashscope/qwen-max
api_key: sk-xxx
- model_name: local-qwen
litellm_params:
model: ollama/qwen2.5:7b
api_base: http://localhost:11434
router_settings:
routing_strategy: "usage-based" # Load balancing
allowed_fails: 3
num_retries: 2
fallbacks:
- deepseek-flash: [qwen-max, local-qwen]
general_settings:
master_key: sk-litellm-admin
Start:
litellm --config litellm_config.yaml --port 4000
Using with Claude Code
{
"apiKey": "sk-litellm-admin",
"baseURL": "http://localhost:4000/v1",
"model": "deepseek-flash"
}
Cost Tracking
LiteLLM automatically logs token usage and cost for each call. Visit http://localhost:4000/ui to view the Dashboard.
Routing Strategy Trade-offs
Routing strategy determines how requests are distributed across backends, and choosing the wrong one can cause cost or latency to spiral out of control.
| Strategy | Suitable scenario | Trade-off |
|---|---|---|
| usage-based | When backend quotas vary widely and you want to direct volume to those with lower current usage | Latency fluctuates when distribution is uneven |
| least-busy | When there are many backends with similar capabilities | Requires accurate real-time load information |
| latency-based | For latency-sensitive use cases | Requires continuous probing |
| simple-shuffle | When backends are equivalent and simplicity is desired | Does not consider usage and may hit quota limits |
Note that if models of different versions or from different providers are placed under the same alias, response quality can drift, and users may find it hard to notice. In practice: only truly equivalent deployments should share an alias; models with different quality or capabilities should use different aliases, leaving the upper layer to choose explicitly. This way, routing only handles availability, while quality differences are controlled by the caller.
Key Configuration Points for Fallback, Retry, and Rate Limiting
Retry count: Too few will cause outright failures when the backend has intermittent errors, while too many will amplify latency and increase backend load. Start with a conservative value and adjust after observing logs.
Allowed failure count: Once the threshold is exceeded, the deployment is temporarily removed from rotation to prevent requests from continuing to hit an already unhealthy backend. The cooldown period must match the backend recovery time. Too short will repeatedly run into the same problem; too long will waste available capacity.
Fallback list: It is triggered only for specific errors (for example, rate limiting or server errors), so confirm that the error types are matched correctly. A common mistake is pointing the fallback to the same failed backend, or having model aliases in the list inconsistent with upstream definitions, causing silent failures.
Timeout settings: Without timeouts, a single slow request can hold up the entire response chain. Set a reasonable limit for each backend and coordinate it with the higher level retry strategy to avoid retry stacking causing an avalanche.
In addition, error code semantics vary by provider, and the same number may represent different causes on different platforms. Therefore, error handling must be differentiated by provider, and a single rule should not be applied to all backends.
Cost Tracking and Budget Control
LiteLLM records token usage and calculates cost for each call, but the calculation depends on its built-in price list. If the model being used is not included, the cost will be inaccurate, and you need to add the unit price yourself. Therefore, the first task is to verify whether the figures in the Dashboard are consistent with the actual bill. If there is a discrepancy, correct the pricing first.
The next issue is attribution: distinguish users, projects, or environments with different keys so that costs can be categorized correctly instead of being mixed into a single general ledger. Combined with budgets and alert thresholds, this allows you to be notified before a single project uses up its quota.
Note that local models (for example, those served through Ollama) do not incur API fees, but they still consume local compute and electricity. If cloud and local models are mixed on the same route, the cost figures reflect only the cloud portion, so local resources must also be included when evaluating total expenditure.
Secrets and Deployment Security
The api_key and master_key in the configuration file are the risk points that are most easily overlooked.
- Do not write real keys into a configuration file and commit it to version control; inject them through environment variables or protected storage.
- The master key has the highest privileges, so issue it only for administrative purposes; for daily calls, use virtual keys with restricted permissions.
- Add authentication to externally exposed proxies, and do not leave
:4000directly open on the public network. - The Dashboard may display usage and key information, so access must be restricted.
- If a local service (for example, Ollama) is bound to
0.0.0.0, other devices on the same network segment can also call it, so confirm whether this matches expectations.
Verification Methods and Common Errors
Verification Methods
First confirm the model list: call /v1/models to check whether all aliases appear as expected. Then send the simplest possible request to each different alias one by one, confirming that requests actually reach the correct backend and are not all routed to the same one. Next, conduct a failure drill: intentionally make one backend unavailable (for example, by entering the wrong key or disconnecting the local service), and confirm that fallback takes over and responses can still complete. Finally, compare Dashboard costs with actual billing to confirm that the tracked numbers are trustworthy.
Common Errors
- Exposing the master key or writing it directly into a shared configuration file.
- Hardcoding real keys in configuration files and committing them to the version control repository.
- The fallback list points to the same failed backend, which means there is no redundancy.
- No timeout is configured, so a slow backend drags down overall responses.
- Model aliases do not match the names pasted downstream, causing requests to be rejected or routed to unintended models.
- Mixing models with different capabilities under the same alias, causing quality to fluctuate.
- Ignoring privacy boundaries when mixing cloud and local, routing sensitive data to external providers.
- Only looking at total cost without breaking it down by project or user, making it impossible to locate the source of spend.
Checklist
- Confirmed that the alias list from
/v1/modelsmatches the plan. - Each alias has been tested to confirm it reaches the correct backend.
- Retry count, allowed failure count, and cooldown time have been defined.
- Timeouts have been configured for each backend.
- At least one failure drill has been completed to verify that fallback works.
- Separate keys are used to distinguish users or projects.
- Dashboard costs have been reconciled with actual billing.
- Budget and alert thresholds have been configured.
- Confirmed that sensitive data will not be routed to disallowed external providers.
- Keys are managed via environment variables or protected storage and are not written to the version control repository.
Recommended Reading
More in Tools
- PaddleOCR in Practice: Extracting Hong Kong Stock Annual Report Financial Data in 83 Seconds
- Webb-Site: The Essential Hidden Treasure for Hong Kong Stock Research, a One-Click Tool to Get Annual Report PDFs for All Listed Companies
- Academic Research Skills Deep Technical Breakdown: How 45+ Agents Collaborate to Complete the Full Workflow from Literature Review to Peer Review
- AI Engineering from Scratch Deep Dive: 435 Lessons × 20 Stages