Agentic Research

LiteLLM Proxy Multi Model Routing in Practice

2026/05/1027 min readBryan Chan閱讀中文原文
TopicsInferenceAPI

What is LiteLLM?

LiteLLM is an LLM API proxy that provides a standard OpenAI-compatible interface and can route to 100+ LLM providers behind the scenes.

Client → LiteLLM Proxy (:4000) → DeepSeek / ModelStudio / Ollama / ...

Installation

pip install litellm[proxy]

Configuration

litellm_config.yaml:

model_list:
  - model_name: deepseek-flash
    litellm_params:
      model: deepseek/deepseek-chat
      api_key: sk-xxx
  - model_name: deepseek-pro
    litellm_params:
      model: deepseek/deepseek-reasoner
      api_key: sk-xxx
  - model_name: qwen-max
    litellm_params:
      model: dashscope/qwen-max
      api_key: sk-xxx
  - model_name: local-qwen
    litellm_params:
      model: ollama/qwen2.5:7b
      api_base: http://localhost:11434

router_settings:
routing_strategy: "usage-based"  # Load balancing
  allowed_fails: 3
  num_retries: 2
  fallbacks:
    - deepseek-flash: [qwen-max, local-qwen]

general_settings:
  master_key: sk-litellm-admin

Start:

litellm --config litellm_config.yaml --port 4000

Using with Claude Code

{
  "apiKey": "sk-litellm-admin",
  "baseURL": "http://localhost:4000/v1",
  "model": "deepseek-flash"
}

Cost Tracking

LiteLLM automatically logs token usage and cost for each call. Visit http://localhost:4000/ui to view the Dashboard.


Routing Strategy Trade-offs

Routing strategy determines how requests are distributed across backends, and choosing the wrong one can cause cost or latency to spiral out of control.

StrategySuitable scenarioTrade-off
usage-basedWhen backend quotas vary widely and you want to direct volume to those with lower current usageLatency fluctuates when distribution is uneven
least-busyWhen there are many backends with similar capabilitiesRequires accurate real-time load information
latency-basedFor latency-sensitive use casesRequires continuous probing
simple-shuffleWhen backends are equivalent and simplicity is desiredDoes not consider usage and may hit quota limits

Note that if models of different versions or from different providers are placed under the same alias, response quality can drift, and users may find it hard to notice. In practice: only truly equivalent deployments should share an alias; models with different quality or capabilities should use different aliases, leaving the upper layer to choose explicitly. This way, routing only handles availability, while quality differences are controlled by the caller.


Key Configuration Points for Fallback, Retry, and Rate Limiting

Retry count: Too few will cause outright failures when the backend has intermittent errors, while too many will amplify latency and increase backend load. Start with a conservative value and adjust after observing logs.

Allowed failure count: Once the threshold is exceeded, the deployment is temporarily removed from rotation to prevent requests from continuing to hit an already unhealthy backend. The cooldown period must match the backend recovery time. Too short will repeatedly run into the same problem; too long will waste available capacity.

Fallback list: It is triggered only for specific errors (for example, rate limiting or server errors), so confirm that the error types are matched correctly. A common mistake is pointing the fallback to the same failed backend, or having model aliases in the list inconsistent with upstream definitions, causing silent failures.

Timeout settings: Without timeouts, a single slow request can hold up the entire response chain. Set a reasonable limit for each backend and coordinate it with the higher level retry strategy to avoid retry stacking causing an avalanche.

In addition, error code semantics vary by provider, and the same number may represent different causes on different platforms. Therefore, error handling must be differentiated by provider, and a single rule should not be applied to all backends.

Cost Tracking and Budget Control

LiteLLM records token usage and calculates cost for each call, but the calculation depends on its built-in price list. If the model being used is not included, the cost will be inaccurate, and you need to add the unit price yourself. Therefore, the first task is to verify whether the figures in the Dashboard are consistent with the actual bill. If there is a discrepancy, correct the pricing first.

The next issue is attribution: distinguish users, projects, or environments with different keys so that costs can be categorized correctly instead of being mixed into a single general ledger. Combined with budgets and alert thresholds, this allows you to be notified before a single project uses up its quota.

Note that local models (for example, those served through Ollama) do not incur API fees, but they still consume local compute and electricity. If cloud and local models are mixed on the same route, the cost figures reflect only the cloud portion, so local resources must also be included when evaluating total expenditure.


Secrets and Deployment Security

The api_key and master_key in the configuration file are the risk points that are most easily overlooked.

  • Do not write real keys into a configuration file and commit it to version control; inject them through environment variables or protected storage.
  • The master key has the highest privileges, so issue it only for administrative purposes; for daily calls, use virtual keys with restricted permissions.
  • Add authentication to externally exposed proxies, and do not leave :4000 directly open on the public network.
  • The Dashboard may display usage and key information, so access must be restricted.
  • If a local service (for example, Ollama) is bound to 0.0.0.0, other devices on the same network segment can also call it, so confirm whether this matches expectations.

Verification Methods and Common Errors

Verification Methods

First confirm the model list: call /v1/models to check whether all aliases appear as expected. Then send the simplest possible request to each different alias one by one, confirming that requests actually reach the correct backend and are not all routed to the same one. Next, conduct a failure drill: intentionally make one backend unavailable (for example, by entering the wrong key or disconnecting the local service), and confirm that fallback takes over and responses can still complete. Finally, compare Dashboard costs with actual billing to confirm that the tracked numbers are trustworthy.

Common Errors

  • Exposing the master key or writing it directly into a shared configuration file.
  • Hardcoding real keys in configuration files and committing them to the version control repository.
  • The fallback list points to the same failed backend, which means there is no redundancy.
  • No timeout is configured, so a slow backend drags down overall responses.
  • Model aliases do not match the names pasted downstream, causing requests to be rejected or routed to unintended models.
  • Mixing models with different capabilities under the same alias, causing quality to fluctuate.
  • Ignoring privacy boundaries when mixing cloud and local, routing sensitive data to external providers.
  • Only looking at total cost without breaking it down by project or user, making it impossible to locate the source of spend.

Checklist

  • Confirmed that the alias list from /v1/models matches the plan.
  • Each alias has been tested to confirm it reaches the correct backend.
  • Retry count, allowed failure count, and cooldown time have been defined.
  • Timeouts have been configured for each backend.
  • At least one failure drill has been completed to verify that fallback works.
  • Separate keys are used to distinguish users or projects.
  • Dashboard costs have been reconciled with actual billing.
  • Budget and alert thresholds have been configured.
  • Confirmed that sensitive data will not be routed to disallowed external providers.
  • Keys are managed via environment variables or protected storage and are not written to the version control repository.

Recommended Reading