Complete Guide to vLLM Deployment and Performance Tuning
What Is vLLM?
vLLM is one of the fastest open-source LLM inference engines available. Its core innovation is PagedAttention, which is similar to virtual memory management in operating systems and significantly reduces KV Cache waste.
- Throughput: 24 times faster than HuggingFace Transformers
- Concurrency: Supports high-concurrency requests
- OpenAI-compatible API: Seamless migration
Installation
pip install vllm
Deploying a Model
# Start OpenAI-compatible server
vllm serve Qwen/Qwen2.5-7B-Instruct --port 8000
# Quantized deployment (to reduce memory)
vllm serve Qwen/Qwen2.5-7B-Instruct --quantization awq --port 8000
# Specify GPU memory
vllm serve Qwen/Qwen2.5-7B-Instruct --gpu-memory-utilization 0.9
Using with Claude Code
{
"apiKey": "not-needed",
"baseURL": "http://localhost:8000/v1",
"model": "Qwen/Qwen2.5-7B-Instruct"
}
Performance Tuning
| Parameter | Suggested Value | Description |
|---|---|---|
--max-model-len | 8192 | Maximum context length |
--gpu-memory-utilization | 0.9 | GPU memory utilization |
--tensor-parallel-size | Number of GPUs | Multi-GPU tensor parallelism |
Trade-offs Between PagedAttention and KV Cache
Understanding the memory behavior of KV Cache is a prerequisite for tuning. During generation, each request must retain key-value cache for processed tokens. Long contexts and high concurrency can quickly exhaust this memory. Early approaches reserved a full contiguous block for each request, causing significant internal fragmentation; PagedAttention divides the cache into fixed-size blocks and allocates on demand, providing exactly as much as is used, so it can accommodate more concurrent requests with less memory.
Block size: The smaller the block, the less fragmentation but the higher management overhead; the larger the block, the simpler management but increased tail waste. In most cases, default values are sufficient unless significant memory waste is observed under specific workloads.
Prefix caching: If many requests share the same prefix (for example, the same system prompt or the same long document), enabling prefix caching avoids recomputation and directly reuses already computed cache blocks, which clearly helps throughput. The cost is additional management of cache lifecycle and eviction.
Chunked prefill: Splitting the prefill of a long prompt into multiple segments allows generation and prefill to interleave, reducing blocking of other requests by long prompts and improving overall latency distribution. However, the finer the split, the slightly lower the efficiency of a single computation, requiring a trade-off between throughput and latency.
Quantization: Quantization such as AWQ or GPTQ can significantly reduce memory used by weights at the cost of some accuracy. The principle is to first confirm that output quality is acceptable, then consider how much concurrency the saved memory can buy.
Practical Tuning of Deployment Parameters
| Parameter | Adjustment Direction | Trade-off |
|---|---|---|
--max-model-len | Set based on actual requirements; do not set it to the maximum all at once | Setting it too high reserves too much KV Cache space |
--gpu-memory-utilization | Adjust based on the GPU and other processes on the same machine | If set too high, it may be preempted by other processes and fail |
--tensor-parallel-size | Set to the number of GPUs actually in use | If set incorrectly, it may fail to start or efficiency may decrease |
--max-num-seqs | Adjust the concurrency limit according to latency goals | Higher values improve throughput but increase single-request latency |
--dtype | Choose based on hardware support | Balance between precision and speed |
| Quantization options | Consider when memory is insufficient | Precision may decrease |
The practical workflow is to first get the model running and confirm that it can serve, then gradually increase load to find the stable upper limit. --gpu-memory-utilization is not a case where higher is better: if other processes on the same machine (for example, another service or a monitoring tool) need VRAM, leaving some headroom can prevent failures during startup or execution. For multi-GPU deployments, --tensor-parallel-size must match the number of GPUs actually visible, and you must confirm that the interconnect bandwidth between GPUs in the machine is sufficient; otherwise, the gains from tensor parallelism will be diminished.
Performance Benchmarking Methodology
Without measurement, there is no tuning. The following practices make comparison results credible.
- Keep input and output lengths fixed so that test conditions do not differ from run to run.
- Warm up first, and begin timing only after model loading and caches have stabilized.
- Perform a concurrency sweep: start from a single request and gradually increase the number of concurrent requests to observe how throughput and latency change.
- Measure first token latency (wait time) and per-token generation time separately; they affect the experience in different ways.
- Record throughput (tokens processed per second) and latency percentiles. Do not look only at averages, because averages mask the long tail.
- Change only one parameter at a time, and keep raw data for later review.
- Watch for cooling and thermal throttling; long continuous test runs can make later data points lower.
It is recommended to script the tests and keep the environment fixed, so that when changing models, quantization, or parameters, you can quickly rerun and compare. If you look at throughput and latency together, you will usually see a trade-off curve: higher concurrency means higher throughput, but each request waits longer. Ultimately, the operating point must be determined by the acceptable latency for your business.
Common Mistakes and Checklist
Common Mistakes
- Setting
--max-model-lendirectly to the model maximum, which causes the number of serviceable concurrent requests to drop significantly. - Setting
--gpu-memory-utilizationtoo high, causing contention for GPU memory with other processes on the same machine and leading to startup failure. --tensor-parallel-sizedoes not match the actual number of GPUs, so the service cannot start.- Looking only at throughput numbers and ignoring the latency distribution results in a poor real-world user experience.
- Comparing two configurations without warming up leads to unreliable conclusions.
- Enabling quantization without rechecking output quality.
- Treating
apiKeyas a required field and mistakenly assuming that a local service also requires authentication to connect. - Binding the service to a public interface without any access control.
Checklist
- Confirmed that the model can load and respond to requests normally.
- Set
--max-model-lenaccording to actual requirements. - Reserved a reasonable amount of GPU memory for other processes.
- Confirmed that
--tensor-parallel-sizematches the number of GPUs. - Decided whether prefix caching needs to be enabled.
- Measured first token latency and per-token generation time.
- Recorded throughput and latency percentiles, and performed a concurrency sweep.
- Confirmed that the output quality after quantization is acceptable.
- Confirmed that the service's network binding and access control meet expectations.
Recommended Reading
More in Tools
- PaddleOCR in Practice: Extracting Hong Kong Stock Annual Report Financial Data in 83 Seconds
- Webb-Site: The Essential Hidden Treasure for Hong Kong Stock Research, a One-Click Tool to Get Annual Report PDFs for All Listed Companies
- Academic Research Skills Deep Technical Breakdown: How 45+ Agents Collaborate to Complete the Full Workflow from Literature Review to Peer Review
- AI Engineering from Scratch Deep Dive: 435 Lessons × 20 Stages