Agentic Research

Complete Guide to vLLM Deployment and Performance Tuning

2026/05/1025 min readBryan Chan閱讀中文原文
TopicsInference

What Is vLLM?

vLLM is one of the fastest open-source LLM inference engines available. Its core innovation is PagedAttention, which is similar to virtual memory management in operating systems and significantly reduces KV Cache waste.

  • Throughput: 24 times faster than HuggingFace Transformers
  • Concurrency: Supports high-concurrency requests
  • OpenAI-compatible API: Seamless migration

Installation

pip install vllm

Deploying a Model

# Start OpenAI-compatible server
vllm serve Qwen/Qwen2.5-7B-Instruct --port 8000

# Quantized deployment (to reduce memory)
vllm serve Qwen/Qwen2.5-7B-Instruct --quantization awq --port 8000

# Specify GPU memory
vllm serve Qwen/Qwen2.5-7B-Instruct --gpu-memory-utilization 0.9

Using with Claude Code

{
  "apiKey": "not-needed",
  "baseURL": "http://localhost:8000/v1",
  "model": "Qwen/Qwen2.5-7B-Instruct"
}

Performance Tuning

ParameterSuggested ValueDescription
--max-model-len8192Maximum context length
--gpu-memory-utilization0.9GPU memory utilization
--tensor-parallel-sizeNumber of GPUsMulti-GPU tensor parallelism

Trade-offs Between PagedAttention and KV Cache

Understanding the memory behavior of KV Cache is a prerequisite for tuning. During generation, each request must retain key-value cache for processed tokens. Long contexts and high concurrency can quickly exhaust this memory. Early approaches reserved a full contiguous block for each request, causing significant internal fragmentation; PagedAttention divides the cache into fixed-size blocks and allocates on demand, providing exactly as much as is used, so it can accommodate more concurrent requests with less memory.

Block size: The smaller the block, the less fragmentation but the higher management overhead; the larger the block, the simpler management but increased tail waste. In most cases, default values are sufficient unless significant memory waste is observed under specific workloads.

Prefix caching: If many requests share the same prefix (for example, the same system prompt or the same long document), enabling prefix caching avoids recomputation and directly reuses already computed cache blocks, which clearly helps throughput. The cost is additional management of cache lifecycle and eviction.

Chunked prefill: Splitting the prefill of a long prompt into multiple segments allows generation and prefill to interleave, reducing blocking of other requests by long prompts and improving overall latency distribution. However, the finer the split, the slightly lower the efficiency of a single computation, requiring a trade-off between throughput and latency.

Quantization: Quantization such as AWQ or GPTQ can significantly reduce memory used by weights at the cost of some accuracy. The principle is to first confirm that output quality is acceptable, then consider how much concurrency the saved memory can buy.


Practical Tuning of Deployment Parameters

ParameterAdjustment DirectionTrade-off
--max-model-lenSet based on actual requirements; do not set it to the maximum all at onceSetting it too high reserves too much KV Cache space
--gpu-memory-utilizationAdjust based on the GPU and other processes on the same machineIf set too high, it may be preempted by other processes and fail
--tensor-parallel-sizeSet to the number of GPUs actually in useIf set incorrectly, it may fail to start or efficiency may decrease
--max-num-seqsAdjust the concurrency limit according to latency goalsHigher values improve throughput but increase single-request latency
--dtypeChoose based on hardware supportBalance between precision and speed
Quantization optionsConsider when memory is insufficientPrecision may decrease

The practical workflow is to first get the model running and confirm that it can serve, then gradually increase load to find the stable upper limit. --gpu-memory-utilization is not a case where higher is better: if other processes on the same machine (for example, another service or a monitoring tool) need VRAM, leaving some headroom can prevent failures during startup or execution. For multi-GPU deployments, --tensor-parallel-size must match the number of GPUs actually visible, and you must confirm that the interconnect bandwidth between GPUs in the machine is sufficient; otherwise, the gains from tensor parallelism will be diminished.


Performance Benchmarking Methodology

Without measurement, there is no tuning. The following practices make comparison results credible.

  • Keep input and output lengths fixed so that test conditions do not differ from run to run.
  • Warm up first, and begin timing only after model loading and caches have stabilized.
  • Perform a concurrency sweep: start from a single request and gradually increase the number of concurrent requests to observe how throughput and latency change.
  • Measure first token latency (wait time) and per-token generation time separately; they affect the experience in different ways.
  • Record throughput (tokens processed per second) and latency percentiles. Do not look only at averages, because averages mask the long tail.
  • Change only one parameter at a time, and keep raw data for later review.
  • Watch for cooling and thermal throttling; long continuous test runs can make later data points lower.

It is recommended to script the tests and keep the environment fixed, so that when changing models, quantization, or parameters, you can quickly rerun and compare. If you look at throughput and latency together, you will usually see a trade-off curve: higher concurrency means higher throughput, but each request waits longer. Ultimately, the operating point must be determined by the acceptable latency for your business.


Common Mistakes and Checklist

Common Mistakes

  • Setting --max-model-len directly to the model maximum, which causes the number of serviceable concurrent requests to drop significantly.
  • Setting --gpu-memory-utilization too high, causing contention for GPU memory with other processes on the same machine and leading to startup failure.
  • --tensor-parallel-size does not match the actual number of GPUs, so the service cannot start.
  • Looking only at throughput numbers and ignoring the latency distribution results in a poor real-world user experience.
  • Comparing two configurations without warming up leads to unreliable conclusions.
  • Enabling quantization without rechecking output quality.
  • Treating apiKey as a required field and mistakenly assuming that a local service also requires authentication to connect.
  • Binding the service to a public interface without any access control.

Checklist

  • Confirmed that the model can load and respond to requests normally.
  • Set --max-model-len according to actual requirements.
  • Reserved a reasonable amount of GPU memory for other processes.
  • Confirmed that --tensor-parallel-size matches the number of GPUs.
  • Decided whether prefix caching needs to be enabled.
  • Measured first token latency and per-token generation time.
  • Recorded throughput and latency percentiles, and performed a concurrency sweep.
  • Confirmed that the output quality after quantization is acceptable.
  • Confirmed that the service's network binding and access control meet expectations.

Recommended Reading