vllm-server

Compare original and translation side by side

🇺🇸

Original

English
🇨🇳

Translation

Chinese

vLLM Server Management

vLLM 服务器管理

Deploy production-grade LLM inference servers with vLLM — the fastest open-source LLM serving engine with PagedAttention and continuous batching.
使用vLLM部署生产级LLM推理服务器——这是一款采用PagedAttention和连续批处理技术的最快开源LLM服务引擎。

When to Use This Skill

何时使用该技能

Use this skill when:
  • Serving open-source LLMs (Llama, Mistral, Qwen, Gemma) at scale
  • Building an OpenAI-compatible API endpoint for self-hosted models
  • Optimizing LLM throughput and latency for production traffic
  • Running multi-GPU inference with tensor or pipeline parallelism
  • Deploying quantized models to reduce GPU memory requirements
在以下场景中使用本技能:
  • 大规模部署开源LLM(Llama、Mistral、Qwen、Gemma)
  • 为自托管模型构建兼容OpenAI的API端点
  • 针对生产流量优化LLM的吞吐量与延迟
  • 通过张量或流水线并行运行多GPU推理
  • 部署量化模型以降低GPU内存需求

Prerequisites

前置条件

  • NVIDIA GPU(s) with CUDA 12.1+ (A100/H100 recommended for production)
  • Docker or Python 3.9+ with pip
  • 40GB+ VRAM for 70B models; 8GB+ for 7B models
  • nvidia-container-toolkit
    for Docker GPU passthrough
  • 搭载CUDA 12.1+的NVIDIA GPU(生产环境推荐使用A100/H100)
  • Docker或Python 3.9+及pip
  • 70B模型需40GB+显存;7B模型需8GB+显存
  • 用于Docker GPU透传的
    nvidia-container-toolkit

Quick Start

快速开始

bash
undefined
bash
undefined

Install vLLM

Install vLLM

pip install vllm
pip install vllm

Serve a model (OpenAI-compatible API)

Serve a model (OpenAI-compatible API)

vllm serve meta-llama/Llama-3.1-8B-Instruct
--host 0.0.0.0
--port 8000
--api-key your-secret-key
vllm serve meta-llama/Llama-3.1-8B-Instruct
--host 0.0.0.0
--port 8000
--api-key your-secret-key

Test the endpoint

Test the endpoint

curl http://localhost:8000/v1/chat/completions
-H "Content-Type: application/json"
-H "Authorization: Bearer your-secret-key"
-d '{ "model": "meta-llama/Llama-3.1-8B-Instruct", "messages": [{"role": "user", "content": "Hello!"}] }'
undefined
curl http://localhost:8000/v1/chat/completions
-H "Content-Type: application/json"
-H "Authorization: Bearer your-secret-key"
-d '{ "model": "meta-llama/Llama-3.1-8B-Instruct", "messages": [{"role": "user", "content": "Hello!"}] }'
undefined

Docker Deployment

Docker部署

bash
docker run --runtime nvidia --gpus all \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -p 8000:8000 \
  --ipc=host \
  vllm/vllm-openai:latest \
  --model meta-llama/Llama-3.1-8B-Instruct \
  --api-key your-secret-key
bash
docker run --runtime nvidia --gpus all \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -p 8000:8000 \
  --ipc=host \
  vllm/vllm-openai:latest \
  --model meta-llama/Llama-3.1-8B-Instruct \
  --api-key your-secret-key

Docker Compose (Production)

Docker Compose(生产环境)

yaml
services:
  vllm:
    image: vllm/vllm-openai:latest
    runtime: nvidia
    environment:
      - NVIDIA_VISIBLE_DEVICES=all
      - HUGGING_FACE_HUB_TOKEN=${HF_TOKEN}
    volumes:
      - model-cache:/root/.cache/huggingface
    ports:
      - "8000:8000"
    ipc: host
    command: >
      --model meta-llama/Llama-3.1-70B-Instruct
      --tensor-parallel-size 2
      --max-model-len 32768
      --gpu-memory-utilization 0.90
      --api-key ${VLLM_API_KEY}
    restart: unless-stopped
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
      interval: 30s
      timeout: 10s
      retries: 3

volumes:
  model-cache:
yaml
services:
  vllm:
    image: vllm/vllm-openai:latest
    runtime: nvidia
    environment:
      - NVIDIA_VISIBLE_DEVICES=all
      - HUGGING_FACE_HUB_TOKEN=${HF_TOKEN}
    volumes:
      - model-cache:/root/.cache/huggingface
    ports:
      - "8000:8000"
    ipc: host
    command: >
      --model meta-llama/Llama-3.1-70B-Instruct
      --tensor-parallel-size 2
      --max-model-len 32768
      --gpu-memory-utilization 0.90
      --api-key ${VLLM_API_KEY}
    restart: unless-stopped
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
      interval: 30s
      timeout: 10s
      retries: 3

volumes:
  model-cache:

Key Configuration Options

核心配置选项

Multi-GPU Tensor Parallelism

多GPU张量并行

bash
undefined
bash
undefined

Split one model across 4 GPUs

Split one model across 4 GPUs

vllm serve meta-llama/Llama-3.1-70B-Instruct
--tensor-parallel-size 4
--gpu-memory-utilization 0.90
undefined
vllm serve meta-llama/Llama-3.1-70B-Instruct
--tensor-parallel-size 4
--gpu-memory-utilization 0.90
undefined

Quantization (Lower VRAM)

量化(降低显存占用)

bash
undefined
bash
undefined

AWQ quantization (70B on 2x A100 40GB)

AWQ quantization (70B on 2x A100 40GB)

vllm serve casperhansen/llama-3-70b-instruct-awq
--quantization awq
--tensor-parallel-size 2
vllm serve casperhansen/llama-3-70b-instruct-awq
--quantization awq
--tensor-parallel-size 2

GPTQ quantization

GPTQ quantization

vllm serve TheBloke/Llama-2-70B-Chat-GPTQ
--quantization gptq
vllm serve TheBloke/Llama-2-70B-Chat-GPTQ
--quantization gptq

FP8 (H100 NVL native)

FP8 (H100 NVL native)

vllm serve meta-llama/Llama-3.1-405B-Instruct
--quantization fp8
--tensor-parallel-size 8
undefined
vllm serve meta-llama/Llama-3.1-405B-Instruct
--quantization fp8
--tensor-parallel-size 8
undefined

Structured Output & Tools

结构化输出与工具调用

bash
vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --enable-auto-tool-choice \
  --tool-call-parser llama3_json \
  --guided-decoding-backend outlines
bash
vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --enable-auto-tool-choice \
  --tool-call-parser llama3_json \
  --guided-decoding-backend outlines

LoRA Adapters

LoRA适配器

bash
vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --enable-lora \
  --lora-modules sql-lora=/path/to/sql-lora \
                 code-lora=/path/to/code-lora \
  --max-lora-rank 64
bash
vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --enable-lora \
  --lora-modules sql-lora=/path/to/sql-lora \
                 code-lora=/path/to/code-lora \
  --max-lora-rank 64

Performance Tuning

性能调优

bash
undefined
bash
undefined

Maximize throughput for batch workloads

Maximize throughput for batch workloads

vllm serve <model>
--max-num-seqs 256 \ # max concurrent sequences --max-num-batched-tokens 8192 \ # tokens per batch --gpu-memory-utilization 0.95 \ # use 95% VRAM --swap-space 4 # CPU swap (GiB)
vllm serve <model>
--max-num-seqs 256 \ # max concurrent sequences --max-num-batched-tokens 8192 \ # tokens per batch --gpu-memory-utilization 0.95 \ # use 95% VRAM --swap-space 4 # CPU swap (GiB)

Minimize latency for interactive use

Minimize latency for interactive use

vllm serve <model>
--max-num-seqs 32
--enforce-eager # disable CUDA graph capture
undefined
vllm serve <model>
--max-num-seqs 32
--enforce-eager # disable CUDA graph capture
undefined

Benchmarking

基准测试

bash
undefined
bash
undefined

Install benchmark tool

Install benchmark tool

pip install vllm
pip install vllm

Run throughput benchmark

Run throughput benchmark

python -m vllm.entrypoints.openai.run_batch
--model meta-llama/Llama-3.1-8B-Instruct
--input-file prompts.jsonl
--output-file results.jsonl
python -m vllm.entrypoints.openai.run_batch
--model meta-llama/Llama-3.1-8B-Instruct
--input-file prompts.jsonl
--output-file results.jsonl

Benchmark with vllm bench

Benchmark with vllm bench

vllm bench throughput
--model meta-llama/Llama-3.1-8B-Instruct
--num-prompts 1000
--input-len 512
--output-len 128
undefined
vllm bench throughput
--model meta-llama/Llama-3.1-8B-Instruct
--num-prompts 1000
--input-len 512
--output-len 128
undefined

Monitoring

监控

bash
undefined
bash
undefined

Check running server stats

Check running server stats

curl http://localhost:8000/metrics # Prometheus metrics
curl http://localhost:8000/metrics # Prometheus metrics

Key metrics to watch:

Key metrics to watch:

vllm:num_requests_running - active requests

vllm:num_requests_running - active requests

vllm:gpu_cache_usage_perc - KV cache utilization

vllm:gpu_cache_usage_perc - KV cache utilization

vllm:generation_tokens_per_s - throughput

vllm:generation_tokens_per_s - throughput

vllm:time_to_first_token_ms - TTFT latency

vllm:time_to_first_token_ms - TTFT latency

vllm:e2e_request_latency_seconds - end-to-end latency

vllm:e2e_request_latency_seconds - end-to-end latency

undefined
undefined

Common Issues

常见问题

IssueCauseFix
CUDA out of memory
Model too large for VRAMAdd
--quantization awq
or reduce
--gpu-memory-utilization
Slow cold startModel not cachedPre-pull with
huggingface-cli download <model>
Low throughputToo few concurrent requestsIncrease
--max-num-seqs
KV cache full errorsContext length too longSet
--max-model-len
lower
tokenizer error
Tokenizer mismatchUse
--tokenizer
to specify correct tokenizer
问题原因解决方法
CUDA out of memory
模型显存占用超出容量添加
--quantization awq
参数或降低
--gpu-memory-utilization
冷启动缓慢模型未缓存使用
huggingface-cli download <model>
预拉取模型
吞吐量低并发请求数过少提高
--max-num-seqs
参数值
KV缓存已满错误上下文长度过长降低
--max-model-len
参数值
tokenizer error
分词器不匹配使用
--tokenizer
指定正确的分词器

Best Practices

最佳实践

  • Use
    --gpu-memory-utilization 0.90
    to leave headroom for CUDA kernels.
  • Pin model versions with
    --revision
    for reproducible deployments.
  • Set
    HF_HUB_OFFLINE=1
    in production to prevent unexpected downloads.
  • Use AWQ or GPTQ quantization before tensor parallelism — lower VRAM first.
  • Enable
    --enable-chunked-prefill
    for long-context workloads.
  • Monitor
    gpu_cache_usage_perc
    — above 95% causes queuing.
  • 使用
    --gpu-memory-utilization 0.90
    为CUDA内核预留空间。
  • 通过
    --revision
    固定模型版本,确保部署可复现。
  • 生产环境中设置
    HF_HUB_OFFLINE=1
    ,避免意外下载。
  • 优先使用AWQ或GPTQ量化,再考虑张量并行——先降低显存占用。
  • 针对长上下文工作负载启用
    --enable-chunked-prefill
  • 监控
    gpu_cache_usage_perc
    ——超过95%会导致请求排队。

Related Skills

相关技能

  • llm-inference-scaling - Auto-scaling vLLM deployments
  • gpu-server-management - GPU driver setup
  • llm-gateway - Load balancing across vLLM instances
  • llm-cost-optimization - Cost management
  • model-serving-kubernetes - K8s deployment
  • llm-inference-scaling - vLLM部署自动扩容
  • gpu-server-management - GPU驱动配置
  • llm-gateway - vLLM实例负载均衡
  • llm-cost-optimization - 成本管理
  • model-serving-kubernetes - K8s部署