vllm-prefix-cache-bench
Original:🇺🇸 English
Translated
This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns. Use when the user asks to benchmark prefix caching hit rate, caching efficiency, or repeated-prompt performance in vLLM.
16installs
Sourcevllm-project/vllm-skills
Added on
NPX Install
npx skill4agent add vllm-project/vllm-skills vllm-prefix-cache-benchTags
Translated version includes tags in frontmatterSKILL.md Content
View Translation Comparison →vLLM Prefix Caching Benchmark
Benchmark the efficiency of vLLM's automatic prefix caching (APC) feature. The offline script runs directly against the vLLM engine (no server required). For online/serving tests, use with the dataset.
benchmarks/benchmark_prefix_caching.pyvllm bench serveprefix_repetitionWhen to use
- User wants to measure the performance impact of prefix caching for repeated or partially-shared prompts.
- User wants to compare throughput/latency with and without .
--enable-prefix-caching - User wants to test prefix caching using a fixed synthetic prompt, a real dataset (e.g. ShareGPT), or a synthetic prefix/suffix repetition pattern.
Option 1 (default). Fixed Prompt with Prefix Caching
Runs a synthetic benchmark with a fixed prompt repeated multiple times to directly measure cache hit efficiency. No dataset download required.
bash
python3 benchmarks/benchmark_prefix_caching.py \
--model Qwen/Qwen3-8B \
--enable-prefix-caching \
--num-prompts 1 \
--repeat-count 100 \
--input-length-range 128:256To compare against the baseline without caching:
bash
python3 benchmarks/benchmark_prefix_caching.py \
--model Qwen/Qwen3-8B \
--no-enable-prefix-caching \
--num-prompts 1 \
--repeat-count 100 \
--input-length-range 128:256Option 2. ShareGPT Dataset with Prefix Caching
Uses real-world conversational data from ShareGPT to evaluate prefix caching with naturally occurring prompt sharing.
First, download the dataset:
bash
wget https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered/resolve/main/ShareGPT_V3_unfiltered_cleaned_split.jsonThen run the benchmark:
bash
python3 benchmarks/benchmark_prefix_caching.py \
--model Qwen/Qwen3-8B \
--dataset-path ShareGPT_V3_unfiltered_cleaned_split.json \
--enable-prefix-caching \
--num-prompts 20 \
--repeat-count 5 \
--input-length-range 128:256Option 3. Prefix Repetition Dataset (Online)
Uses with the synthetic dataset to test caching via the serving API. This requires a running vLLM server.
vllm bench serveprefix_repetitionFirst, start the server:
bash
vllm serve Qwen/Qwen3-8BThen run the benchmark:
bash
vllm bench serve \
--backend openai \
--model Qwen/Qwen3-8B \
--dataset-name prefix_repetition \
--num-prompts 100 \
--prefix-repetition-prefix-len 512 \
--prefix-repetition-suffix-len 128 \
--prefix-repetition-num-prefixes 5 \
--prefix-repetition-output-len 128Key parameters for :
prefix_repetition| Parameter | Description |
|---|---|
| Number of tokens in the shared prefix portion |
| Number of tokens in the unique suffix portion |
| Number of distinct prefixes to cycle through |
| Number of output tokens to generate per request |
Notes
- Run all commands from the root of the vLLM repository ().
cd vllm - Keep the default model () unless the user specifies a different one or the model is unavailable; change only
Qwen/Qwen3-8B.--model - in Option 1 and 2 controls how many times each sampled prompt is replayed; higher values increase cache hit rate.
--repeat-count - accepts a
--input-length-rangetoken range, e.g.min:max.128:256 - For multi-GPU setups, add .
--tensor-parallel-size <N> - To test different hash algorithms for prefix caching internals, use (requires
--prefix-caching-hash-algo xxhash).pip install xxhash
Arguments for benchmark_prefix_caching.py
benchmark_prefix_caching.py| Argument | Required | Description |
|---|---|---|
| Yes | Model name or path (HuggingFace ID or local path) |
| Yes | Number of prompts to process |
| Yes | Token length range for inputs, e.g. |
| No | Number of times each prompt is repeated (default: 1) |
| No | Path to a dataset file (e.g. ShareGPT JSON). Omit for synthetic fixed-prompt mode |
| No | Fixed prefix token length to prepend to every prompt |
| No | Number of output tokens to generate per request |
| No | Sort prompts by length before benchmarking |
| No | Toggle APC (recommended: enable to test caching) |
| No | Hash algorithm: |
| No | Number of GPUs for tensor parallelism |
| No | Skip detokenization to reduce overhead |
Troubleshooting
- If reports file not found, locate your local vLLM repository first and run the command from that repo root.
python3 benchmarks/*.py - If you do not have the repository yet, clone it and continue:
bash
git clone https://github.com/vllm-project/vllm
cd vllm- If HuggingFace model download fails due to access restrictions, set your token: or pass
export HF_TOKEN=<your_token>.--hf-token <your_token> - If or
xxhashis not installed and you use those hash algorithms, install them first:cbor2.pip install xxhash cbor2