Skip to content
vLLM
Summary
Initializing search
GitHub
Home
User Guide
Developer Guide
Benchmarking
API Reference
CLI Reference
Community
vLLM
GitHub
Home
User Guide
User Guide
Getting Started
Getting Started
Quickstart
Installation
Installation
GPU
CPU
TPU
Examples
Examples
Applications
Applications
API Server
Chatbot
Rag
Basic
Basic
Offline Inference
Online Serving
Deployment
Deployment
Async LLM Streaming
Helm Charts
LLM Engine Example
Sagemaker-Entrypoint
Disaggregated
Disaggregated
Disaggregated Encoder
Disaggregated Serving
Ec Both Encoder
Disaggregated Prefill V1
Flexkv Connector
KV Load Failure Recovery Test
LMCache Examples
Mooncake Connector
Features
Features
Automatic Prefix Caching
Batch Invariance
Context Extension
Data Parallel
Kv Events
Logging Configuration
Custom Logits Processors
LoRA
Offline Inference with the OpenAI Batch file format
Pause Resume
Profiling
Prompt Embed
Reset Kv
Sharded State
Speculative Decoding
Structured Outputs
Tensorize vLLM Model
Torchrun
Generate
Generate
Batched Chat Completions Online
Multimodal
Qwen 1M Offline
Observability
Observability
Monitoring Dashboards
Metrics
Setup OpenTelemetry POC
Prometheus and Grafana
Pooling
Pooling
Classify
Embed
Plugin
Reward
Score
Token Classify
Token Embed
Ray Serving
Ray Serving
Batch LLM Inference
Elastic Ep
Multi-Node-Serving
Ray Serve Deepseek
Run Cluster
Reasoning
Reasoning
OpenAI Chat Completion Tool Calls With Reasoning
OpenAI Chat Completion With Reasoning
OpenAI Chat Completion With Reasoning Streaming
OpenAI Responses Client
RL
RL
RLHF Async New APIs
RLHF Http IPC
RLHF Http NCCL
RLHF IPC
RLHF IPC Fsdp Ep
RLHF NCCL
RLHF NCCL Fsdp Ep
RLHF Sparse NCCL
Routed Experts E2E
Skip Loading Weights In Engine Init
Scale Out
Scale Out
Init
Example Mm Serve
Token Generation Client
Speech To Text
Speech To Text
Lid
OpenAI
Realtime
Tool Calling
Tool Calling
Chat With Tools Offline
OpenAI Chat Completion Client With Tools
OpenAI Chat Completion Client With Tools Required
OpenAI Chat Completion Client With Tools Xlam
OpenAI Chat Completion Client With Tools Xlam Streaming
OpenAI Responses Client With Mcp Tools
OpenAI Responses Client With Tools
General
General
vLLM V1
Frequently Asked Questions
Production Metrics
Reproducibility
Security
Troubleshooting
Usage Stats Collection
Inference and Serving
Inference and Serving
Offline Inference
Online Serving
Online Serving
Derenderer APIs
Generative Scoring
OpenAI-Compatible Server
Renderer APIs
Speech to Text APIs
Context Parallel Deployment
Data Parallel Deployment
Troubleshooting distributed deployments
Expert Parallel Deployment
Parallelism and Scaling
Integrations
Integrations
Claude Code
Codex
LangChain
LlamaIndex
Deployment
Deployment
Using Docker
Using Kubernetes
Using Nginx
Frameworks
Frameworks
Anyscale
AnythingLLM
AutoGen
BentoML
Cerebrium
Chatbox
Dify
dstack
Haystack
Helm
Hugging Face Inference Endpoints
LiteLLM
Lobe Chat
LWS
Modal
Open WebUI
Retrieval-Augmented Generation
RunPod
SkyPilot
Streamlit
NVIDIA Triton
Integrations
Integrations
AIBrix
NVIDIA Dynamo
KAITO
KServe
Kthena
KubeAI
KubeRay
Llama Stack
llm-d
llmaz
Production stack
Training
Training
Async Reinforcement Learning
What is Layerwise (Re)loading?
Reinforcement Learning from Human Feedback
Transformers Reinforcement Learning
Weight Transfer
Weight Transfer
Base Class and Custom Engines
IPC Engine
NCCL Engine
Configuration
Configuration
Conserving Memory
Engine Arguments
Environment Variables
Model Resolution
Optimization and Tuning
Server Arguments
TPU
Models
Models
Supported Models
Generative Models
Pooling Models
Pooling Models
Classification Usages
Embedding Usages
Reward Usages
Scoring Usages
Specific Model Examples
Token Classification Usages
Token Embedding Usages
Extensions
Extensions
Loading model weights with fastsafetensors
Loading Model Weights with InstantTensor
Loading models with Run:ai Model Streamer
Loading models with CoreWeave's Tensorizer
Hardware Supported Models
Hardware Supported Models
CPU - Intel® Xeon®
XPU - Intel® GPUs
TPU
Features
Features
Automatic Prefix Caching
Batch Invariance
Context Extension
Custom Arguments
Custom Logits Processors
Disaggregated Encoder
Disaggregated Prefilling (experimental)
IndexCache
Interleaved Thinking
KV Offloading Usage Guide
LoRA Adapters
MooncakeConnector Usage Guide
MooncakeStoreConnector Usage Guide
MoRIIOConnector Usage Guide
Multimodal Inputs
NixlConnector Compatibility Matrix
NixlConnector Usage Guide
Per-Request Metrics
Prompt Embedding Inputs
Reasoning Outputs
Sleep Mode
Structured Outputs
Tool Calling
Quantization
Quantization
AutoAWQ
BitsAndBytes
FP8 ViT Encoder Attention
GGUF
GPTQModel
Intel Quantization Support
NVIDIA Model Optimizer
Online Quantization
Quantized KV Cache
AMD Quark
TorchAO
LLM Compressor
LLM Compressor
FP8 W8A8
INT4 W4A16
INT8 W4A8
INT8 W8A8
Speculative Decoding
Speculative Decoding
Draft Models
Dynamic Speculative Decoding
EAGLE Draft Models
Hidden State Extraction
MLP Draft Models
MTP (Multi-Token Prediction)
N-Gram Speculation
Parallel Draft Models
vLLM-Project/Speculators
Suffix Decoding
Developer Guide
Developer Guide
General
General
Deprecation Policy
Dockerfile
Editing Agent Instructions
Incremental Compilation Workflow
Profiling vLLM
Vulnerability Management
Model Implementation
Model Implementation
Basic Model
Registering a Model
Unit Testing
Multi-Modal Support
Speech-to-Text (Transcription/Translation) Support
CI
CI
CI Failures
Nightly Builds of vLLM Wheels
Update PyTorch version on vLLM OSS CI/CD
Design Documents
Design Documents
Plugins
Plugins
Endpoint Plugins
IO Processor Plugins
LoRA Resolver Plugins
Plugin System
Architecture Overview
Attention Backend Feature Support
CUDA Graphs
Vision Encoder (ViT) CUDA Graphs
CustomOp
Dual Batch Overlap
How to debug the vLLM-torch.compile integration
Fused MoE Modular Kernel
Fusion torch.compile passes
Integration with Hugging Face
Hybrid KV Cache Manager
Logits Processors
Metrics
Multi-Modal Data Processing
Model Runner V2 Design Document
Fused MoE Kernel Features
Python Multiprocessing
NIXL KV Cache Lease Renewal
NIXL push-mode KV transfer
Optimization Levels
Paged Attention
Automatic Prefix Caching
torch.compile integration
torch.compile with Multimodal Encoders
vLLM IR: Functional Intermediate Representation
Benchmarking
Benchmarking
Benchmark CLI
Parameter Sweeps
Performance Dashboard
API Reference
API Reference
vllm
vllm
collect_env
connections
env_override
envs
exceptions
forward_context
logger
logits_process
logprobs
model_inspection
outputs
pooling_params
sampling_params
scalar_type
scripts
sequence
tasks
version
assets
assets
audio
base
image
video
benchmarks
benchmarks
latency
mm_processor
plot
serve
startup
throughput
datasets
datasets
create_txt_slices_dataset
datasets
utils
lib
lib
endpoint_request_func
ready_checker
utils
sweep
sweep
cli
param_sweep
plot
plot_pareto
serve
serve_workload
server
startup
utils
compilation
compilation
backends
base_static_graph
breakable_cudagraph
caching
codegen
compiler_interface
counter
cuda_graph
decorators
monitor
partition_rules
piecewise_backend
wrapper
passes
passes
fx_utils
inductor_pass
pass_manager
vllm_inductor_pass
fusion
fusion
act_quant_fusion
add_rms_fusion
allreduce_rms_fusion
attn_quant_fusion
collective_fusion
matcher_utils
mla_attn_quant_fusion
mla_rope_kvcache_cat_fusion
qk_norm_rope_fusion
qk_norm_rope_kvcache_fusion
rms_quant_fusion
rocm_aiter_fusion
rope_kvcache_fusion
sequence_parallelism
ir
ir
clone_elimination
inplace_functionalization
lowering_pass
utils
utility
utility
fix_functionalization
noop_elimination
post_cleanup
scatter_split_replace
split_coalescing
config
config
attention
cache
compilation
device
diffusion
ec_manager_config
ec_transfer
fault_tolerance
kernel
kv_events
kv_transfer
load
lora
mamba
model
model_arch
multimodal
observability
offload
parallel
pooler
profiler
quantization
reasoning
scheduler
speculative
speech_to_text
structured_outputs
utils
vllm
weight_transfer
cute_utils
cute_utils
cvt
device_allocator
device_allocator
cumem
sleep_mode_backend
xpumem
distributed
distributed
communication_op
kv_events
nixl_utils
parallel_state
stateless_coordinator
utils
device_communicators
device_communicators
aiter_custom_all_reduce
all2all
all_reduce_utils
base_device_communicator
cpu_communicator
cuda_communicator
cuda_wrapper
custom_all_reduce
flashinfer_all_reduce
mnnvl_compat
pynccl
pynccl_allocator
pynccl_wrapper
quick_all_reduce
ray_communicator
shm_broadcast
shm_object_storage
symm_mem
xpu_communicator
ec_transfer
ec_transfer
ec_transfer_state
ec_connector
ec_connector
base
example_connector
factory
cpu
cpu
common
connector
ec_shared_region
scheduler
scheduler
embedding_cache
step_tracker
worker
worker
descriptor_buffers
elastic_ep
elastic_ep
elastic_execute
elastic_state
standby_state
eplb
eplb
async_worker
eplb_communicator
eplb_state
eplb_utils
rebalance_execute
policy
policy
abstract
default
kv_transfer
kv_transfer
kv_transfer_state
kv_connector
kv_connector
base
factory
utils
v1
v1
base
decode_bench_connector
example_connector
example_hidden_states_connector
flexkv_connector
lmcache_connector
lmcache_mp_connector
metrics
multi_connector
offloading_connector
simple_cpu_offload_connector
ssm_conv_transfer_utils
hf3fs
hf3fs
hf3fs_client
hf3fs_connector
hf3fs_metadata_server
utils
utils
common
gather_scatter_helper
hf3fs_mock_client
lmcache_integration
lmcache_integration
multi_process_adapter
utils
vllm_v1_adapter
mooncake
mooncake
mooncake_connector
mooncake_utils
rdma_utils
stats
store
store
connector
coordinator
data
metrics
protocol
scheduler
worker
moriio
moriio
moriio_common
moriio_connector
moriio_engine
moriio_layout
nixl
nixl
base_scheduler
base_worker
connector
metadata
pull_scheduler
pull_worker
push_scheduler
push_worker
scheduler
stats
tp_mapping
utils
worker
offloading
offloading
canonical_mapping
common
config
events
metrics
scheduler
worker
weight_transfer
weight_transfer
base
clients
factory
ipc_engine
nccl_common
nccl_engine
packed_tensor
sparse_nccl_engine
engine
engine
arg_utils
async_llm_engine
llm_engine
protocol
entrypoints
entrypoints
chat_utils
grpc_server
launcher
llm
offline_utils
anthropic
anthropic
api_router
protocol
serving
cli
cli
collect_env
launch
main
openai
run_batch
serve
types
benchmark
benchmark
base
latency
main
mm_processor
serve
startup
sweep
throughput
cohere
cohere
api_router
cohere_chat_message
protocol
serving
generate
generate
api_router
factories
base
base
serving
beam_search
beam_search
offline
online
utils
generative_scoring
generative_scoring
api_router
serving
mcp
mcp
tool
tool_server
openai
openai
api_server
cli_args
dp_supervisor
run_batch
chat_completion
chat_completion
api_router
batch_serving
protocol
serving
completion
completion
api_router
protocol
serving
engine
engine
protocol
models
models
api_router
protocol
serving
parser
parser
harmony_utils
responses
responses
api_router
context
harmony
protocol
serving
streaming_events
utils
pooling
pooling
factories
offline
typing
utils
base
base
io_processor
protocol
serving
classify
classify
api_router
io_processor
protocol
serving
embed
embed
api_router
io_processor
protocol
serving
pooling
pooling
api_router
io_processor
protocol
serving
scoring
scoring
api_router
io_processor
protocol
serving
typing
utils
scale_out
scale_out
factories
derender
derender
api_router
serving
render
render
api_router
serving
token_in_token_out
token_in_token_out
api_router
mm_serde
protocol
serving
serve
serve
dev
dev
cache
cache
api_router
rlhf
rlhf
api_router
rpc
rpc
api_router
server_info
server_info
api_router
sleep
sleep
api_router
elastic_ep
elastic_ep
api_router
middleware
engine
engine
serving
typing
fault_tolerance
fault_tolerance
api_router
instrumentator
instrumentator
basic
health
metrics
offline_docs
lora
lora
api_router
protocol
profile
profile
api_router
sagemaker
sagemaker
api_router
tokenize
tokenize
api_router
protocol
serving
utils
utils
api_utils
constants
error_response
fingerprint
orca_metrics
request_logger
server_utils
ssl
tool_calls_utils
speech_to_text
speech_to_text
factories
base
base
protocol
serving
utils
realtime
realtime
api_router
connection
metrics
protocol
serving
transcription
transcription
api_router
protocol
serving
translation
translation
api_router
protocol
serving
inputs
inputs
engine
llm
preprocess
ir
ir
op
tolerances
util
ops
ops
layernorm
kernels
kernels
aiter_ops
oink_ops
vllm_c
helion
helion
case_key
config_manager
register
utils
ops
ops
dynamic_per_token_scaled_fp8_quant
fused_qk_norm_rope
per_token_group_fp8_quant
rms_norm_dynamic_per_token_quant
rms_norm_per_block_quant
silu_and_mul_per_block_quant
silu_mul_fp8
triton
triton
qkv_padded_fp8_quant
logging_utils
logging_utils
access_log_filter
dump_input
formatter
lazy
log_time
torch_tensor
lora
lora
lora_model
lora_weights
model_manager
peft_helper
request
resolver
utils
worker_manager
layers
layers
base
base_linear
column_parallel_linear
fused_moe
logits_processor
replicated_linear
row_parallel_linear
utils
vocal_parallel_embedding
ops
ops
torch_ops
torch_ops
lora_ops
triton_ops
triton_ops
fp8_kernel_utils
fused_moe_lora_fp8_op
fused_moe_lora_op
kernel_utils
lora_expand_fp8_op
lora_expand_op
lora_kernel_metadata
lora_shrink_fp8_op
lora_shrink_op
utils
xpu_ops
xpu_ops
lora_ops
punica_wrapper
punica_wrapper
punica_base
punica_cpu
punica_gpu
punica_selector
punica_xpu
utils
model_executor
model_executor
custom_op
parameter
utils
kernels
kernels
attention
attention
dsa
dsa
dcp_indexer_cutedsl
linear
linear
base
zentorch_utils
cute_dsl
cute_dsl
ll_bf16
skinny_gemm
mixed_precision
mixed_precision
allspark
conch
cpu
cutlass
dynamic_4bit
exllama
humming
MPLinearKernel
machete
marlin
rdna3_w4a16
rdna_hybrid_w4a16
triton_w4a16
xpu
zentorch
mxfp4
mxfp4
aiter
base
emulation
flashinfer
humming
marlin
xpu
mxfp6
mxfp6
base
emulation
mxfp8
mxfp8
emulation
flashinfer
humming
Mxfp8LinearKernel
marlin
rocm_native
xpu
nvfp4
nvfp4
base
cutlass
emulation
fbgemm
flashinfer
humming
marlin
scaled_mm
scaled_mm
aiter
BlockScaledMMLinearKernel
cpu
cutlass
deep_gemm
flashinfer
humming
marlin
pytorch
rocm
ScaledMMLinearKernel
triton
xpu
zentorch
mhc
mhc
aiter
tilelang
tilelang_kernels
torch
triton
layers
layers
activation
attention_layer_base
batch_invariant
conv
fused_allreduce_gemma_rms_norm
fused_qk_norm_rope
layernorm
lightning_attn
linear
logits_processor
mhc
mla
resampler
sparse_attn_indexer
utils
vocab_parallel_embedding
attention
attention
attention
chunked_local_attention
cross_attention
encoder_only_attention
kv_transfer_utils
mla_attention
mm_encoder_attention
pcp
prefill_prefix_lm_attention
rswa_attention
sparse_mla_attention
sparse_mla_mask
static_sink_attention
fused_moe
fused_moe
activation
all2all_utils
config
deep_gemm_utils
eep_reconfigure
expert_map_manager
fused_flydsl_moe
fused_moe
fused_moe_method_base
fused_moe_modular_method
hpc_moe
layer
modular_kernel
moe_align_block_size
moe_fused_mul_sum
moe_permute_unpermute
routed_experts
routed_experts_capturer
topk_weight_and_reduce
unquantized_fused_moe_method
utils
experts
experts
aiter_mxfp4_w4a8_moe
aiter_mxfp8_moe
batched_deep_gemm_moe
cpu_int4_moe
cpu_moe
cutlass_moe
deep_gemm_moe
fallback
flashinfer_b12x_moe
flashinfer_cutedsl_batched_moe
flashinfer_cutedsl_moe
flashinfer_cutlass_moe
fused_batched_moe
fused_humming_moe
gpt_oss_triton_kernels_moe
int4_emulation_moe
lora_context
lora_experts_mixin
marlin_moe
mxfp8_emulation_moe
mxfp8_native_moe
nvfp4_emulation_moe
ocp_mx_emulation_moe
rocm_aiter_moe
triton_cutlass_moe
triton_deep_gemm_moe
triton_moe
trtllm_bf16_moe
trtllm_fp8_moe
trtllm_lora_moe
trtllm_mxfp4_moe
trtllm_mxint4_moe
trtllm_nvfp4_moe
xpu_moe
oracle
oracle
base
fp8
int8
int_wna16
mxfp4
mxfp8
nvfp4
unquantized
w4a8
w4a8_int8
prepare_finalize
prepare_finalize
batched
deepep_ht
deepep_ll
deepep_v2
flashinfer_nvlink_one_sided
flashinfer_nvlink_two_sided
mori
naive_dp_ep
nixl_ep
no_dp_ep
router
router
aiter_shared_routed_fused_moe_router
base_router
bf16x3_router_gemm_cutedsl
custom_routing_router
dsv4_topk
fused_moe_router