Skip to content
vLLM
Summary
Initializing search
GitHub
Home
User Guide
Developer Guide
Benchmarking
API Reference
CLI Reference
Community
vLLM
GitHub
Home
User Guide
User Guide
Getting Started
Getting Started
Quickstart
Installation
Installation
GPU
CPU
TPU
Examples
Examples
Applications
Applications
API Server
Chatbot
Rag
Basic
Basic
Offline Inference
Online Serving
Deployment
Deployment
Async LLM Streaming
Helm Charts
LLM Engine Example
Sagemaker-Entrypoint
Disaggregated
Disaggregated
Disaggregated Encoder
Disaggregated Serving
Ec Both Encoder
Disaggregated Prefill V1
Flexkv Connector
KV Load Failure Recovery Test
LMCache Examples
Mooncake Connector
Features
Features
Automatic Prefix Caching
Batch Invariance
Context Extension
Data Parallel
Kv Events
Logging Configuration
Custom Logits Processors
LoRA
Offline Inference with the OpenAI Batch file format
Pause Resume
Profiling
Prompt Embed
Reset Kv
Sharded State
Speculative Decoding
Structured Outputs
Tensorize vLLM Model
Torchrun
Generate
Generate
Batched Chat Completions Online
Multimodal
Qwen 1M Offline
Observability
Observability
Monitoring Dashboards
Metrics
Setup OpenTelemetry POC
Prometheus and Grafana
Pooling
Pooling
Classify
Embed
Plugin
Reward
Score
Token Classify
Token Embed
Ray Serving
Ray Serving
Batch LLM Inference
Elastic Ep
Multi-Node-Serving
Ray Serve Deepseek
Run Cluster
Reasoning
Reasoning
OpenAI Chat Completion Tool Calls With Reasoning
OpenAI Chat Completion With Reasoning
OpenAI Chat Completion With Reasoning Streaming
OpenAI Responses Client
RL
RL
RLHF Async New APIs
RLHF Http IPC
RLHF Http NCCL
RLHF IPC
RLHF IPC Fsdp Ep
RLHF NCCL
RLHF NCCL Fsdp Ep
RLHF Sparse NCCL
Routed Experts E2E
Skip Loading Weights In Engine Init
Scale Out
Scale Out
Init
Example Mm Serve
Token Generation Client
Speech To Text
Speech To Text
Lid
OpenAI
Realtime
Tool Calling
Tool Calling
Chat With Tools Offline
OpenAI Chat Completion Client With Tools
OpenAI Chat Completion Client With Tools Required
OpenAI Chat Completion Client With Tools Xlam
OpenAI Chat Completion Client With Tools Xlam Streaming
OpenAI Responses Client With Mcp Tools
OpenAI Responses Client With Tools
General
General
vLLM V1
Frequently Asked Questions
Production Metrics
Reproducibility
Security
Troubleshooting
Usage Stats Collection
Inference and Serving
Inference and Serving
Offline Inference
Online Serving
Online Serving
Derenderer APIs
Generative Scoring
OpenAI-Compatible Server
Renderer APIs
Speech to Text APIs
Context Parallel Deployment
Data Parallel Deployment
Troubleshooting distributed deployments
Expert Parallel Deployment
Parallelism and Scaling
Integrations
Integrations
Claude Code
Codex
LangChain
LlamaIndex
Deployment
Deployment
Using Docker
Using Kubernetes
Using Nginx
Frameworks
Frameworks
Anyscale
AnythingLLM
AutoGen
BentoML
Cerebrium
Chatbox
Dify
dstack
Haystack
Helm
Hugging Face Inference Endpoints
LiteLLM
Lobe Chat
LWS
Modal
Open WebUI
Retrieval-Augmented Generation
RunPod
SkyPilot
Streamlit
NVIDIA Triton
Integrations
Integrations
AIBrix
NVIDIA Dynamo
KAITO
KServe
Kthena
KubeAI
KubeRay
Llama Stack
llm-d
llmaz
Production stack
Training
Training
Async Reinforcement Learning
What is Layerwise (Re)loading?
Reinforcement Learning from Human Feedback
Transformers Reinforcement Learning
Weight Transfer
Weight Transfer
Base Class and Custom Engines
IPC Engine
NCCL Engine
Configuration
Configuration
Conserving Memory
Engine Arguments
Environment Variables
Model Resolution
Optimization and Tuning
Server Arguments
TPU
Models
Models
Supported Models
Generative Models
Pooling Models
Pooling Models
Classification Usages
Embedding Usages
Reward Usages
Scoring Usages
Specific Model Examples
Token Classification Usages
Token Embedding Usages
Extensions
Extensions
Loading model weights with fastsafetensors
Loading Model Weights with InstantTensor
Loading models with Run:ai Model Streamer
Loading models with CoreWeave's Tensorizer
Hardware Supported Models
Hardware Supported Models
CPU - Intel® Xeon®
XPU - Intel® GPUs
TPU
Features
Features
Automatic Prefix Caching
Batch Invariance
Context Extension
Custom Arguments
Custom Logits Processors
Disaggregated Encoder
Disaggregated Prefilling (experimental)
IndexCache
Interleaved Thinking
KV Offloading Usage Guide
LoRA Adapters
MooncakeConnector Usage Guide
MooncakeStoreConnector Usage Guide
MoRIIOConnector Usage Guide
Multimodal Inputs
NixlConnector Compatibility Matrix
NixlConnector Usage Guide
Per-Request Metrics
Prompt Embedding Inputs
Reasoning Outputs
Sleep Mode
Structured Outputs
Tool Calling
Quantization
Quantization
AutoAWQ
BitsAndBytes
FP8 ViT Encoder Attention
GGUF
GPTQModel
Intel Quantization Support
NVIDIA Model Optimizer
Online Quantization
Quantized KV Cache
AMD Quark
TorchAO
LLM Compressor
LLM Compressor
FP8 W8A8
INT4 W4A16
INT8 W4A8
INT8 W8A8
Speculative Decoding
Speculative Decoding
Draft Models
Dynamic Speculative Decoding
EAGLE Draft Models
Hidden State Extraction
MLP Draft Models
MTP (Multi-Token Prediction)
N-Gram Speculation
Parallel Draft Models
vLLM-Project/Speculators
Suffix Decoding
Developer Guide
Developer Guide
General
General
Deprecation Policy
Dockerfile
Editing Agent Instructions
Incremental Compilation Workflow
Profiling vLLM
Vulnerability Management
Model Implementation
Model Implementation
Basic Model
Registering a Model
Unit Testing
Multi-Modal Support
Speech-to-Text (Transcription/Translation) Support
CI
CI
CI Failures
Nightly Builds of vLLM Wheels
Update PyTorch version on vLLM OSS CI/CD
Design Documents
Design Documents
Plugins
Plugins
Endpoint Plugins
IO Processor Plugins
LoRA Resolver Plugins
Plugin System
Architecture Overview
Attention Backend Feature Support
CUDA Graphs
Vision Encoder (ViT) CUDA Graphs
CustomOp
Dual Batch Overlap
How to debug the vLLM-torch.compile integration
Fused MoE Modular Kernel
Fusion torch.compile passes
Integration with Hugging Face
Hybrid KV Cache Manager
Logits Processors
Metrics
Multi-Modal Data Processing
Model Runner V2 Design Document
Fused MoE Kernel Features
Python Multiprocessing
NIXL KV Cache Lease Renewal
NIXL push-mode KV transfer