Most inference libraries treat GPU memory as one more thing for you to manage. vLLM, Apache-2.0-licensed and originally developed in the Sky Computing Lab at UC Berkeley, is the serving engine that made memory management the product: its PagedAttention scheme, published in the SOSP 2023 paper, treats attention key and value memory like virtual memory pages so that thousands of requests can share one physical pool without fragmentation. Grown into one of the most active open-source AI projects, with over 2000 contributors across academia and industry, it describes itself in three words: easy, fast, and cheap LLM serving. The README feature list backs the slogan up with specifics: continuous batching of incoming requests, chunked prefill, prefix caching, piecewise and full CUDA and HIP graphs, quantization from FP8 and NVFP4 down to INT4 and GGUF, attention kernels from FlashAttention, FlashInfer, and Triton, speculative decoding with n-gram, suffix, and EAGLE methods, tensor, data, expert, and context parallelism, and an OpenAI-compatible API server that also speaks the Anthropic Messages API and gRPC. The engine serves more than 200 Hugging Face model architectures, from decoder-only LLMs and Mixture-of-Experts giants like DeepSeek-V3 to multimodal, embedding, and reward models, and it runs on NVIDIA, AMD, and Intel GPUs plus x86, ARM, and PowerPC CPUs, with plugins reaching TPUs, Intel Gaudi, Huawei Ascend, and Apple Silicon.

The layout mirrors the split between Python orchestration and device code: the vllm package holds the API front ends and the v1 engine core, csrc holds the C++ and CUDA kernel sources, a rust directory carries components wired in through setuptools-rust, and tests, examples, benchmarks, docs, and requirements folders complete the picture. Python 3.11 through 3.14 are supported, the build pins torch 2.13.0, and the release version itself is generated at build time by setuptools-scm rather than hardcoded. Inside the package, the load-bearing directories are entrypoints, engine, v1, model_executor, distributed, lora, multimodal, and platforms, and the console script installed as vllm resolves to the CLI there. As always in this series, what follows is an educational tour of published source code.

vLLM overview architecture diagram

vLLM at a glance: the server launcher mounts OpenAI-compatible routes that feed the AsyncLLM front-end, requests cross into the EngineCore process where the continuous batching scheduler allocates PagedAttention blocks, the executor fans the batch out to GPU workers running model code over hand-tuned kernels, and results stream back through the output processor as detokenized deltas.

Reading the overview from left to right:

Why You Need This

The first reason is that throughput here is an architecture, not a flag you hope stays on. The KV cache manager at vllm/v1/core/kv_cache_manager.py allocates fixed-size blocks per request, and the block pool at vllm/v1/core/block_pool.py tracks free, cached, and touched blocks so that two prompts sharing a system prompt share physical pages, which is the prefix caching idea implemented through the block hashing in vllm/v1/core/kv_cache_utils.py. Above that, the scheduler at vllm/v1/core/sched/scheduler.py implements continuous batching: it admits new requests into a running batch whenever blocks free up, pauses or preempts low-priority work under memory pressure, and consults the priority request queues in vllm/v1/core/sched/request_queue.py to decide who goes next. This is the machinery that turns a GPU into a shared, saturated resource.

The second reason is that every decode-speed trick you have read about is actually in here, wired end to end. Speculative decoding lives under vllm/v1/spec_decode with n-gram and EAGLE-style drafters verified by the target model in one batched step. Structured generation is enforced by the grammar compilers under vllm/v1/structured_output, which feed token masks so models emit valid JSON against a schema. CUDA graph capture and replay, including piecewise compilation that interleaves graphs with attention, live at vllm/compilation/cuda_graph.py, and the model runner at vllm/v1/worker/gpu_model_runner.py is where batched forwards, logprobs, and multimodal embeddings meet the device. Reading these files is a masterclass in serving speed.

The third reason is the operational surface. The executor layer at vllm/v1/executor/multiproc_executor.py spawns one worker process per GPU for a single node, while the Ray-based path at vllm/v1/executor/ray_executor.py stretches the same stepping loop across machines. Disaggregated prefill and decode, where one pool of GPUs handles prompts and another handles generation, is plumbed through the KV transfer connectors under vllm/distributed/kv_transfer. Multi-LoRA serving of many adapters against one base model is implemented in vllm/lora, multimodal image and audio ingestion in vllm/multimodal, and the whole stack negotiates hardware differences through the platform backends at vllm/platforms.

vLLM detail architecture diagram

The detail view: the API layer with its routers, Anthropic adapter, gRPC server, MCP tools, and CLI on the left; the engine front-end with AsyncLLM, processors, detokenizer, admission control, and the ZMQ-based core client in the middle; scheduling and KV cache internals plus structured output and speculative decoding below that; and the execution layer with executors, workers, CUDA graphs, the model registry, layers, and kernels on the right.

Walking the detail diagram, the API layer is broader than most people expect. Beyond the OpenAI-compatible chat and completion routers, the repository ships an Anthropic Messages adapter under vllm/entrypoints/anthropic, a gRPC server at vllm/entrypoints/grpc_server.py, an MCP tool integration under vllm/entrypoints/mcp, and a CLI package at vllm/entrypoints/cli that backs the vllm serve command. For offline batch work there is the synchronous LLM class at vllm/entrypoints/llm.py, the API from a thousand quickstart scripts, which wraps the same async engine rather than a separate path.

In the engine front-end, the AsyncLLM class separates concerns cleanly: the input processor at vllm/v1/engine/input_processor.py renders prompts and multimodal features into engine-ready form, the admission controller at vllm/v1/engine/admission_control.py refuses work the queue cannot absorb, and the output processor at vllm/v1/engine/output_processor.py pairs the incremental detokenizer at vllm/v1/engine/detokenizer.py with per-request collectors so streaming clients see text the moment tokens are final. The core client at vllm/v1/engine/core_client.py carries requests to the EngineCore over a message channel, letting the API process and the GPU-driving core process crash and restart independently.

The execution layer closes the loop. The weight loader under vllm/model_executor/model_loader materializes checkpoints into architectures chosen from vllm/model_executor/models/registry.py, the quantized and fused layer implementations under vllm/model_executor/layers dispatch to kernels from the CUDA sources under csrc, and the GPU worker at vllm/v1/worker/gpu_worker.py hosts the model runner that owns the device, KV tensors, and CUDA graphs for its shard.

From Install to Serving

Installation is deliberately boring: uv pip install vllm, optionally followed by vllm download-kernels for a faster first launch. The environment contract in pyproject.toml is worth reading, however: Python 3.11 through 3.14 and a build pinned to torch 2.13.0, with the package version stamped by setuptools-scm from git tags rather than a static string. After install, the vllm console script is your entry point: vllm serve brings up the API server, and the launcher at vllm/entrypoints/launchers/api_server/entry.py builds the FastAPI application, registers the routers, and initializes the engine client before serving. From there the model runner warms up, captures CUDA graphs for common batch shapes, and the server begins answering OpenAI-shaped requests with streaming deltas.

Honest limits: the device coupling is real, because kernels are compiled for specific GPU architectures and building from source means a full CUDA toolchain and a long CMake run; the Python and torch pins in pyproject.toml constrain where the engine can live; and the internals move quickly, as the old vllm.entrypoints.openai.api_server module now emits a DeprecationWarning pointing at the launchers package, a reminder that internal import paths shift between releases. Single-node execution is the default, the Ray executor is required for multi-node deployments, and disaggregated serving expects additional connector infrastructure. None of this dims the engineering: this is the codebase that taught the industry how to serve LLMs.

Watch PyShine on YouTube

Contents