Serving systems usually rediscover the same insight over and over: most tokens you generate have been generated before. SGLang, Apache-2.0-licensed and hosted by the LMSYS non-profit, turned that insight into its founding idea. The project grew out of the RadixAttention research line, which stores attention key and value tensors in a radix tree keyed by token prefixes so that any request sharing a prefix with past work reuses the computed pages instead of paying for them again. Today the README describes SGLang as an open-source inference framework for large language, vision-language, and diffusion models, optimized for agentic workloads, large-scale serving, and RL rollouts, and the code delivers all three: the same repository contains the SGLang Diffusion engine for image and video generation, a Rust model gateway for cluster routing, EAGLE and n-gram speculative decoding, prefill and decode disaggregation, and deep hooks for reinforcement-learning frameworks like slime, verl, and AReaL that use SGLang for rollout generation. Hardware reach is unusually wide for a serving engine: NVIDIA from B200 and GB200 down to A100, AMD Instinct MI300X and MI355X, Google TPUs, Intel Arc GPUs and Xeon CPUs, Apple Silicon through Metal and MLX, Huawei Ascend NPUs, and Moore Threads GPUs.
The repository splits into Python orchestration and a Rust edge. The python/sglang package is the installable core, with the srt directory, the SGLang runtime, holding almost everything that matters: managers for the multi-process control plane, mem_cache for the radix tree and allocators, model_executor for device work, layers for attention, mixture-of-experts, and quantization kernels, plus speculative, constrained, function_call, and disaggregation modules. The sgl-model-gateway directory at the root is the Rust router, and proto carries gRPC definitions. The build pins torch 2.14.1, supports Python 3.10 and newer, and uses setuptools-scm for versioning and setuptools-rust for the gateway side. The README’s own acknowledgment section credits Guidance, vLLM, LightLLM, FlashInfer, Outlines, and LMQL as design influences, which tells you exactly which ideas this codebase digested. As always in this series, what follows is an educational tour of published source code.
SGLang at a glance: the FastAPI server and OpenAI-compatible adapters feed the TokenizerManager, the DataParallel controller fans requests to scheduler replicas, cache-aware policies match prefixes against the RadixAttention tree before every batch, and the TP worker drives the ModelRunner over fused attention and MoE layers.
Reading the overview from left to right:
- Serving starts at the FastAPI application in python/sglang/srt/entrypoints/http_server.py, which mounts the OpenAI-compatible routers under python/sglang/srt/entrypoints/openai.
- The offline-programmatic path is the Engine class at python/sglang/srt/entrypoints/engine.py, which reuses the same machinery without HTTP.
- Requests land in the TokenizerManager at python/sglang/srt/managers/tokenizer_manager.py, the async front-end that tokenizes, tracks, and answers every call.
- The DataParallel controller at python/sglang/srt/managers/data_parallel_controller.py balances load across scheduler replicas.
- Each replica runs the Scheduler at python/sglang/srt/managers/scheduler.py, which orders work through the cache-aware policies in python/sglang/srt/managers/schedule_policy.py.
- Prefix reuse comes from the RadixAttention tree at python/sglang/srt/mem_cache/radix_cache.py, backed by token allocators and eviction policies.
- Batches execute in the TP worker at python/sglang/srt/managers/tp_worker.py, which owns the ModelRunner at python/sglang/srt/model_executor/model_runner.py.
- The runner composes architectures from python/sglang/srt/models on top of the attention, MoE, and quantization modules in python/sglang/srt/layers.
Why You Need This
The first reason is the prefix cache, and it is worth understanding at the data-structure level. The radix tree in python/sglang/srt/mem_cache/radix_cache.py maps token sequences to physical KV pages, so a hundred users asking about the same document after the same system prompt hit one cached path. The allocators under python/sglang/srt/mem_cache/allocator hand out paged token slots, and the eviction policies at python/sglang/srt/mem_cache/evict_policy.py decide which leaves to drop under pressure. What makes this operational rather than academic is the scheduler integration: the policies in python/sglang/srt/managers/schedule_policy.py sort waiting requests by how much cache they would reuse, which means the batch you run is the batch the tree wants. For multi-turn chat and agent loops, this is the difference between paying for prefill once and paying for it on every single call.
The second reason is the depth of the execution tricks, all wired into one loop. The overlap modules under python/sglang/srt/batch_overlap hide CPU scheduling cost behind GPU work; the decode CUDA graph runner at python/sglang/srt/model_executor/runner/decode_cuda_graph_runner.py replays captured graphs for steady-state tokens, with a prefill sibling beside it; and the speculative module under python/sglang/srt/speculative implements EAGLE-style drafting, an n-gram drafter with C++ helpers, and the DFlash variants, each verified by the target model in one pass. Prefill and decode disaggregation under python/sglang/srt/disaggregation lets you dedicate one GPU pool to prompts and another to generation, shipping KV tensors between them. These are the levers that separate a demo server from production throughput.
The third reason is the breadth of surfaces and hardware. The entrypoints directory ships an OpenAI-compatible API, Anthropic and Ollama adapters, a gRPC server, and the in-process Engine, so client code rarely needs rewriting to try SGLang. The tool-call story is handled by the parsers under python/sglang/srt/function_call, constrained generation by the grammar modules in python/sglang/srt/constrained, and multi-adapter serving by python/sglang/srt/lora. The hardware matrix in the README covers seven vendor families, and the platform code under srt reflects it, which makes the project a practical choice when your fleet is not uniform. Add the Rust gateway for routing across many engines and you have a complete serving tier in one repository.
The detail view: the serving front-end with launch entry, HTTP server, API adapters, gRPC, config validation, and the Rust gateway on the left; manager processes with tokenizer, ZMQ message structs, detokenizer, and DataParallel control in the middle; scheduling and cache internals with the radix tree, allocators, overlap, and disaggregation below; and the execution layer with the TP worker, ModelRunner, CUDA graphs, model zoo, layers, and collectives on the right.
Walking the detail diagram, the front end is wider than the README suggests. The bootstrap starts at python/sglang/launch_server.py, which validates hundreds of options through ServerArgs at python/sglang/srt/server_args.py before anything boots, and the Rust model gateway under sgl-model-gateway sits in front of engine replicas for cluster deployments. Between client and engine, the TokenizerManager at python/sglang/srt/managers/tokenizer_manager.py speaks the ZMQ message vocabulary defined at python/sglang/srt/managers/io_struct.py, supports many tokenizers at once through the mixin at python/sglang/srt/managers/multi_tokenizer_mixin.py, and receives detokenized text back from the DetokenizerManager at python/sglang/srt/managers/detokenizer_manager.py.
In the execution layer, the TP worker at python/sglang/srt/managers/tp_worker.py prepares a ForwardBatch, whose fields live at python/sglang/srt/model_executor/forward_batch_info.py, and hands it to the ModelRunner at python/sglang/srt/model_executor/model_runner.py. The runner chooses between eager execution and the captured graph runners under python/sglang/srt/model_executor/runner, coordinates tensor and expert parallelism through the collectives in python/sglang/srt/distributed, and delegates weight math to the quantization methods under python/sglang/srt/layers/quantization and the MoE dispatch code in python/sglang/srt/layers/moe.
From Install to Serving
The README’s recommended route is the Docker image: docker pull lmsysorg/sglang:latest, which carries the torch and CUDA stack prebuilt, and the development guide expects the lmsysorg/sglang:dev container with an editable install of the python directory. If you prefer a bare environment, the package installs with uv pip install –prerelease=allow sglang, per the README, and the environment contract in pyproject.toml is explicit: Python 3.10 or newer and a build pinned to torch 2.14.1. Launching is one command, python -m sglang.launch_server, which resolves to the bootstrap at python/sglang/launch_server.py and brings up the HTTP server, the TokenizerManager, and the scheduler processes for your chosen model and parallelism flags. From the first request onward, the radix tree starts filling, and the cache-aware scheduler quietly reuses every prefix it has seen.
Honest limits: the recommended install path is a multi-gigabyte Docker image, and a source environment inherits a pinned torch plus CUDA dependencies that make lightweight setups awkward; the cluster-grade features, disaggregation, the DataParallel controller, and elastic expert parallelism, assume real multi-GPU infrastructure and configuration effort; and parts of the tree are explicitly experimental, which is the cost of a project that ships diffusion, audio, and speculative research directions at once. None of that changes the core verdict: the radix cache and the scheduler around it are some of the most instructive serving code being published today. Enjoyed this post? Never miss out on future posts by following us