Embeddings are only half of a semantic search system; the other half is a database that can store billions of them, filter them by arbitrary metadata, and return nearest neighbors in milliseconds under concurrent writes. Qdrant, an Apache-2.0 vector database written in Rust, is one of the most deployed answers to that problem, and its README makes the pitch concretely: store, search, and manage points β vectors with an attached JSON payload β with extended filtering support, hybrid search that fuses dense and sparse results through Reciprocal Rank Fusion or Distribution-Based Score Fusion, built-in quantization that cuts RAM usage by up to 97 percent, and horizontal scaling with sharding and replication and zero-downtime resizes. One docker command starts it, an OpenAPI 3.0 specification documents the REST surface, gRPC serves the production tier, and a built-in web UI lets you browse collections in a browser.
What makes the repository a rewarding read is that it is a textbook layered Rust workspace. The binary crate under src contains the API servers and the Raft consensus code, while the serious machinery lives in library crates under lib: collection for logical containers, shard for query planning and updates, segment for physical storage, plus wal, quantization, sparse, bm25, and gpu as focused side crates. The README notes that development happens on the dev branch and master carries releases, which is worth knowing before you browse. As always in this series, what follows is an educational tour of published source code.
Qdrant at a glance: startup mounts the REST and gRPC APIs and joins the Raft cluster, the dispatcher resolves each request to a collection via the content manager, collections fan out into shards, shards read and write segments, and every mutation is journaled through the write-ahead log before it lands.
Reading the overview from left to right:
- The process boots in src/startup.rs, mounting the Actix-based REST layer under src/actix and the Tonic gRPC services under src/tonic.
- Cluster membership and metadata replication live in src/consensus.rs.
- Both API shapes converge on the dispatcher at lib/storage/src/dispatcher.rs.
- The collections table β which collection exists, with what shard layout β is owned by the content manager in lib/storage/src/content_manager.
- The logical data path starts in the collection crate at lib/collection/src, which fans out into the shard crate at lib/shard/src.
- Physical storage is the segment crate at lib/segment/src, backed by the write-ahead log in lib/wal/src.
- Compression and keyword search are side crates: quantization at lib/quantization/src and sparse vectors at lib/sparse/src.
Why You Need This
The first reason is filtered vector search done properly, because nearest-neighbor search without metadata filtering is a demo, not a product. Qdrant attaches an arbitrary JSON payload to every point and lets you combine keyword matching, full-text, numeric ranges, and geo conditions with should, must, and must_not clauses β and crucially, the filter is applied inside the index traversal rather than as a post-filter that empties your result set. Query planning exploits stored payload indexes to pick an execution strategy, and faceting aggregates results by payload values. Multitenancy is treated as a first-class partitioning scheme, so one collection can serve thousands of users safely. This combination is precisely what recommendation, retrieval-augmented generation, and semantic cache workloads need, and it is why so many frameworks in this series β from LangChain to Mem0 β ship a Qdrant adapter.
The second reason is the storage engineering, which is where the Rust pays for itself. The segment crate under lib/segment/src separates vector storage, the HNSW graph index in lib/segment/src/index, payload storage, and the ID tracker that maps external point ids onto internal records; distance computations run through hand-tuned SIMD kernels in lib/segment/src/spaces for x86-64 and ARM Neon. Quantization compresses vectors β scalar, binary, and product variants β so a machine with a fraction of the RAM can still serve the full corpus, and GPU support in lib/gpu/src accelerates index construction for NVIDIA and AMD hardware. Durability comes from the write-ahead log in lib/wal/src, which the README credits with surviving power outages, plus io_uring-based asynchronous disk I/O for throughput on network storage.
The third reason is distributed operation without downtime. The consensus code in src/consensus.rs replicates collection metadata through Raft, while data itself is sharded and replicated across peers by the machinery in the collection and shard crates; resharding and replica moves are orchestrated as background operations rather than export-import rituals. Background optimizers in lib/shard/src/optimizers rebuild segments and merge them continuously, and during a rebuild the proxy segment in lib/shard/src/proxy_segment forwards writes so nothing is lost mid-optimization. That combination β live traffic while the index reorganizes underneath it β is the hard part of running a database, and Qdrant has it as a designed-in behavior rather than an operations manual.
The detail view: startup, settings, REST and gRPC handlers with auth and the web UI at the top, the control plane of dispatcher, content manager, and RBAC, the sharding core of collections, registries, update workers, optimizers, and query modules, the segment internals from vectors to SIMD, and the WAL and accelerator crates below.
Walking the detail diagram from the API tier, both protocol stacks are thin. The Actix handlers in src/actix/api and the Tonic services in src/tonic/api parse requests, verify them through src/actix/auth.rs β which consults role-based access control from lib/storage/src/rbac β and hand the operation to the dispatcher in lib/storage/src/dispatcher.rs. The built-in console served by src/actix/web_ui.rs reuses the same handlers, which is why the web UI can browse and mutate real collections. Configuration comes from the YAML loader in src/settings.rs.
The control plane and data plane meet in the collection abstraction. Cluster-committed changes arrive at the content manager in lib/storage/src/content_manager from the Raft member in src/consensus.rs, typed as the operation models under lib/collection/src/operations. A collection instance in lib/collection/src/collection delegates shard placement to lib/collection/src/collection_manager and keeps local and remote replicas in lib/collection/src/shards, while writes flow through the parallel executors of lib/collection/src/update_workers. Queries compile against the DSL in lib/shard/src/query and execute through the fan-in logic of lib/shard/src/retrieve.
The segment layer is where vectors actually live. Each shard keeps its working set in lib/shard/src/segment_holder, which hands out locked handles to segment facades from lib/segment/src/entry backed by the appendable implementation in lib/segment/src/segment; construction and version migration happen in lib/segment/src/segment_constructor. Inside a segment, vectors go to lib/segment/src/vector_storage, graph search to lib/segment/src/index, payloads to lib/segment/src/payload_storage, and point identity to lib/segment/src/id_tracker. Every mutation is journaled to the WAL first and replayed on boot, and sparse search pairs the inverted indexes of lib/sparse/src with the BM25 scoring of lib/bm25/src for hybrid retrieval.
From Install to a Filtered Search
The on-ramp is one command: docker run -p 6333:6333 qdrant/qdrant, which starts an unauthenticated instance meant for experimentation β the README points to its security guide before production. From there, any of the official clients in Python, Go, Rust, JavaScript, .NET, or Java talks REST or gRPC: create a collection with a vector size and distance function, upsert points with ids, vectors, and payloads, then search with a query vector plus a filter block. The interesting knobs come next β enable quantization on the collection for a fraction of the memory, add payload indexes for the fields you filter on, configure hybrid queries that fuse dense and sparse scores, and shard across a cluster of peers when one machine stops being enough. The web UI on the same port lets you watch collections, points, and metrics while you do it.
Honest limits: this is a systems codebase, and the Rust workspace rewards readers who are comfortable with traits, lifetimes, and lock hierarchies β the segment facade alone exists to keep concurrency sane. Operational complexity is real once you add sharding and replication, and the teamβs development branch is dev while master carries releases, so fresh features appear there first. Quantization trades recall for memory, and the right configuration is workload-specific despite the seductive headline numbers. But as a study of how a modern vector database is actually built β API surface, consensus, sharding, indexing, durability β qdrant/qdrant is one of the most complete open-source references you can read. Enjoyed this post? Never miss out on future posts by following us