Some databases earn adoption by being effortless, and Chroma is the clearest example in the vector search world. Its README reduces the entire workflow to four functions — create a collection, add documents with metadata, query by text, and get by id — with tokenization, embedding, and indexing handled automatically unless you bring your own vectors. Chroma, Apache-2.0 licensed and installable with pip or npm, runs three ways from one codebase: embedded in your Python process for prototyping, as a single-node server started with the run command for a team, and as Chroma Cloud, the hosted service that the open-source core is increasingly built to power. That last point is what makes the repository unusually interesting: it is mid-transition into a polyglot system where a Rust workspace of some forty crates under rust coexists with the classic Python engine under chromadb.

Reading it means holding two architectures in mind at once. The Python side is a self-contained database: a FastAPI server, SQLite metadata, in-process HNSW segments, and a component framework that wires everything together from a config. The Rust side is the distributed future: query and compaction services, a replicated log called WAL3, S3-style block storage, and protobuf contracts in idl that both languages share. As always in this series, what follows is an educational tour of published source code.

Chroma overview architecture diagram

Chroma at a glance: Python and JavaScript clients talk to the FastAPI server, which delegates to the embedded SegmentAPI engine with SQLite metadata and local HNSW segments; the Rust core runs query and compaction services over the log service and block storage for the distributed deployment of the same data model.

Reading the overview from left to right:

Why You Need This

The first reason is the embedded path, which removes the entire operations conversation from prototyping. Import chromadb, construct a client, and you have a working vector database inside your notebook: the component framework in chromadb/app.py assembles the system from typed definitions in chromadb/config.py, the SegmentAPI in chromadb/api/segment.py implements the same API the server exposes, and persistence is just a directory. When your prototype becomes a product, the same code flips to server mode without a rewrite, and when traffic outgrows one box, the managed cloud takes over. Very few databases offer a genuinely continuous path from import statement to production cluster.

The second reason is that the source is a complete lesson in how a vector database is actually layered. Metadata lives in SQLite through the system database interface at chromadb/db/system.py and its implementation in chromadb/db/impl/sqlite.py; vector data lives in segments, created and managed by chromadb/segment/impl/manager/local.py, with the HNSW graph in chromadb/segment/impl/vector/local_hnsw.py and its memory-mapped persistence sibling at chromadb/segment/impl/vector/local_persistent_hnsw.py. Queries are planned and filtered through the executor modules in chromadb/execution/executor and the expression trees of chromadb/execution/expression, which is where the where and where_document filter clauses of the README become an evaluated AST. The seam between logical collections and physical segments is the abstraction that lets one codebase serve laptop and cluster.

The third reason is the Rust core, which shows where the project is going. The workspace under rust contains the query frontend in rust/frontend/src, service binaries in rust/worker/src/bin for query serving, compaction, and the work queue, and the compactor in rust/worker/src/compactor that converts log records into queryable segments and indexes from rust/segment/src and rust/index/src. Durability is delegated to WAL3 in rust/wal3/src and cloud files in the block storage of rust/blockstore/src, with distance kernels in rust/distance/src doing the numerical work. For teams evaluating infrastructure bets, watching this half of the repository is a preview of the next five years of the product.

Chroma detail architecture diagram

The detail view: clients and the FastAPI server on the left, the Python engine of system database, segment managers, HNSW, and execution below it, the Rust services of frontend, workers, compactor, and log service in the middle, the data path of WAL3, segments, indexes, storage, and distance kernels on the right, and the PyO3 bridge and protobuf IDL as shared contracts.

Walking the detail diagram from the client side, the transport layer is a lesson in itself. The synchronous client at chromadb/api/client.py and the asyncio variant at chromadb/api/async_client.py both speak through the FastAPI adapter in chromadb/api/fastapi.py, while the server in chromadb/server/fastapi/init.py mounts the same routes over the assembled System built in chromadb/app.py. Because the API surface is shared, embedding Chroma in-process or pointing at a server changes one import, not one line of application logic. Cross-cutting concerns are wired at assembly time: quota enforcement from chromadb/quota/simple_quota_enforcer and product telemetry via chromadb/telemetry/product/posthog.py are components like any other.

The embedded engine is where reads and writes actually execute. The SegmentAPI at chromadb/api/segment.py resolves collections through the metadata system in chromadb/db/system.py backed by chromadb/db/impl/sqlite.py, routes vectors through the local manager to chromadb/segment/impl/vector/local_hnsw.py, and keeps attribute records in the metadata segment modules of chromadb/segment/impl/metadata. The distributed manager in chromadb/segment/impl/manager/distributed.py implements the same interfaces against remote segments, which is how the Python engine can front a Rust cluster. Query execution composes the plan builders of chromadb/execution/executor with the filter evaluation of chromadb/execution/expression.

The Rust side mirrors the responsibilities at cluster scale. The system assembly in rust/chroma/src wires components through the registry in rust/system/src; the frontend in rust/frontend/src resolves topology via rust/sysdb/src and tails writes from the log service in rust/log-service/src, which appends to the replicated WAL3 in rust/wal3/src. The compactor at rust/worker/src/compactor drains the log into segments and indexes, flushing through the object storage abstraction in rust/storage/src and the chunked IO of rust/blockstore/src. The bridge back to Python is the PyO3-based client at chromadb/api/rust.py, and both languages share wire types generated from the protobuf definitions at idl/chromadb/proto/chroma.proto.

From Install to a Working Collection

The zero-config path is the README’s own snippet: pip install chromadb, construct the default client, create a collection, and add documents — raw text, optional metadata, and ids — with embeddings computed for you. Query with text and a result count, add where filters on metadata or document contents, and when you want durability pass a path or start the server with the run command and connect over HTTP. The asynchronous client exists for applications that need concurrent requests, and the JavaScript SDK mirrors the surface for Node and the browser. Scaling up means switching to server mode behind your application, and the distributed Rust deployment is what powers the hosted offering. Throughout, the same four operations carry the day, which is the design’s whole point.

Honest limits: the repository is genuinely two codebases in transition, and the Python engine remains the simpler, better-documented path while much of the Rust core targets the cloud product rather than casual self-hosting. Version migrations between storage formats appear in chromadb/migrations/metadb and are worth respecting before upgrading a persistent directory. And the four-function API’s simplicity has edges — complex hybrid search and multi-stage retrieval may push you toward lower-level configuration faster than expected. But as a study of database architecture at two scales sharing one contract, chroma-core/chroma is one of the most instructive codebases in this series.

Watch PyShine on YouTube

Contents