GitHub Copilot popularized AI code completion, but it also taught a generation of developers an uncomfortable lesson: your code leaves your machine. Tabby, built by TabbyML in Rust under an open source license, is the counter-argument - a self-hosted AI coding assistant that the README positions as an on-premises alternative to Copilot, self-contained with no database management system or cloud service required, running on consumer-grade GPUs. One binary serves completions and chat over an OpenAPI-compatible interface, indexes your repositories locally, and plugs into VS Code, JetBrains, and Vim through a shared agent.

The Rust engineering is what deserves the tour. The repository is a Cargo workspace of small crates with sharp boundaries: an axum HTTP server that mounts completion, chat, health, and event routes; an inference layer that dispatches to llama.cpp locally or OpenAI-style APIs remotely; a code-indexing subsystem that parses repositories with tree-sitter, stores full-text in Tantivy, and embeds what matters; and an enterprise layer in the ee/ folder that adds GraphQL, SQLite persistence, authentication, and an answer engine. Almost every layer can be read in an afternoon, and each one models a real-world decision about privacy, latency, and cost.

As with every tour in this series, the goal is education, not a claim that any tool makes your data magically safe. Tabby gives you the machinery to keep code and embeddings on hardware you control, but running it responsibly is still your job: pick models you trust, keep the indexed repositories authorized, and read the model specification before pointing the server at a GGUF file you downloaded from somewhere.

Tabby overview architecture diagram

Tabby at a glance: editor clients through a shared agent, an axum API with completion services, local and remote inference backends, a tree-sitter plus Tantivy knowledge layer, and an enterprise webserver.

Reading the overview from left to right:

Why You Need This

The first reason is sovereignty over the completion loop. Tabby shows exactly what a Copilot-style assistant needs to work: a fill-in-the-middle prompt builder, a streaming inference backend, a debounced client agent, and a small HTTP API between them. Because all four pieces live in one readable workspace, you can see the whole round trip from a keystroke in VS Code to a token produced by llama.cpp and back. If you have ever wondered where your code goes when you accept a suggestion, Tabby’s answer is a localhost port you control.

The second reason is the knowledge layer. Code completion alone forgets everything the moment it leaves your file, so Tabby pairs the model with an index: repositories are cloned by the git sync crate, parsed into tags with tree-sitter’s tag queries in crates/tabby-index/src/code/intelligence.rs, chunked and stored in Tantivy for full-text and lexical search, and embedded for semantic retrieval. The same machinery powers the answer engine in the enterprise layer, which is why chat can cite your actual code instead of guessing. This is one of the cleanest Rust examples of retrieval-augmented generation for source code.

The third reason is the model specification. Instead of hard-coding a zoo of providers, Tabby defines models as directories - a tabby.json describing a FIM prompt template like <PRE>{prefix}<SUF>{suffix}<MID> and an optional chat template, plus the GGUF weights in a ggml/ folder - documented in MODEL_SPEC.md. The download command fetches them with a parallel range-request downloader, and the device flag targets cpu, cuda, metal, or vulkan. For teams building anything model-adjacent, that little contract is a lesson in keeping infrastructure model-agnostic.

How It Works

Tabby detailed architecture diagram

Inside Tabby: the client stack, CLI entry, axum routes, services, inference backends, the indexing crates, and the enterprise webserver with its GraphQL schema and SQLite database.

Understanding the Architecture

A two-command CLI with a typed device list. The clap parser in crates/tabby/src/main.rs exposes Serve and Download subcommands, an OpenTelemetry endpoint flag, and an explicit Device enum covering cpu, cuda, metal, and vulkan. Serve arguments live in crates/tabby/src/serve.rs and download hands off to the aim-downloader crate for fast parallel weight fetching. Everything else hangs off those two entry points.

An axum server with a REST heart and a GraphQL shell. The routes module mounts completions, chat, health, events, and metrics endpoints on axum 0.8. The community server is the API surface editor extensions need; the enterprise webserver in ee/tabby-webserver/src embeds the same API and adds a juniper GraphQL schema from ee/tabby-schema/src/lib.rs with subscriptions, OpenAPI generation through utoipa, OIDC login via openidconnect, email through lettre, and WebSocket support. Persistence is SQLite accessed with sqlx, which is how Tabby honors its no-DBMS promise.

The completion service and its prompts. crates/tabby/src/services/completion.rs receives a completion request, assembles context from the code service, and renders one of two prompt builders: the FIM completion prompt or the newer next-edit prompt, both under crates/tabby/src/services/completion/. The request then streams through the inference dispatcher, and the service applies language-aware post-processing before responding, with events recorded for quality tracking.

Inference dispatch with pluggable backends. The dispatcher in crates/tabby-inference/src/lib.rs routes completion, code generation, chat, and embedding requests to a backend chosen by configuration: the bundled llama.cpp server process for local GGUF models, OpenAI-compatible HTTP bindings for remote endpoints, or Ollama bindings for locally managed models. The llama-cpp-server crate supervises the child process and translates tabby’s internal request types into llama.cpp calls, which keeps the GPU code out of the Rust tree.

Indexing as a scheduled background job. The indexer in crates/tabby-index/src/indexer.rs coordinates builds across the code index and the structured-document index. Repositories arrive through the git sync crate, documentation through the crawler. The code index extracts tags via tree-sitter, stores searchable documents in Tantivy through a shared utilities crate, and keeps per-repository snapshots so rebuilds are incremental. The embedding service feeds vector representations through the same inference layer, keeping all model access behind one door.

A shared agent for every editor. The clients folder contains a VS Code extension, a JetBrains plugin, a Vim plugin, and an Eclipse client, all delegating to the tabby-agent TypeScript package. The agent owns connection management, completion debouncing, and status reporting, and it renders chat through the tabby-chat-panel component, communicating across extension boundaries with the tabby-threads RPC library. Because the intelligence lives in the agent rather than in each editor SDK, the thinnest plugin still gets the full experience.

The end-to-end flow. A keystroke triggers the agent, which calls the completions endpoint; the route dispatches to the completion service, which builds a FIM prompt enriched with indexed context; the inference dispatcher streams tokens from llama.cpp or a remote API; and post-processed suggestions return over the same connection. Background jobs - repository sync, crawling, index builds, embeddings - run on their own schedule so the hot path never waits on the cold one.

Advantages

  • Single-binary self-hosting. SQLite storage and an embedded server mean no external database, message broker, or cloud account to stand up.
  • Consumer GPU support. Explicit cpu, cuda, metal, and vulkan targets plus llama.cpp serving run useful models on workstation hardware.
  • Rust performance discipline. Axum, tokio, and small crates keep memory low and startup fast, even with indexing jobs in the background.
  • Model-agnostic by design. The tabby.json model contract plus OpenAI-style HTTP bindings let you swap local GGUF models for hosted endpoints without code changes.
  • One agent, four editors. VS Code, JetBrains, Vim, and Eclipse share the tabby-agent core, so fixes propagate everywhere at once.
  • OpenAPI-first integration. The documented REST interface makes Tabby embeddable in cloud IDEs and internal tooling beyond the shipped clients.

Benefits

  • Code never leaves your network. Completions, embeddings, and indexes all live on infrastructure you run, which satisfies the strictest source-code policies.
  • Predictable cost. No per-seat subscription; the bill is hardware and electricity, and modest models remain genuinely useful for completion.
  • Latency you can tune. Local serving removes round trips to remote APIs, and the next-edit experience builds on that tight loop.
  • Searchable team knowledge. Indexed repositories, docs from the crawler, and the answer engine turn the server into a codebase Q&A endpoint, not just autocomplete.
  • A teaching-grade Rust codebase. Each crate demonstrates one concern - downloading, indexing, inference, serving - making the project an excellent blueprint for production Rust services.
  • Enterprise-ready when you need it. GraphQL, OIDC, email, and scheduled job notifications arrive in the ee/ layer without burdening the minimal community deployment.

Usage

The README’s one-minute start runs the Docker image with a completion model and a chat model:

docker run -it \
  --gpus all -p 8080:8080 -v $HOME/.tabby:/data \
  tabbyml/tabby \
  serve --model StarCoder-1B --device cuda --chat-model Qwen2-1.5B-Instruct

Bring your own weights by following the model directory contract from MODEL_SPEC.md:

my-model/
  tabby.json          # prompt_template for FIM, chat_template for instruct
  ggml/
    model-00001-of-00001.gguf

Then point an editor extension at http://localhost:8080 and start typing. Health, events, and metrics endpoints make the server easy to monitor, and the same CLI offers tabby download for pre-fetching models onto air-gapped machines.

Conclusion

Tabby proves that a self-hosted coding assistant does not have to be a compromise. The workspace layout shows genuine Rust craft - small crates, honest boundaries, supervised subprocesses - while the feature set covers the full loop from editor keystroke to indexed, cited answers on your own hardware. Read the completion service first, then the index crates, and you will understand both this tool and the anatomy of every code-completion product you will ever evaluate.

Links:

Watch PyShine on YouTube

Contents