Before AI coding assistants consolidated into single-vendor products, Continue was one of the projects that proved an open source alternative could match them: the same agent engine available in VS Code, JetBrains, and the terminal, connected to whichever model you prefer. The maintainers describe it as a pioneering open-source coding agent, and the repository - Apache-2.0 and written almost entirely in TypeScript - now stands as a finished artifact after their final 2.0.0 release, kept public so others can build on it. That makes it a perfect subject for a source tour: complete, famous, and cleanly layered.

The architecture is the lesson here. One core package in core/ hosts the config loader, the tab-autocomplete engine, the codebase indexer, the MCP tool layer, and an LLM gateway that routes to 64 provider modules covering everything from Anthropic and OpenAI to Ollama on your laptop. Thin extensions in VS Code, JetBrains, and a CLI attach to that core, and a React GUI renders the chat panel. Codebase understanding comes from an indexing subsystem that chunks your repository, embeds it, and stores the vectors in a local LanceDB database, so retrieval stays on your machine.

As always in this series, this is an educational tour of how a well-known tool is built, not an endorsement of pasting proprietary code into any model. Continue’s own story makes the responsible-use point for us: the final release removed anonymous telemetry and pulled out authentication, and the codebase is offered as a foundation. Study the design, respect your provider’s terms, and keep your source code decisions your own.

Continue overview architecture diagram

Continue at a glance: three IDE surfaces on one core engine, an AI layer of providers and autocomplete, and a knowledge layer of context providers and LanceDB indexes.

Reading the overview from left to right:

Why You Need This

The first reason is the provider problem, solved more completely here than almost anywhere else. Supporting one LLM is a demo; supporting dozens is a subsystem. Continue’s core/llm/llms folder is exactly that subsystem, with a module each for Anthropic, OpenAI, Bedrock, VertexAI, Azure, Gemini, Groq, Mistral, Ollama, OpenRouter, Together, vLLM, xAI, WatsonX, and dozens more, all normalized by a shared gateway that handles streaming, token counting, FIM support detection, and error parsing once so each adapter stays tiny. If you are building anything that calls models, this is the best annotated example of the shape such a layer should take.

The second reason is that Continue shows how to give an assistant real knowledge of a codebase without a server. The indexing subsystem walks your repository, respects .continueignore, builds a refresh index of file changes, and materializes three kinds of knowledge: full-text search, a code-snippet index, and vector embeddings stored locally in LanceDB keyed by the embedding provider. Retrieval then happens through context providers, thirty of them, ranging from the current file and open editors to git commits, GitHub issues, Postgres schemas, terminal output, and MCP servers. Reading these two folders together teaches the full stack of practical retrieval-augmented generation for code.

The third reason is the multi-surface discipline. The same engine answers a slash command in the terminal, a chat message in the IDE panel, and an inline completion at your cursor, because every surface speaks one typed protocol to the same Core. The CLI in extensions/cli is especially instructive: sessions, slash commands, history compaction, and file tools show what an agent loop looks like when you strip the GUI away. Anyone planning a coding agent should study how little each surface actually has to implement.

How It Works

Continue detailed architecture diagram

Inside Continue: the surface adapters, the core hub with its messenger protocol and data layer, the config layer, the AI layer, indexing, context plus MCP, and the CLI agent loop.

Understanding the Architecture

Thin surfaces, one dynamic import. The VS Code entrypoint is deliberately two lines of real logic: it dynamically imports the activation module and hands it the extension context, which keeps the extension host startup fast. The JetBrains plugin and the CLI attach to the same Core, and the React GUI in gui/src communicates over the typed messenger protocol defined in core/protocol. The Core class in core/core.ts holds the five long-lived services: config handler, codebase indexer, completion provider, next-edit provider, and docs service.

A YAML config layer with a schema package. Configuration lives in config.yaml, validated and loaded by packages/config-yaml, which can also pull reusable blocks from a registry through its registry client. The core-side core/config handler merges that YAML with IDE settings and exposes a compiled config to every other service, so models, rules, context providers, and MCP servers are all declared in one place rather than scattered through UI options.

A gateway that speaks every provider. The abstract BaseLLM in core/llm/index.ts implements complete, chat, embed, rerank, and token counting once, using a tiktoken worker pool plus a bundled llama tokenizer. Concrete providers subclass it, and when a provider is OpenAI-shaped the gateway delegates to the shared adapters in packages/openai-adapters. Streaming is centralized in core/llm/streamChat.ts, which parses SSE chunks into a uniform event stream for chat and FIM completions alike.

Tab autocomplete as a staged assembly line. core/autocomplete/CompletionProvider.ts orchestrates the whole experience: a debouncer decides when to fire, prefiltering and snippet stages assemble context, the templating stage renders the fill-in-the-middle prompt, generation streams candidates, and post-processing cleans them with bracket matching and filtering before display. An LRU cache deduplicates requests, the default temperature is pinned to 0.01 for determinism, and every outcome flows into a logging service that tracks acceptances. The newer next-edit experience in core/nextEdit predicts where you will edit next rather than only what comes after the cursor.

Indexing that respects your ignore files. core/indexing/CodebaseIndexer.ts computes which indexes need building from a refresh index of file tags and times, then delegates to pluggable indexes: the LanceDB vector index in core/indexing/LanceDbIndex.ts, the code-snippets index, and full-text search. Ignore logic from .gitignore and .continueignore is shared across the whole engine through a single shouldIgnore utility, so what the indexer skips is exactly what the rest of the agent never sees.

Context providers and MCP as the tool surface. Every source of knowledge is a provider under core/context/providers: current file, file tree, open files, diff, git commits, GitHub issues, GitLab merge requests, docs, database schemas, Postgres, terminal contents, OS details, problems, debug locals, clipboard, URLs, web search, Google, Greptile, a repo map, and more. The MCP layer in core/context/mcp manages connections to MCP servers, including the OAuth handshake, and surfaces their tools, resources, and prompts to the agent and the chat.

The CLI agent loop. The CLI in extensions/cli/src/session.ts runs an agentic session against the same core: slash commands in extensions/cli/src/slashCommands.ts switch modes and manage context, file tools let the model read and edit, and a compaction module in extensions/cli/src/compaction.ts summarizes older turns so long sessions survive the context window. Onboarding, system prompt assembly, and authentication round out the loop.

The end-to-end flow. A user message enters through a surface, crosses the protocol into Core, gets compiled with retrieved context from the providers and indexes, and streams out through the gateway and the chosen provider module. Inline completions take the parallel autocomplete route, while agent tools execute through MCP. Every subsystem reads the same compiled config, which is what keeps three different IDEs behaving identically.

Advantages

  • True model freedom. 64 provider modules plus OpenAI-compatible adapters mean nearly any cloud or local model works, including through Ollama and vLLM.
  • Local-first codebase intelligence. LanceDB embeddings, full-text search, and snippets all live on your machine, so retrieval works offline and keeps source code private.
  • One engine, three surfaces. VS Code, JetBrains, and CLI share the core, so behavior stays consistent and each new surface is cheap to add.
  • MCP-native tooling. First-class MCP connection management, including OAuth, puts an entire ecosystem of tools one config block away.
  • Staged autocomplete design. Debouncing, prefiltering, templating, and post-processing as separate stages make each step tunable and testable.
  • A finished, readable codebase. The final release froze the project deliberately, which makes the repository a stable textbook rather than a moving target.

Benefits

  • No vendor lock-in. Switching models or providers is a config edit, not a migration, protecting your workflow from pricing and deprecation churn.
  • Privacy by architecture. Telemetry was removed and indexing stays local, which matters for teams with strict code confidentiality requirements.
  • Faster onboarding for agent builders. The monorepo layout lets you lift a single subsystem, such as the provider gateway or the MCP manager, into your own project.
  • Consistent experience across editors. Teams split between VS Code and JetBrains get the same agent instead of maintaining two configurations.
  • CLI automation ready. The agent loop with slash commands and file tools scripts cleanly for code review bots and repository maintenance tasks.
  • Educational value that compounds. Config schema, protocol, gateway, and indexing are each clean enough to teach from, which is rare at this scale.

Usage

The final release ships as a CLI on npm and as IDE extensions. From the README:

# install the VS Code extension from the marketplace
code --install-extension Continue.continue

# or use the CLI package
npm install -g @continuedev/cli

Configuration lives in a config.yaml under your .continue directory, where you declare models, rules, context providers, and MCP servers:

name: Local Assistant
version: 1.0.0
models:
  - name: claude
    provider: anthropic
    model: claude-sonnet-4-5
    apiKey: <key>
rules:
  - Prefer TypeScript strict mode idioms.

Launch the CLI in a repository and work through slash commands, or open the IDE panel and let tab autocomplete and the agent share the same configured models. Changes to config.yaml hot-reload, so adding an MCP server or swapping a model never requires a restart.

Conclusion

Continue is what an open-source coding agent looks like when it is taken to completion: a provider gateway wide enough to be future-proof, local indexing deep enough to be genuinely useful, and a core so cleanly separated that VS Code, JetBrains, and a CLI are all just clients. The maintainers have moved on, but they left the blueprint. Read the Core class, then the provider folder, then the indexer, and you will have absorbed more practical agent architecture than most courses teach.

Links:

Watch PyShine on YouTube

Contents