Cline started life as a VS Code extension and grew into one of the most popular open source coding agents, and its repository has now been rebuilt into something bigger: a layered TypeScript monorepo where one agent core serves an IDE extension, a JetBrains plugin, a terminal CLI, and a desktop hub. The Apache-2.0 licensed codebase is organized around a Bun workspace whose root package.json pulls together two families of code: an SDK under sdk/packages and a set of host applications under apps. The promise spelled out in the SDK’s ARCHITECTURE.md is simple to state and hard to deliver: plan and act modes, MCP servers, checkpoints, rules, and provider configuration behave identically no matter which surface you drive the agent from.

The layering is the interesting part. At the bottom sits a shared package of contracts and schemas, above it a model layer that isolates every provider behind one gateway, above that a stateless agent loop, and on top a core package that owns everything stateful: sessions, checkpoints, settings, a WebSocket hub, and scheduled routines. Hosts such as the terminal UI or the editor extension never touch providers directly; they talk to the core, the core starts the loop, and the loop calls the gateway. Reading the source top to bottom feels less like browsing an app and more like studying a deliberately stratified system where each floor can only call the floor beneath it.

As with every project in this series, this is an educational tour of published source code, not an endorsement of handing an agent unsupervised access to your machine. Cline runs commands and edits files on your behalf, and its own design answers that risk with approval gates, a subprocess sandbox, and human-in-the-loop checkpoints. If you study or deploy it, do so the way its authors intend: on work you own, with the approval settings you actually understand, and with the same care you would give any tool that can rewrite your repository.

Cline overview architecture diagram

Cline at a glance: a Bun workspace binding four SDK packages and three host apps, with one agent loop fed by a gateway of eleven provider vendors.

Reading the overview from left to right:

Why You Need This

The first reason is architectural literacy. Most coding agents grow by accretion, and you feel it the moment you try to reuse their internals. Cline took the opposite path: its ARCHITECTURE.md is a genuine source of truth that names each package’s responsibility and its dependency direction. Shared may not depend on anything above it, agents may not own storage, and hosts must route settings changes through core services rather than writing files themselves. If you have ever wondered what a clean layering looks like for an LLM application, this repository is a working answer you can read in an afternoon.

The second reason is provider portability done right. The model layer hides eleven vendor implementations behind a single gateway contract, so the agent loop never learns whether your model lives on Anthropic, OpenAI, AWS Bedrock, Google Vertex, Mistral, or a local Ollama server. There is even a community vendor for OpenAI-compatible endpoints and a first-party cline vendor for the project’s own account service. Handler creation flows through a factory registry, retries and error classification are shared, and model metadata comes from a dedicated catalog. Any application you write that calls more than one LLM vendor will eventually need exactly this seam.

The third reason is the hub model of agent infrastructure. Cline’s core can run its sessions inside a detached hub daemon, let multiple clients attach and detach without killing the run, and expose the whole thing over a WebSocket protocol with explicit run, approval, and schedule commands. That is the architecture of a team tool, not a toy: one workspace, one authority runtime, and many windows into it. Studying how the daemon publishes discovery records, how clients register their version and identity, and how detached shell commands keep draining into capped logs will teach you patterns that go far beyond coding agents.

How It Works

Cline detailed architecture diagram

Inside cline: host surfaces, the core facade, runtime host selection, the agent loop with its approval and sandbox gates, the model layer, and the hub services.

Understanding the Architecture

A monorepo with a strict ladder. The root package.json names six SDK packages, three app workspaces, and a set of scripts that build, typecheck, and test everything through Bun. The SDK carries a uniform version, 0.0.91 at the time of writing, while the CLI app sits at 3.0.69 and ships platform binaries for macOS, Linux, and Windows on arm64 and x64. The dependency rule is enforced by review rather than folklore: llms consumes shared, agents consumes llms and shared, core consumes all three, and apps consume only core. Nothing at the bottom knows the name of anything at the top.

A facade that hosts actually see. Hosts construct a ClineCore through the core package’s entrypoint, and the facade normalizes broad host configuration into a runtime session config before any agent exists. The core then picks a runtime host: a local in-process host, a hub-backed host that talks to a shared daemon, or a remote host for cloud sessions. Telemetry, settings mutation, cron routines, plugin installation, and marketplace features all hang off this facade, which is why a VS Code extension and a terminal process can expose the same capabilities with different chrome.

A stateless loop with explicit seams. The agent runtime in the agents package is deliberately free of persistence: it prepares each model request, runs hooks before and after the model call, executes tools, emits structured runtime events, and returns a run result. Hook interfaces named before-model, before-tool, and after-tool-result give hosts and extensions observation points without wrapping transports. When a session ends, the local runtime looks for an explicit completion tool call and emits exactly one task-completed event, with a teardown fallback for sessions that finish cleanly without one, so hosts never double-count a finished task.

A gateway with eleven vendors and one contract. The llms package resolves provider settings, consults its model catalog, and builds a handler through the factory registry. The vendors directory holds implementations for anthropic, openai, openai-compatible, bedrock, vertex, google, mistral, ollama, minimax-thinking, community, and cline, each with its own wire-format tests. Execution funnels through an AI SDK integration that normalizes streaming, tool calls, reasoning tokens, and request IDs, and a separate transcription service handles audio input. Error classification decides what is retryable, which is what keeps long agent runs alive across provider hiccups.

Nine built-in tools behind one approval gate. The core package assembles the default tool set from definitions.ts: read_files, search_codebase, run_commands, fetch_web_content, apply_patch, an editor tool, a skills loader, ask_question for clarifications, and submit_and_exit, which is the SDK’s successor to the classic attempt_completion. Every call passes through a tool-approval controller on the local runtime, and run_commands executes inside a subprocess sandbox with lifecycle tracking. Extensions registered through the shared extension registry add their own tools to the same flow, which is also how MCP servers become first-class tools.

Sessions that survive anything. The session directory owns stores, snapshots, versioning, and checkpoint restore, so a conversation can be forked, restored to an earlier checkpoint, or handed to a different host without losing state. Checkpoint diffing and restore are first-class files, not afterthoughts, and history origin tracking records where each session’s transcript came from. The design rule that hosts retain title and transcript policies while core owns fork ancestry keeps the boundary honest.

A hub for shared authority. Under core’s hub directory, a WebSocket server brokers commands and events, a daemon entrypoint runs detached from any editor, and a discovery package publishes endpoint records so clients can find a compatible hub. Client adapters named NodeHubClient, HubSessionClient, and HubUIClient translate the socket protocol into host-facing APIs. Session status is reported, never fabricated: prompt-bearing starts begin running, idle interactive starts stay idle until their first turn, and detached shell commands drain into size-capped logs that a reconciliation pass keeps honest even after the launching process exits.

Terminal and editor surfaces. The CLI app builds a full terminal UI with plan and act toggles, slash commands, file mentions, and live tool approvals, plus an ACP adapter for editor integration and a headless JSON mode that streams NDJSON events for scripting. The VS Code app separates hosts from integrations, with the host layer bridging editor state into core and the integrations layer handling editor-specific observations. The desktop hub app wraps the same hub server behind its own webview. Three surfaces, one behavior contract.

End to end. A prompt typed in any host reaches the core facade, which starts or attaches to a runtime, which drives the agent loop, which calls models through the gateway and executes tools through the approval gate and sandbox. Every turn is persisted by the session layer, every event can be relayed through the hub, and every host sees the same stream of structured events. That is the whole system, and its consistency across surfaces is exactly what the layering buys.

Advantages

  • True multi-surface consistency. The same core powers CLI, VS Code, JetBrains, and desktop, so behavior you verify in one place holds in the others.
  • Provider freedom without glue code. Eleven vendor handlers behind one gateway mean switching models is a settings change, not a refactor.
  • Safety designed in, not bolted on. Approval gates, a subprocess sandbox, and checkpoint restore are part of the core tool path, not optional decorations.
  • Hub-ready architecture. A detached daemon, discovery records, and multi-client attach make team and long-running use cases first-class citizens.
  • Readable source of truth. The SDK ARCHITECTURE.md documents package boundaries and design rules, which is rare candor for a project this size.
  • Extensible at the seams. Hooks and an extension registry let you observe and extend the loop without forking the runtime.

Benefits

  • Learn a reference layering. The shared-llms-agents-core ladder is a template you can lift into any LLM application that outgrows a single script.
  • Write provider-agnostic tools. Studying the gateway and factory registry shows how to normalize streaming, retries, and error classes across vendors.
  • Borrow the event discipline. Structured runtime events with explicit lifecycle boundaries are reusable in any UI that renders agent progress.
  • Understand agent safety engineering. Approval flows, sandboxed command execution, and checkpoint restore are patterns every agent builder eventually needs.
  • Automate with confidence. The headless NDJSON mode and cron routines show what scripted, unattended agent runs look like when designed deliberately.
  • Study a real hub protocol. The WebSocket command, event, and reconciliation design is a free education in multi-client runtime plumbing.

Usage

Install the CLI from npm as the README directs:

npm install -g cline

Run interactively, one-shot, or with piped input:

cline
cline "Audit this package and propose fixes"
cat file.txt | cline "Summarize this"

Sign in interactively or wire a provider straight from flags:

cline auth
cline auth cline
cline auth --provider anthropic --apikey sk-... --modelid claude-sonnet-4-6

For scripting and CI, stream structured events instead of rendering the TUI:

cline --json "Refactor the config loader and run the tests"

Nightly builds are available as cline@nightly, and cline --help prints the full flag reference.

Conclusion

Cline’s rebuild is a case study in scaling an open source agent past its original shell. The repository now reads as four clean layers and three thin surfaces, with one loop, one gateway, one session model, and one hub protocol shared by all of them. The payoff is practical: provider changes are configuration, new hosts are thin, and the safety machinery sits on the same path as every tool call. Read it for the layering, borrow the gateway and hook seams for your own agents, and drive it, as its authors intend, with approvals on and checkpoints within reach.

Links:

Watch PyShine on YouTube

Contents