Running an LLM on your own machine used to mean choosing between a CLI tool and a science project. Jan, self-described as an open-source ChatGPT replacement, occupies the middle ground: a real desktop product, downloadable from the Microsoft Store, Flathub, and its own site, that runs Llama, Gemma, Qwen, and GPT-oss style models from HuggingFace locally, talks to cloud providers like OpenAI, Anthropic, Mistral, Groq, and MiniMax when you let it, exposes an OpenAI-compatible API on localhost:1337 for other applications, and integrates MCP so the assistant can actually do things rather than just talk. The repository at janhq/jan is Apache-licensed and currently at version 0.8.5, and the striking thing about its source tree is that it is a second act: the app was famously Electron-based, and the current main branch is a full rebuild on Tauri, with the acknowledgements page naming the three giants it stands on, llama.cpp, Tauri, and Scalar.

The rebuild is worth studying precisely because it is a rebuild. Desktop AI apps sit at a messy intersection of concerns, a React UI that must feel native, a local inference engine that must speak GPU dialects, a server that must look like OpenAI’s, an agent layer that must be sandboxed, and an extension system so the community can add engines and behaviors without forking the app. Jan v2 answers each concern in its own crate or package, and the boundaries are clean enough that you can trace one chat message from a React container through a TypeScript browser API, across the Tauri IPC bridge, into Rust session state, out to a spawned llama.cpp worker process, and back again. Node.js 20 and up, Yarn 4, Make, and Rust are the build prerequisites, and make dev is the documented entry point. As always in this series, what follows is an educational tour of published source code.

Jan overview architecture diagram

Jan at a glance: the React web-app calls @janhq/core, which crosses the Tauri IPC into the Rust app shell; the core modules run threads, the agent loop, MCP clients, and the local server; engines load through the llamacpp plugin into the jan-llama-worker, with MLX as the Apple Silicon path; extensions register through the core API and the agent-tools plugin supplies sandboxed capabilities.

Reading the overview from left to right:

Why You Need This

The first reason is that Jan shows how to isolate native inference safely, and the answer is a worker process, not a library call. The llamacpp plugin does not link llama.cpp into the app binary; it vendors the engine source, pinned by the plugin’s build script, and compiles a separate jan-llama-worker binary whose GPU backends are selected by a single build token, with Vulkan always present as the fallback, CPU compiled as every x86-64 and ARM microarchitecture variant and scored at load time, and CUDA 12, CUDA 13, Metal, and HIP as optional add-ons. The plugin’s Rust side, with its engine session management in src-tauri/plugins/tauri-plugin-llamacpp/src/engine and GGUF metadata parsing in src-tauri/plugins/tauri-plugin-llamacpp/src/gguf, talks to that worker across a process boundary, so a crashed model load takes down a worker, not the whole chat app. Apple Silicon gets a parallel path: the MLX plugin plus an mlx-server written in Swift, with the Metal shaders compiled by Xcode’s PrepareMetalShaders step rather than a plain swift build, a detail the Makefile documents after clearly being learned the hard way.

The second reason is the agent architecture, which has grown past simple chat into something you can borrow ideas from wholesale. The agent module in src-tauri/src/core/agent is a directory of single-purpose Rust files: a loop that drives turns, subagents with a builtin_subagents folder, skills, memory, todos, plans, sessions, transcripts, and compaction, the context-squeezing step long conversations eventually need, plus OpenTelemetry hooks in an otel folder. Tool execution goes through the agent-tools plugin, whose lib wires skills, permissions, memory, workspace, and preview, with the concrete built-ins in src-tauri/plugins/tauri-plugin-agent-tools/src/tools; MCP servers registered through the MCP module extend that tool surface with external capabilities. The same agent core is reachable outside the GUI: the jan-cli crate builds a standalone jan binary, and the repository commits a protocol schema plus an RPC schema with a Make target whose only job is to fail the build when Rust types and the committed schema drift apart, and it speaks the Agent Client Protocol so other editors can drive Jan’s agent directly.

The third reason is the extension seam, which is the difference between an app and a platform. The TypeScript side defines the contract in @janhq/core: an extension manager in core/src/browser/extension.ts that loads and registers packages, an inference glue layer in core/src/browser/extensions/inference.ts that routes chat requests to whichever engine backs them, and packages under extensions for assistants, conversation handling, downloads, llamacpp, mlx, RAG, and vector storage. For third-party developers the same surface is exported as the Agent Development Kit in packages/adk/src/index.js, with RPC types generated from the Rust protocol so agents built against the ADK stay compatible with the app. Because the local server module in src-tauri/src/core/server exposes the OpenAI-compatible API on port 1337, the ADK and any OpenAI client can treat a running Jan instance as just another endpoint.

Jan detail architecture diagram

The detail view: web frontend and @janhq/core on the left, the Tauri core modules in the middle, agent tools and the engine stack on the right, the extension layer bridging them, and memory, search, and hardware plugins underneath.

Walk the detail diagram and the frontend is deliberately boring, which is a compliment. React containers in web-app/src/containers read state from the stores in web-app/src/stores, and both reach native capability only through @janhq/core, so no component knows whether a model list came from a local folder scan or a cloud provider. The Rust shell in src-tauri/src/lib.rs wires the plugin registry at startup, and the core modules own their domains: threads own chat sessions, the agent owns turns and tools, downloads own model fetching with the hardware plugin supplying GPU detection so the app can recommend engine variants, and the server module re-exports the session state as OpenAI-shaped routes so the ADK, curl, and other apps share one surface.

The engine stack is where the platform thinking pays off. The llamacpp plugin parses GGUF metadata to learn a model’s context length and chat template before starting anything, spins the worker with the backend set baked at build time, and keeps sessions alive across chat turns; the MLX plugin mirrors the same contract for Apple’s framework, which is why the inference glue can switch between them without the UI changing. The RAG and vector-db plugins pair with their TypeScript counterparts to give assistants document memory, and the websearch plugin gives the agent an escape hatch beyond local knowledge. Every piece registers through the same plugin and extension mechanisms the diagram shows, which is the property that makes the codebase pleasant to extend: nothing in the core knows the full list of capabilities, it only knows the seams.

From Install to a Working App

The supported path is the download: installers exist for Windows, macOS, and Linux including AppImage, deb, Flathub, and the Microsoft Store, and the app ships with its engines prebuilt. For source builds the README is candid about the weight: Node.js 20 or later, Yarn 4, Make, and Rust for Tauri, then make dev, which installs workspaces, builds core and extensions, downloads engine binaries, and launches the app. GPU variants are a build-time choice through a single environment token, so a CUDA 13 build names itself, and the Makefile even relocates the engine build tree to a short path under the user’s local app data to dodge Windows path-length limits with nvcc, the kind of scar tissue that tells you real people build this on real machines.

Honest limits: this is a heavy desktop codebase, two languages and a process boundary, so a first build is a genuinely long compile and the v2 rebuild is younger than the years of history behind the Jan brand, which means rough edges are still being sanded. Local model quality is bounded by what your hardware can hold, and the agent layer’s tool permissions deserve a careful read before you let it loose on your filesystem. But as a working blueprint for a privacy-first local AI product, with a clean IPC seam, a sandboxed engine worker, an OpenAI-compatible server, MCP, and an extension platform that outsiders can actually build against, the rebuilt Jan is one of the most complete open references in its category, and its source reads like a checklist for anyone assembling their own local assistant.

Watch PyShine on YouTube

Contents