Most coding agents wake up with amnesia every session: context is assembled, work is done, and everything learned evaporates when the process exits. Letta Code, from the team behind the MemGPT research paper, attacks that premise directly. Its README describes it as a stateful agent harness for creating agents that are more like people than tools β€” agents with memory, identity, and a sense of experience over time, which learn and evolve across long horizons by rewriting their own memory, skills, prompts, and even the harness itself through mods. The project is Apache-2.0 licensed, ships as the npm package @letta-ai/letta-code at version 0.34.10, and is written in TypeScript running on Bun. You can talk to the same agent through a local terminal UI, a desktop app for macOS, Windows, and Linux, the hosted chat.letta.com including mobile, or messaging integrations for Telegram, Slack, Discord, and custom channels.

The repository layout makes the ambition legible. Under src you find cli, an Ink-and-React terminal application; headless and gateway-core for non-interactive and server operation; a backend directory that literally forks into local and api implementations; agent, the runtime that owns lifecycle and memory; tools, providers, and skills; and an infra band of channels, websocket, web, queue, cron, and sandbox directories that turn a laptop process into an always-on service. The python directory is a packaging sidecar rather than the product, which tells you this is one of the few agents in this series whose core is not Python. As always in this series, what follows is an educational tour of published source code.

Letta Code overview architecture diagram

Letta Code at a glance: CLI, headless, and gateway entry points pick a backend, which drives the agent runtime; the runtime reads and writes git-tracked MemFS, loads skills, dispatches tools, and calls providers, while channels, websockets, and queues feed work in from the outside world.

Reading the overview from left to right:

Why You Need This

The first reason is memory that actually persists and is actually editable β€” by the agent. MemFS treats all agent context, including memory blocks, as files tracked with git, implemented in src/agent/memory-filesystem.ts; pointing it at your own repository with a memory-repository setting syncs an agent’s evolving mind to GitHub like any other codebase. Because memory is data the model can read and write, improvement becomes a first-class action rather than a research project: the README’s sleeptime feature configures periodic dreaming, the palace command renders the current memory structure, and doctor investigates why an agent behaved a certain way. Guards around what memory may contain are enforced separately in src/memory-confinement.ts. This is the MemGPT lineage showing up as product: instead of a context window you stuff, you get a workspace the agent curates.

The second reason is one agent on every surface. The same stateful identity answers in your terminal, in the desktop app, in a browser at chat.letta.com, and in group chats, because channel adapters under src/channels β€” with src/channels/telegram as one example β€” feed messages into the same headless machinery. Agents are not pinned to your laptop either: running the server command with a computer name registers the machine so cloud-stored agents can be routed onto it, and headless invocations accept a computer flag to execute on a specific machine. The backend switch in src/backend/backend.ts is what makes this mundane: local mode keeps state on device, cloud mode keeps agents, memory, and conversations in Letta Cloud while compute stays wherever you pointed it.

The third reason is governance and extensibility as installed defaults rather than DIY. Tool calls flow through a permissions system in src/permissions with modes and allow or deny rules, and interactive confirmation is handled by src/agent/approval-execution.ts. Lifecycle scripts run from src/hooks, commands execute inside the isolation of src/sandbox, and external capability arrives through the MCP client in src/mcp-client.ts. Model access is bring-your-own-key through the connections managed in src/providers, and skills install from GitHub, ClawHub, or Hermes into global, project, or agent scopes. Cron schedules and queues round it out so an agent can wake up and work without a human typing first.

Letta Code detail architecture diagram

The detail view: entry surfaces and their stream and settings plumbing on the left, the backend fork and agent runtime with subagents and prompts in the middle, tools with permissions, sandbox, MCP, and hooks below, the MemFS and skills memory column, and channels, websockets, and queues on the right.

Walking the detail diagram from the entry side, the interactive app in src/cli/App.tsx and the one-shot path in src/headless.ts share the same settings and state plumbing: persisted preferences flow through src/settings-manager.ts, agent state queries go through src/app-server-client.ts, and machine-readable output is emitted by src/stream-json-writer.ts. That last file is what makes letta usable as a component in other programs rather than only as a chat experience.

The runtime core is where the stateful design becomes concrete. src/backend/backend.ts picks between the local implementation under src/backend/local and the cloud client under src/backend/api, and the agent runtime in src/agent works against that seam. Creation starts at src/agent/create.ts; system prompt composition draws on the library in src/agent/prompts, with bundled defaults becoming cloud-managed through src/agent/cloud-managed-system-prompt.ts. Subagents β€” the general-purpose, forked, and recall helpers described in the README β€” live in src/agent/subagents, and a shared execution context is carried in src/runtime-context.ts.

The tool and memory columns close the loop. Tool schemas are declared in src/tools/define-tool.ts with handlers under src/tools/impl; before anything dangerous runs, the approval module in src/agent/approval-execution.ts consults the rules in src/permissions, and execution lands in the sandbox at src/sandbox. On the memory side, everything the agent reads or writes about itself passes through src/agent/memory-filesystem.ts and the confinement guards in src/memory-confinement.ts, while skill content is resolved in src/skills from global, project, and agent-scoped sources.

From Install to a Self-Editing Agent

The on-ramp is genuinely short: install globally with npm, run letta inside your project directory, and choose to sign in with Letta or proceed locally β€” the choice is saved, and can be flipped later with the backend subcommand or overridden per run with a flag. Connect your own API keys with the connect command, swap models with the model command, and try the tutorial agent with a personality flag to see memory in action. From there the interesting commands are all about the long horizon: schedule dreaming with sleeptime, inspect what the agent knows with the palace command, sync its memory to your own repository with a memory-repository setting, install a skill straight from a GitHub URL, and register the machine as a named computer so the same agent can follow you to another machine. An always-on variant is the same binary under the server command, where cron schedules and queues keep it working between conversations.

Honest limits: the stateful model asks you to trust a new abstraction β€” memory that the agent edits β€” and the boundary enforcement in src/memory-confinement.ts is only as good as your configuration of it. Cloud is the default backend, so privacy-conscious users must consciously pick local mode. Automatic dreaming is disabled by default on native Windows, per the README, with manual commands still available. And because the harness is TypeScript on Bun with many moving parts, from channels to crons to mods, there is more surface area here than a minimal coding agent β€” which is precisely the point, but it is a real cost if all you wanted was a stateless autocomplete in your terminal.

Watch PyShine on YouTube

Contents