Codex is OpenAI’s coding agent, and unlike most products from the company it is fully open source under Apache-2.0, with its entire engine readable in the repository. The agent runs locally on your machine: it reads your code, proposes and applies edits, executes commands, and reports back, whether you drive it from a terminal UI, a single non-interactive command, your IDE, the desktop app, or the cloud service. What makes the repository remarkable is its form: the heart of the project is a Rust workspace of 120 crates under codex-rs, decomposed with unusual discipline into engines, protocol types, sandboxes, providers, and UI layers, each an independently buildable unit.

The engineering story is worth studying even if you never ship a coding agent. One agent core serves every surface: the interactive TUI, the scriptable codex exec mode, an app server that bridges IDEs, and cloud task browsing all drive the same thread engine through the same event protocol. Safety is engineered as its own layer, with dedicated crates for Linux, Windows, and macOS sandboxing, an executable policy evaluator, and a network proxy that filters the agent’s traffic. Context management, from AGENTS.md project instructions to automatic compaction of long histories, lives in the core where it can be audited crate by crate.

As with every project in this series, this is an educational tour of published source code, not a guide to delegating dangerous work blindly. Codex executes model-generated commands on your machine, which is exactly why its sandbox layer exists and why its docs ask you to sign in and review what the agent runs. Study the sandboxing design, keep approvals where they matter, and treat the agent the way its own architecture does: as a powerful collaborator that still needs containment.

Codex overview architecture diagram

Codex at a glance: one dispatcher routes to TUI, exec, and app-server surfaces, all feeding one Rust agent core guarded by a sandbox layer and persisted as rollouts.

Reading the overview from left to right:

  • The binary starts at codex-rs/cli/src/main.rs, a clap dispatcher with subcommands from exec and login to mcp, sandbox, and cloud.
  • The interactive surface is the terminal UI under codex-rs/tui, with its own chat widget, keymaps, and onboarding flow.
  • Scripted and CI-friendly runs go through codex-rs/exec, the non-interactive engine driver.
  • IDEs talk to the codex-rs/app-server, which exposes the agent over a JSON-RPC style protocol.
  • The engine itself lives in codex-rs/core, where threads, turns, and tool orchestration are coordinated.
  • Tool implementations, from shell to apply-patch to plan updates, sit in codex-rs/core/src/tools.
  • codex-rs/protocol defines the event and message types every surface shares.
  • The safety layer pairs codex-rs/sandboxing with the platform policy checks of codex-rs/execpolicy.
  • External tool servers connect through the MCP client in codex-rs/rmcp-client.
  • Completed sessions are persisted as rollout files under codex-rs/rollout, which power resume and fork.

Why You Need This

The first reason is that this is OpenAI’s engineering, in the open, at production scale. Most descriptions of how a frontier lab builds an agent are marketing; this one is a Cargo workspace you can compile. Reading codex-core shows how OpenAI structures a turn loop, how tool calls are routed through an orchestrator with parallel execution and approvals, and how model clients are abstracted from providers. Whether you agree with every choice, you are reading decisions that ship to millions of terminals, which is a rare kind of reference material.

The second reason is the sandbox engineering. Agent safety is usually a paragraph in a blog post; here it is a set of crates. The sandboxing crate defines the policy model, Linux uses a dedicated sandbox binary, Windows gets its own sandbox service, macOS uses seatbelt-style profiles, and the execpolicy crate evaluates whether specific commands are allowed, with a network proxy crate standing between the agent and the internet. If you are building anything that lets a model execute code, this is the most complete blueprint you can legally copy.

The third reason is the multi-surface discipline. Codex shows what it takes to keep one agent behavior consistent across a TUI with vim keybindings and inline visualizations, a one-shot CLI for automation, a JSON protocol for IDE integration, and cloud task applications. The answer in the source is a strict event protocol crate plus thread management that every surface consumes. Any team building an agent that must appear in more than one place will recognize the problem and can borrow the solution wholesale.

How It Works

Codex detailed architecture diagram

Inside Codex: the CLI surfaces, the thread engine with its tool orchestrator, the model layer, the platform-specific sandbox stack, and the persistence crates.

Understanding the Architecture

A dispatcher with thirty subcommands. The main.rs of the cli crate maps clap subcommands to subsystems: exec for non-interactive runs, review for code review, login and logout for auth, mcp for external tool servers, plugin for managing plugins, apply for applying a produced diff to your working tree, resume, fork, queue, and archive for session management, cloud for browsing Codex Cloud tasks, plus doctor, sandbox, completion, and update. Surfaces like the TUI and app-server are launched from here but implemented in their own crates, keeping the dispatcher thin and startup fast.

One engine, many threads. The codex-core crate owns the agent loop. A thread manager spawns and coordinates Codex threads, each thread runs turns queued as tasks, and every turn assembles context, streams a model response, and executes tool calls. Session state lives in a dedicated session module, project context arrives through an AGENTS.md loader that reads the repository’s instruction files, and long conversations are summarized by a compaction module with token budgeting so a thread can run far longer than the context window. All of this communicates through the protocol crate’s event types.

Tools as an orchestrated subsystem. The tools module routes every call the model makes. Handlers exist for the shell and unified execution, apply-patch edits, plan updates, image viewing, MCP tool passthrough, tool search, and multi-agent delegation, each with a spec that describes it to the model. Parallel execution lets independent calls run together, an approvals layer decides what needs a human, and lifecycle plus metadata modules track what happened for replay and analytics. The design keeps adding a tool, whether built in or from an MCP server, a matter of registering one more handler.

A model layer behind a seam. The client module in codex-core streams completions, while the model-provider and model-provider-info crates describe providers, their wire formats, and their capabilities. This is how the same engine talks to OpenAI models with ChatGPT-plan auth through the login crate, or to alternative providers declared in config. Compaction fallbacks and error handling live near this seam so that provider hiccups degrade gracefully instead of killing a run.

Sandboxes that differ by platform. The sandboxing crate defines the policy: read-only, workspace-write, or full access, applied to file system and network. Backends are per-platform crates: a Linux sandbox helper, a Windows sandbox service, plus the execpolicy crate that decides command-by-command whether something is permitted and the network-proxy crate that enforces egress rules. The TUI can then offer escalation prompts with a clear model of what is being asked, because the policy and the enforcement are separate crates with explicit interfaces.

Persistence and the session lifecycle. Every interaction is written as a rollout file by the rollout crate, coordinated with the state crate that manages stored sessions. That persistence is what powers resume to continue a conversation, fork to branch one, queue to send a follow-up to a running session, and migrate-rollouts to upgrade old formats. The cloud-tasks crate applies the same session model to tasks started in Codex Cloud, letting you browse and apply their diffs locally through the chatgpt crate’s apply command.

The terminal as a first-class surface. The TUI crate is large and deliberate: a chat widget with streaming diffs, history cells, a pager overlay, onboarding, keymaps, clipboard handling, and IDE context detection. It consumes the same protocol events as the IDE bridge, which is why both surfaces show the same approval requests and the same tool output. End to end, a prompt enters through any surface, becomes a thread in the core, drives model calls and sandboxed tool executions, emits protocol events, and leaves behind a rollout you can resume or fork tomorrow.

Advantages

  • Fully open engine. Apache-2.0 with the entire 120-crate workspace readable, buildable, and modifiable.
  • Sandbox-first execution. Platform-native isolation, executable policy checks, and a network proxy are separate crates, not afterthoughts.
  • One core, every surface. TUI, exec, IDE app-server, and cloud all drive the same thread engine through one protocol.
  • Session superpowers. Resume, fork, queue, archive, and replayable rollouts make long-running work durable.
  • MCP-native extensibility. External tool servers plug in through a standard client crate alongside rich built-in handlers.
  • Rust reliability. Memory-safe systems code with per-crate builds, tests, and fast startup for CLI use.

Benefits

  • Study a production turn loop. Thread management, task queuing, and compaction show how to run agents beyond a single context window.
  • Copy the sandbox blueprint. Policy model, per-platform backends, and command-level policy evaluation are directly reusable in your own tooling.
  • Learn protocol-first design. The protocol crate demonstrates how one event schema keeps four surfaces honest.
  • See provider abstraction done simply. Small provider-info types keep the engine portable across models and endpoints.
  • Understand session persistence. Rollout files and the state crate are a clean pattern for durable agent history.
  • Bridge IDEs and terminals. The app-server shows exactly what a tool must expose to live in both worlds.

Usage

Install Codex CLI with the one-liner from the README on macOS or Linux:

curl -fsSL https://chatgpt.com/codex/install.sh | sh

On Windows:

powershell -ExecutionPolicy ByPass -c "irm https://chatgpt.com/codex/install.ps1 | iex"

Package managers work too:

npm install -g @openai/codex
brew install --cask codex

Start the interactive agent and sign in with your ChatGPT plan, or use an API key:

codex
codex login

Run it non-interactive for scripts and CI, and manage external tool servers:

codex exec "explain what this repo does in three bullets"
codex mcp list
codex resume --last

Run codex --help for the full subcommand reference, including review, apply, fork, sandbox, and cloud.

Conclusion

Codex is the rare artifact that is simultaneously a product, a research platform, and a systems textbook. Its 120 crates draw the boundaries you would want in any agent, model calls behind a provider seam, tools behind an orchestrator, actions behind platform sandboxes, and sessions behind durable storage, and its Rust implementation makes those boundaries enforceable rather than aspirational. Read it to see how OpenAI actually builds an agent, borrow its sandbox and protocol patterns for your own systems, and drive it locally with the same containment philosophy its architecture embodies.

Links:

Watch PyShine on YouTube

Contents