Agent frameworks usually promise simplicity and deliver fragility: the moment a workflow needs retries, approval gates, or state that survives a crash, the happy-path abstraction collapses. LangGraph, MIT-licensed and built by LangChain, starts from the opposite end. The README calls it a low-level orchestration framework for building, managing, and deploying long-running, stateful agents, and it borrows its execution model from serious distributed-systems research: the acknowledgments section credits Google’s Pregel paper and Apache Beam, with the public interface drawing on NetworkX. The result is a graph of nodes that exchange state through typed channels, executed in supersteps with checkpoints after each one, so that durable execution, human-in-the-loop pauses, and resumable failures are properties of the runtime rather than features you bolt on. Companies the README names, including Klarna, Replit, and Elastic, use it for exactly this reason.
The repository is a monorepo of small, focused packages. The core lives in libs/langgraph, where the pregel directory holds the execution engine and graph holds the StateGraph builder, channels define the state model, func implements the newer functional API, and stream handles the many streaming modes. Persistence is its own layer: libs/checkpoint defines the saver and store interfaces with in-memory implementations, libs/checkpoint-postgres and libs/checkpoint-sqlite ship production savers, and libs/checkpoint-conformance tests saver correctness. On top sit libs/prebuilt with the ready-made create_react_agent and ToolNode, libs/sdk-py and libs/sdk-js for remote clients, and libs/cli for the local dev server. A separate JavaScript implementation exists as LangGraph.js, and the README points users who want batteries included to the higher-level Deep Agents package. As always in this series, what follows is an educational tour of published source code.
LangGraph at a glance: StateGraph and the functional API compile into the Pregel loop, whose applier reads and writes channels per superstep, prebuilt agents ride the same engine with the ToolNode, interrupt helpers create pause points, and every step lands in a checkpoint saver beside the long-term store.
Reading the overview from left to right:
- Most users author graphs with the StateGraph builder at libs/langgraph/langgraph/graph/state.py, declaring nodes, edges, and reducers.
- The functional API under libs/langgraph/langgraph/func lets you write plain entrypoint and task functions instead of graphs.
- The ready-made agent is create_react_agent at libs/prebuilt/langgraph/prebuilt/chat_agent_executor.py, a StateGraph assembled for you.
- Everything compiles to the Pregel class at libs/langgraph/langgraph/pregel/main.py, the compiled runnable that streams, invokes, and resumes.
- The superstep machinery lives in the loop driver at libs/langgraph/langgraph/pregel/_loop.py and the applier algorithm at libs/langgraph/langgraph/pregel/_algo.py.
- State itself is a bundle of channels under libs/langgraph/langgraph/channels, each with its own update semantics.
- Tool calls run through the ToolNode at libs/prebuilt/langgraph/prebuilt/tool_node.py, and human approval pauses through the interrupt helper at libs/prebuilt/langgraph/prebuilt/interrupt.py.
- Every superstep ends in a checkpoint via the BaseCheckpointSaver interface at libs/checkpoint/langgraph/checkpoint/base/init.py, with long-term memory separated into the BaseStore at libs/checkpoint/langgraph/store/base/init.py.
Why You Need This
The first reason is the state model, which is the actual product. A LangGraph graph is a set of channels, and the channel classes under libs/langgraph/langgraph/channels define how parallel writes resolve: LastValue at libs/langgraph/langgraph/channels/last_value.py keeps the latest write, Topic at libs/langgraph/langgraph/channels/topic.py accumulates lists, BinaryOperatorAggregate at libs/langgraph/langgraph/channels/binop.py folds writes through your reducer, and EphemeralValue at libs/langgraph/langgraph/channels/ephemeral_value.py exists for one step only. Because the Pregel engine reads all inputs before a superstep and applies all writes after it, you get deterministic, inspectable state transitions instead of whatever a tangle of callbacks happened to mutate. This is what makes the debugging story, from time-travel to the LangSmith traces the README highlights, technically possible rather than aspirational.
The second reason is that durability and human oversight are engineered into the loop, not promised in a blog post. The checkpoint hook at libs/langgraph/langgraph/pregel/_checkpoint.py writes a checkpoint tuple through libs/checkpoint/langgraph/checkpoint/base/init.py after each superstep, serialized by the type-aware JsonPlusSerializer at libs/checkpoint/langgraph/checkpoint/serde/base.py, with in-memory, SQLite, and PostgreSQL implementations ready to swap. The interrupt machinery pairs with it: a node raises an interrupt, the run stops and persists, and a human resumes it later with a Command, using the control types defined at libs/langgraph/langgraph/types.py. Node-level retry policies at libs/langgraph/langgraph/pregel/_retry.py and a task runner at libs/langgraph/langgraph/pregel/_runner.py round out the fault-tolerance picture.
The third reason is the gradation of abstraction, which is rare. Beginners start with create_react_agent at libs/prebuilt/langgraph/prebuilt/chat_agent_executor.py, a complete tool-calling agent wired through the ToolNode at libs/prebuilt/langgraph/prebuilt/tool_node.py in a few lines. Intermediate users build custom StateGraphs, mixing the messages preset at libs/langgraph/langgraph/graph/message.py with conditional edges handled by the branch builder at libs/langgraph/langgraph/graph/_branch.py. Advanced users drop to the functional API, where the modules under libs/langgraph/langgraph/func turn ordinary functions into durable tasks with the same checkpointing underneath. And when the graph runs on a server, the langgraph CLI at libs/cli/langgraph_cli serves it locally while the Python SDK at libs/sdk-py/langgraph_sdk manages runs, threads, and background execution remotely.
The detail view: the authoring layer with StateGraph, node and branch builders, the functional API, control types, and prebuilt agents on the left; the Pregel engine with loop, applier, IO, runner, retries, and checkpoint hooks in the middle; channels and streaming plus managed values and runtime context beside them; the persistence layer with savers, serializer, stores, and cache below; and the CLI and SDK for deployment on the right.
Walking the detail diagram, the engine internals reward a close look. The Pregel class at libs/langgraph/langgraph/pregel/main.py is the single Runnable every compile target becomes: it validates the compiled graph, maps inputs through the IO module at libs/langgraph/langgraph/pregel/_io.py, and then hands control to the loop at libs/langgraph/langgraph/pregel/_loop.py, which plans which tasks can run given the current channel versions, executes them concurrently, and applies their writes atomically through the algorithm at libs/langgraph/langgraph/pregel/_algo.py. Managed values defined at libs/langgraph/langgraph/managed/base.py are injected at runtime rather than flowing through channels, which is how things like step counters stay out of persisted state.
On the persistence side, the separation between checkpoint and store matters architecturally. The saver family, from the InMemorySaver at libs/checkpoint/langgraph/checkpoint/memory/init.py to PostgresSaver at libs/checkpoint-postgres/langgraph/checkpoint/postgres/init.py and SqliteSaver at libs/checkpoint-sqlite/langgraph/checkpoint/sqlite/init.py, versions thread state so a run can be rewound or resumed exactly. The store family, BaseStore at libs/checkpoint/langgraph/store/base/init.py with its own in-memory implementation, holds long-term memories keyed by namespace that outlive any single thread, and the runtime context at libs/langgraph/langgraph/runtime.py hands both to nodes. A cache interface at libs/checkpoint/langgraph/cache/base/init.py completes the picture for node-level output caching.
From Install to a Running Graph
The README’s quickstart is one command, pip install -U langgraph, which pulls the core package plus the checkpoint base with its in-memory implementations, enough to build and run a graph locally with zero infrastructure. Authoring starts with StateGraph at libs/langgraph/langgraph/graph/state.py: you declare a typed state schema, attach nodes and conditional edges, and compile, and the same object then supports invoke, stream with multiple modes, and checkpointed resumption. For production, the checkpoint moves to the PostgreSQL saver, the graph is served through the langgraph CLI or the LangSmith Deployment platform the README describes, and clients talk to it through the SDK at libs/sdk-py/langgraph_sdk, which handles threads, runs, and streamed events. The migration from in-memory to hosted is a constructor argument, which is the whole point of the layered design.
Honest limits: low-level is the honest word in the README tagline, because durable execution semantics, channel reducers, and superstep behavior are real concepts you must learn, and teams that wanted a one-function agent API may find the graph model heavy; the prebuilt agent and Deep Agents cover the simple cases, but the moment you customize, the full model applies; and the ecosystem pulls you toward LangSmith for the best observability experience, though tracing integrations are pluggable. The engine itself, Pregel applied to LLM agents, is a genuinely elegant piece of work, and the code reads like the distributed-systems pedigree it claims. Enjoyed this post? Never miss out on future posts by following us