If you have ever watched a ComfyUI workflow animate across the screen — text encoder on the left, sampler churning in the middle, a latent preview streaming into the browser — you have seen a distributed-looking system that is, in fact, a single Python process orchestrating everything. The repository behind it, Comfy-Org/ComfyUI, is a compact, well-factored Python codebase where a visual graph is validated, topologically sorted, cached, and converted into scheduled GPU work — a good place to study how production inference systems manage memory, queues, and partial recomputation.

Functionally, ComfyUI is three things in one process: a node-graph GUI (with the compiled frontend shipped as a PyPI package), an HTTP and WebSocket API that exposes workflows to other programs, and the backend engine that loads diffusion models and runs them on local hardware. It supports image, video, audio, and 3D generation models natively, on NVIDIA, AMD, Intel, Apple Silicon, and several NPU platforms. The core is Python, versioned at 0.37.0 in pyproject.toml, licensed under GPL-3.0.

The source is worth a tour because almost every interesting problem in local diffusion inference has a concrete answer in the tree. How do you re-execute only the changed part of a graph? Look in comfy_execution/caching.py. How do you run a multi-gigabyte model on a small card? Look in comfy/model_management.py and comfy/model_patcher.py. How do you let a large ecosystem of third-party plugins extend a typed node system without breaking the core? Look in nodes.py and comfy_api/latest. This post walks those paths in order, from the HTTP boundary down to the sampling loop.

Architecture overview of the Comfy-Org/ComfyUI repository

Overview of ComfyUI’s architecture: the bootstrap and server layer on the left, the node registry in the middle, and the execution engine plus model pipeline on the right.

Reading the overview from left to right: main.py bootstraps the process and starts the aiohttp PromptServer from server.py, which accepts prompts and feeds them through a queue to the PromptExecutor in execution.py. On the way in, nodes.py registers the built-in node packs from comfy_extras so that any node class referenced by a submitted graph can be found. The executor walks the graph topologically via comfy_execution/graph.py, consulting comfy_execution/caching.py to skip nodes whose outputs are unchanged. Whenever a node needs a model, the executor resolves file names through folder_paths.py, asks comfy/model_management.py to schedule VRAM, and gets a patched, memory-aware wrapper from comfy/model_patcher.py around whatever comfy/sd.py loaded. Sampling nodes then drive comfy/sample.py, which is where the actual denoising loop lives.

Why You Need This

Most diffusion UIs hand you a fixed pipeline: a prompt box, a couple of sliders, one image out. The moment you want a ControlNet pass with a custom LoRA schedule feeding an upscaler that conditions a second generation — you hit the ceiling. ComfyUI’s answer is to make the pipeline itself the programming model: every stage is a node with typed inputs and outputs, and arbitrary graphs are first-class artifacts saved as JSON. Generated PNGs embed the workflow that produced them, so you can drag an image back onto the canvas and recover the exact graph and seeds that created it.

The second problem is hardware. Diffusion checkpoints are large, consumer GPUs are not, and naive loading will either crash or thrash. ComfyUI treats memory management as a core subsystem: comfy/model_management.py classifies the device into a VRAM state, estimates inference memory, and decides what gets loaded fully, partially, or streamed. The README advertises running large open-source models on roughly 4 GB of VRAM plus 8 GB of RAM via asynchronous weight streaming — and whether or not your workload hits that envelope, the scheduler exists, and it is readable.

The third problem is recomputation. In a chat-style workflow you requeue the same graph dozens of times, changing only the final save node or a seed, and re-running the text encoder and every upstream patch each time would be wasteful. ComfyUI’s executor caches node outputs keyed by an input-signature fingerprint, so submitting the same graph twice executes nothing, and changing the last node re-executes only that node and its downstream dependents. This “partial re-execution” behavior, noted in the README, falls directly out of comfy_execution/caching.py and the topological walker.

Finally, there is the extensibility problem. The ecosystem around ComfyUI is largely built on third-party custom nodes, so the core must load arbitrary community plugins and expose their schemas to the frontend. The plugin protocol in nodes.py, the versioned node API in comfy_api/latest, and the built-in packs in comfy_extras are a mature answer to that — one other Python projects can borrow from directly.

How It Works

Everything starts at main.py, which is worth reading top to bottom before anything else.

Detailed architecture of the Comfy-Org/ComfyUI repository

Detailed architecture of ComfyUI: runtime and API layer, node registry, the execution engine in comfy_execution, and the model/pipeline subsystem under comfy/.

Understanding the Architecture

The bootstrap. main.py enables argument parsing before torch is imported (the code even warns if torch got imported too early), applies model-path overrides via apply_custom_paths(), and runs each custom node’s prestartup_script.py so plugins can prepare the environment before heavy imports. It then constructs the PromptServer from server.py on a fresh asyncio loop, calls nodes.init_extra_nodes() to populate the registry, starts the database-backed AssetManager from app/assets/manager.py, and spawns a daemon prompt_worker thread that feeds queued prompts to one PromptExecutor instance.

The server boundary. server.py is an aiohttp application with a PromptQueue hanging off it. POST /prompt validates and enqueues a workflow; GET /object_info returns the schema of every registered node so the frontend can render them; /ws streams progress, cached-node lists, execution errors, and binary preview frames (event types declared in protocol.py). There are also endpoints for /history, /queue, /view (serving generated media), and internal REST routes under api_server/. The UI itself is not in this repository — app/frontend_management.py serves the compiled frontend, which ships separately as a PyPI dependency.

The node registry. nodes.py maintains NODE_CLASS_MAPPINGS and NODE_DISPLAY_NAME_MAPPINGS. init_builtin_extra_nodes() loads the classic built-ins plus the node packs in comfy_extras/ — custom samplers, ControlNet, upscaling, compositing, and model-specific nodes — while comfy_api_nodes/ holds nodes that call hosted partner models. Third-party packs are imported from custom_nodes/, with web directories registered into EXTENSION_WEB_DIRS so plugins can ship frontend JavaScript alongside their Python. Newer nodes declare themselves through the versioned API in comfy_api/latest, which provides typed io schema objects and per-node input validation.

The execution engine. When PromptExecutor.execute_async() in execution.py receives a prompt, it wraps the run in torch.inference_mode(), builds a DynamicPrompt (which allows subgraph expansion through comfy_execution/graph_utils.py), and attaches per-prompt caches from comfy_execution/caching.py. Depending on launch flags the cache is NullCache, LRUCache, HierarchicalCache, or the default RAMPressureCache, which evicts intermediate results when system RAM runs short. Cache keys come from CacheKeySetInputSignature, a fingerprint of each node’s inputs; IsChangedCache consults a node’s IS_CHANGED/fingerprint_inputs hook when inputs cannot be compared structurally.

The topological walk. comfy_execution/graph.py implements TopologicalSort and ExecutionList. The executor calls stage_node_execution() in a loop, which picks the next node whose dependencies are ready, then module-level execute() runs that node’s FUNCTION, handling lazy evaluation via check_lazy_status, asynchronous nodes that resolve later, and cycle detection via comfy_execution/validation.py. Progress flows out through comfy_execution/progress.py and latent_preview.py, which decodes intermediate latents into small preview frames for the browser.

Model and pipeline management. When a loader node runs, folder_paths.py resolves names like checkpoints or loras against a registry of folder paths (folder_names_and_paths), which utils/extra_config.py can extend from extra_model_paths.yaml — that is how you share one model directory across multiple UIs. comfy/sd.py is the loading hub: load_checkpoint_guess_config() inspects state-dict structure via comfy/model_detection.py against the registry in comfy/supported_models.py, then returns the diffusion model as a ModelBase subclass plus CLIP and VAE instances. comfy/model_patcher.py’s ModelPatcher attaches LoRA and other patches (add_patches, calculate_weight) and moves weights in and out of VRAM during patch_model(). comfy/model_management.py owns the device: load_models_gpu() decides what fits, free_memory() and unload_all_models() reclaim it, and quantized weight paths in comfy/ops.py trade precision for footprint. Sampling nodes call comfy/sample.py, which prepares noise, selects a sampler from comfy/samplers.py or comfy/k_diffusion/sampling.py, and steps the latent through the network.

End to end. A queued request travels: browser workflow → POST /prompt on server.py → PromptQueue → prompt_worker thread → PromptExecutor.execute_async() → validation and cache lookup → topological execution of each node class found in NODE_CLASS_MAPPINGS → model loads routed through folder_paths.py and comfy/model_management.py → patched weights applied by ModelPatcher → denoising in comfy/sample.py → outputs and previews streamed back over the WebSocket — with every executed node’s output cached so the next pass can skip straight to whatever changed.

Advantages

  • Partial re-execution is architectural, not bolted on. The cache classes in comfy_execution/caching.py plus the topological walker in comfy_execution/graph.py mean unchanged subgraphs are skipped by construction.
  • Memory management is a first-class subsystem. VRAM state classification, memory estimation, offloading, and pinned-memory release in comfy/model_management.py and comfy/memory_management.py target real consumer-hardware constraints.
  • A genuinely open plugin surface. custom_nodes/ loading with prestartup scripts, per-pack web directories, and the typed, versioned node API in comfy_api/latest let third parties extend backend and frontend cleanly.
  • Everything is inspectable and scriptable. The graph the UI submits is plain JSON over POST /prompt, and script_examples/ shows API clients, so workflows can be driven from any language.
  • Broad model coverage in one core. load_checkpoint_guess_config() and the registry in comfy/supported_models.py let image, video, audio, and 3D models share one execution path.
  • Built-in node library with depth. The more than 140 modules in comfy_extras/ ship reference implementations for samplers, conditioning, compositing, and more, doubling as documentation for writing your own nodes.

Benefits

  • Reproducibility by default. Workflows embed into generated media, so any output image can be dragged back onto the canvas to recover the full graph and seeds that produced it.
  • Faster iteration. Caches key on input signatures, so changing one node re-runs only that node’s downstream slice rather than the whole graph.
  • Runs where the hardware is. Windows, Linux, and macOS support across NVIDIA, AMD, Intel, and Apple Silicon, with optional fully-offline operation via --disable-api-nodes.
  • A learning resource in itself. The engine is a compact, working example of topological scheduling, fingerprint caching, and memory-aware loading — patterns that transfer to any batch-compute system.
  • Production integration path. HTTP and WebSocket endpoints, job history and cancel routes under /api/jobs, and App Mode for exposing complex graphs as simple UIs make the jump from desktop experiment to service straightforward.
  • Extensible toward hosted models. The comfy_api_nodes/ packs show how local workflows can mix in remote partner models behind the same node interface.

Usage

Manual install (after installing PyTorch for your GPU per the README):

git clone https://github.com/comfyanonymous/ComfyUI
cd ComfyUI
pip install -r requirements.txt

Run the server:

python main.py

Or install and manage with the official CLI tool:

pip install comfy-cli
comfy install

Place checkpoint files in models/checkpoints (and VAEs in models/vae, LoRAs in models/loras), optionally point extra_model_paths.yaml at an existing model directory, then open http://127.0.0.1:8188 and start wiring nodes.

Conclusion

ComfyUI’s source tree is a rare thing: a widely used creative tool whose internals actually explain its behavior. The split between server.py (transport), nodes.py (registry), execution.py with comfy_execution/ (scheduling, caching, validation), and comfy/ (models, memory, sampling) is clean enough to read layer by layer, and each layer answers a real systems question with concrete code. If you build anything that queues, caches, or schedules GPU work, an afternoon in this repository will pay for itself.

Links:

Watch PyShine on YouTube

Contents