Dify describes itself as an open-source LLM app development platform, and unlike many projects that use that phrase loosely, it really is a platform: a Next.js console for building, a Python backend that serves chat, completion, agent, and workflow applications, a RAG engine, model management through a plugin system, observability with first-party tracing integrations, and a story that now extends into a dedicated agent runtime written in Go. The monorepo at langgenius/dify ships the API as dify-api version 1.17.1, carries a Linux Foundation project badge in its README, and licenses everything under a modified Apache 2.0 with two famous additional conditions: you may not run it as a multi-tenant service without written authorization, and you may not strip the logo from the frontend. Reading the source makes both conditions concrete, because multi-tenancy is not a bolt-on here, it is the axis around which the data model is organized.

The design idea that carries the architecture is the separation between what a user builds and what executes it. Users compose applications in the console, whether that is a chat assistant, a text generator, or a workflow graph with dozens of nodes. The backend stores those definitions as data, and when a request arrives, the app core materializes the right runner for the application type, drives it through the workflow engine or the agent runtime, resolves models and tools through the plugin layer, and streams results back through a callback system that every observer can listen to. Concerns stay in their lanes: prompt assembly, memory, retrieval, moderation, and tracing are each their own package under api/core.

As always in this series, this is an educational tour of published source code. Two things deserve special attention from a safety standpoint. First, the platform is explicitly multi-tenant, so rbac, workspace scoping, and the trigger system in the API surface are worth studying rather than skipping; this is a codebase where authorization mistakes would be real vulnerabilities. Second, the licensing conditions above are enforceable boundaries, not suggestions, and anyone self-hosting should read the LICENSE file before scaling up. The quickstart itself is honest about requirements: at least two CPU cores, four GiB of RAM, and Docker Compose v2.24.0 or later.

Dify overview architecture diagram

Dify at a glance: the Next.js console calls the Python API's controller layer, services drive the app core, which executes workflow graphs or agent loops, the RAG engine indexes and retrieves knowledge, models and tools arrive through the plugin layer with MCP bridging, everything persists through the data models, and the ops package traces runs to external observability platforms.

Reading the overview from left to right:

  • The visual builder and dashboard is the Next.js console in web, one of the most-complete open-source frontends in this space.
  • Requests enter through the controller packages in api/controllers, split by audience: console, the end-user web API, the service API, files, MCP, and triggers.
  • The service layer in api/services is the backbone, with 158 modules that carry almost all business logic.
  • Application execution lives in api/core/app and the workflow engine in api/core/workflow, whose nodes package holds every runnable node type.
  • Intelligence is served by the agent runtime in api/core/agent, the RAG engine in api/core/rag, and the tools framework in api/core/tools.
  • Model providers arrive through the plugin layer in api/core/plugin, with MCP bridging in api/core/mcp.
  • Persistence centers on api/models and observability on api/core/ops.

Why You Need This

The first reason is that Dify is the reference implementation of “AI application as a product.” A chat assistant here is not a script; it is a configured object with a model setting, an orchestrated prompt, context variables, moderation rules, speech-to-text options, and a published API. The app core packages those concerns explicitly: app configuration, feature orchestration, and the runner machinery each have their own module group, so you can see exactly how a visual configuration becomes an execution plan. If you have ever wondered how tools like this turn a form full of settings into a running, streaming, multi-turn conversation, this codebase answers it in the open.

The second reason is the workflow engine and its node zoo. The nodes package under core/workflow implements every node type the canvas offers, from simple LLM calls and template transforms to knowledge retrieval, HTTP requests, code execution, iterations, and question classifiers, all wired through a graph runner with variable propagation. The generator side of the package turns a graph definition into an executable stream. Combined with the trigger controllers, which let external events start a run, this is effectively an automation engine purpose-built for LLM work, and reading its node implementations is the fastest way to understand the real semantics of visual workflow products.

The third reason is the plugin architecture for models and tools. Rather than shipping every provider in-process, Dify resolves models, embeddings, reranking, and tool providers through the plugin layer, which keeps the core small and lets the ecosystem grow independently, with MCP as an additional bridge for external tool servers. The RAG engine around it is a complete assembly line in the conceptual sense, document extraction, cleaning, splitting, embedding, indexing, post-processing, reranking, and retrieval each live in their own package under core/rag, which makes it one of the best-organized retrieval codebases to read. Whether you build with Dify or not, this separation of model access from application logic is the pattern modern LLM platforms converge on.

Dify detailed architecture diagram

The detailed view: console and client SDKs above the API layer with controllers, services, async tasks, and a scheduler, the app core with its config and feature machinery, the agent group with memory, prompts, and an LLM generator, the RAG engine decomposed into extraction, retrieval, and docstore, tools bridging MCP and plugins, the platform group with data models, migrations, repositories, configs, and rbac, and the observability group with tracing, telemetry, moderation, and callback handlers.

The detailed diagram rewards a slow pass. In the API layer, long-running work leaves the request cycle through api/tasks and recurring work through api/schedule. In the agent group, conversation memory lives in api/core/memory, prompt assembly in api/core/prompt, and the LLM generator that turns instructions into prompts in api/core/llm_generator. In the knowledge group, api/core/rag/extractor parses documents, api/core/rag/retrieval executes searches, and api/core/rag/docstore holds the chunks. In the platform group, api/migrations evolves the schema while api/repositories structures data access and api/core/rbac guards authorization. The observability group pairs tracing in api/core/ops, which ships integrations for platforms like Langfuse, Opik, and Arize Phoenix as the README documents, with telemetry in api/core/telemetry and content screening in api/core/moderation.

The New Agent Frontier

The most interesting recent addition is visible at the repository root: two companion projects alongside the classic api and web. dify-agent is a Python package with its own docs, examples, and tests, introducing what the project calls Agenton alongside a new agent runtime. dify-agent-runtime is a Go implementation of a shell server and runtime utilities: a main server binary served with shellctl serve, a PTY sanitizer that filters tmux pipe-pane output, a SQLite exit recorder for post-drain bookkeeping, and a CLI that talks to the agent backend. That is a terminal-grade execution substrate written in a systems language, and it signals where the project is heading: agents that do real work on real machines need runtimes designed for that, not just another HTTP wrapper. The monorepo layout also carries sdks for client languages, e2e tests, and a skills directory, showing a team that treats the whole developer surface as product.

Try It Yourself

The README quickstart gets a full stack running in minutes:

cd dify/docker
cp .env.example .env
docker compose up -d

Then open http://localhost/install and initialize. After that, do the reading exercise that teaches the most: build one chat application and one workflow application in the console, publish both, and then find in the source where each configuration lands. Follow a chat request from the console API through the service layer into the app core, and follow a workflow run from the graph definition into a specific node implementation. Seeing your own clicks become code paths is the shortest route to understanding platform architecture.

Dify earns its place in this series because it is the most complete open embodiment of the LLM platform idea: visual building without a conceptual ceiling, production features like moderation, rbac, triggers, and observability built into the core, a plugin economy for models and tools, and now a Go-based runtime for serious agent work. Its license asks commercial multi-tenant users to pay, which is a fair bargain for this much open engineering.

Next up in this series: CrewAI, the framework that treats multi-agent orchestration as role-based teamwork. Until then, read the seams, not the slogans.

Watch PyShine on YouTube

Contents