Most agent frameworks give you a loop and a toolbox. MetaGPT gives you an org chart. The repository at geekan/MetaGPT, MIT-licensed and currently at version 1.0.0, sells itself in one sentence: the multi-agent framework, and internally it includes product managers, architects, project managers, and engineers, providing the entire process of a software company along with carefully orchestrated SOPs. The core philosophy is compressed into a formula the README displays proudly, Code = SOP(Team): take the standard operating procedures that make human software teams work, materialize them as message-passing code, and staff them with LLM-driven roles. Type one line, metagpt “Create a 2048 game”, and a relay of agents produces requirement documents, designs, task breakdowns, code, and tests in a workspace folder. The project has also grown beyond the demo: the README points to MGX, a hosted data-ops product built on the same engine, and to AFlow, the agentic-workflow search research it says received an ICLR 2025 oral. Python 3.9 through 3.12 is supported, and installation is a single pip install.

What makes the source worth a slow read is that the metaphor is not marketing paint. The package is literally organized as a company: a Team object that hires Roles, a shared Environment that routes Messages between them, an Action library per job description, a Memory module per employee, and a provider layer that is the office’s electricity supply. There is also a newer RoleZero track that replaces fixed relays with planner-driven autonomy, so two generations of agent architecture coexist in one tree. As always in this series, what follows is an educational tour of published source code.

MetaGPT overview architecture diagram

MetaGPT at a glance: startup.py assembles a Team, the Team picks a SoftwareEnv or the default MGXEnv, the shared Environment broadcasts Messages to the five company roles, each role executes actions that publish artifacts back into the bus, and every action calls the LLM provider layer underneath.

Reading the overview from left to right:

Why You Need This

The first reason is that MetaGPT makes SOPs executable, and that turns out to be the cheapest known cure for multi-agent chaos. A Role in metagpt/roles/role.py does not poll its neighbors or wait for a central scheduler to prompt it; it declares which message types it watches, and when a matching Message lands in the shared Environment, the think-act-observe loop wakes it up, picks the next Action from its repertoire, publishes the artifact, and goes back to sleep. Because the coordination lives in subscriptions rather than in free-text chatter, the ProductManager’s PRD cannot be argued out of existence by an enthusiastic engineer two turns early. Artifacts also land on disk when the Team archives the run, so a session leaves a browsable workspace instead of a chat log.

The second reason is the structured-output machinery, which is the quiet engineering core of the whole framework. LLMs improvise; schemas do not. The ActionNode class in metagpt/actions/action_node.py turns a nested Pydantic-style definition into a typed node tree with keys, instructions, and examples, then runs an ask-and-repair loop that fills the schema and fixes malformed output. ActionGraph in metagpt/actions/action_graph.py composes those nodes into a graph so one big artifact, like a full PRD, becomes many small, checkable fills instead of one giant prompt. Every serious action in the company is built on this substrate, from WritePRD and DesignAPI to the code review action in metagpt/actions/write_code_review.py, and the effect is that outputs are parseable files, not vibes.

The third reason is that the codebase shows the frontier of agent design without hiding the seams. The same tree contains the classic SOP relay and the newer RoleZero layer, represented in the diagram by the TeamLeader in metagpt/roles/di/team_leader.py, where an agent plans its own next action through the Planner in metagpt/strategy/planner.py, searches decision trees with the Tree-of-Thought strategy in metagpt/strategy/tot.py, and consults past episodes through the ExperienceRetriever in metagpt/strategy/experience_retriever.py backed by the experience pool in metagpt/exp_pool/manager.py. Memory gets equal treatment, with a vector-indexed LongTermMemory beside the plain message store. Reading both generations side by side teaches you more about agent architecture than most tutorials, because you see exactly what fixed procedures give up to gain autonomy.

MetaGPT detail architecture diagram

The detail view: orchestration core on top, the pub/sub Environment with its two concrete rooms, the Role layer with its think-act-observe loop and RoleZero offshoot, the Action layer with ActionNode as the structured-output engine, memory stores below, and the LLM-plus-strategy intelligence layer feeding everything.

Walk through the detail diagram and the whole machine snaps into focus. Start at the orchestration core: startup.py parses your idea and constructs a Team in metagpt/team.py, which keeps a shared Context in metagpt/context.py holding configuration, repository, and cost state, and a Message and Artifact vocabulary from metagpt/schema.py. The Team then spawns an environment: the BaseEnvironment in metagpt/base/base_env.py is the generic pub/sub room, SoftwareEnv is the software-company variant, and MGXEnv is the default when use_mgx is true, extending the same base with a data-workspace flavor. Message routing is dashed on purpose: nothing in the room executes logic, it only ferries typed payloads.

The role layer is where Pydantic modeling meets job descriptions. Every role extends BaseRole in metagpt/base/base_role.py, and the Role base class wires the loop: observe the environment, think about which action to take next, act, and remember. The five company roles then bind concrete actions: the ProductManager produces the PRD through WritePRD, the Architect turns requirements into an API design through DesignAPI in metagpt/actions/design_api.py, the ProjectManager splits work through the task actions in metagpt/actions/project_management.py, the Engineer implements through WriteCode in metagpt/actions/write_code.py, reviews its own output, and debugs failures through DebugError in metagpt/actions/debug_error.py, while the QaEngineer exercises the code through RunCode in metagpt/actions/run_code.py. The review action feeds fixes back into WriteCode, the run action feeds errors into debug, and the loop closes exactly the way a human team’s iteration does.

Below the actions sit two support layers. Memory keeps each role’s history in the Memory class of metagpt/memory/memory.py, with LongTermMemory in metagpt/memory/longterm_memory.py persisting and searching through the vector-backed MemoryStorage in metagpt/memory/memory_storage.py, and the experience pool manager harvesting traces for later retrieval. Intelligence is the provider stack: BaseLLM defines the interface, the OpenAI-compatible implementation in metagpt/provider/openai_api.py is the flagship, and sibling adapters cover Anthropic, Gemini, Azure, Ollama, Bedrock, and several Chinese providers. The strategy classes above it let a RoleZero agent decide its own next move instead of following a fixed relay.

From Install to a Working App

The path from zero to a multi-agent deliverable is short. Ensure Python 3.9 or later but below 3.12, install the package with pip install metagpt, configure a model key, and run the one-liner from the README: metagpt “Create a 2048 game”. The process prints the company at work, each role publishing its artifact in turn, and when the run finishes, a workspace folder holds the PRD, the system design, the task list, the generated repository, and the tests, ready to browse as if a small remote team had just delivered a sprint. The Team archive step is what finalizes that workspace, so a session is inspectable after the fact rather than evaporating with the terminal.

For programmatic use, the diagram’s surfaces are public API: construct a Team, hire the roles you want, invest to set the compute budget, then await run_project with your idea and run with a round count; all four methods live in metagpt/team.py. For the modern autonomous style, the README’s own example imports the DataInterpreter from metagpt/roles/di/data_interpreter.py, a RoleZero-style analyst that writes and runs its own Python against your dataset. Between those two styles you can staff anything from a full software relay to a single data analyst, paying only for the rounds you invest.

Honest limits: runs cost real tokens across many rounds, and output quality tracks the model you point it at, so weaker models produce weaker PRDs and buggier code with the same ceremony. Fixed SOPs trade flexibility for legibility, a feature when you want auditable artifacts and a cost when your idea does not fit the relay. But as a blueprint for industrializing LLM collaboration, with schema-driven actions, subscription-based coordination, memory, planning, and provider abstraction in one readable tree, MetaGPT remains one of the most instructive codebases in the multi-agent world, and exactly the kind of source to read before designing your own team.

Watch PyShine on YouTube

Contents