Before the current wave of coding agents, there was GPT Pilot, the project at Pythagora-io/gpt-pilot whose thesis was refreshingly specific: AI can write most of the code for an app, maybe ninety-five percent, but a developer is needed for the rest, so the tool should code step by step with a human in the loop rather than dump a codebase at once. That thesis produced one of the most instructive agent architectures ever published: an orchestrator that walks a project from specification through architecture, development tasks, code edits, review, and debugging, with eleven role-playing agents, every conversation persisted to a database, and every project state snapshotted so a run can resume after a crash. The core package, pythagora-core at version 2.0.10, is Python 3.9 or newer under the Functional Source License 1.1 with an MIT future grant, and it powered the Pythagora VS Code extension as well as the CLI.

One thing must be said plainly before the tour. The repository is no longer maintained, and its README now opens with a security notice: in August 2025 a malicious commit disguised as a routine revert hid a credential-stealing supply-chain worm inside the telemetry package, it was publicly reported in June 2026 and removed three days later, and anyone who ran the tool from source in that window was advised to rotate credentials. The malicious files are gone from the current tree, and this tour covers the published source as it stands today, which makes the project a double lesson: a landmark in agent design, and a case study in why even famous repositories need dependency vigilance. As always in this series, what follows is an educational tour of published source code.

GPT Pilot overview architecture diagram

GPT Pilot at a glance: the CLI drives the orchestrator, which spawns role-playing agents whose conversations flow through the LLM layer and its prompt templates, while project state, specifications, and chat history persist to SQLAlchemy models and the orchestrator edits files and runs commands through the virtual file system and process manager.

Reading the overview from left to right:

Why You Need This

The first reason is the agent decomposition, which remains a masterclass in division of labor. The orchestrator at core/agents/orchestrator.py walks the lifecycle: the specification writer in core/agents/spec_writer.py interrogates a one-line app description into a full specification, the architect in core/agents/architect.py selects technologies and verifies they are installed, the tech lead in core/agents/tech_lead.py breaks the spec into development tasks, and only then does the developer in core/agents/developer.py break each task into step-by-step implementation descriptions for the code monkey in core/agents/code_monkey.py, the agent that actually edits files. Review, troubleshooting, bug hunting in core/agents/bug_hunter.py, and documentation via the tech writer round out the roster, each defined as a small subclass of the BaseAgent in core/agents/base.py. Read these classes and you can see why coarse one-shot generators produced bug farms while this step-wise design debugged as it went.

The second reason is the context engineering that made it scale. The developer class extends the breakdown and relevant-files mixins in core/agents/mixins.py, and before every conversation the system filters the project down to the files relevant to the current task, so the model never carries the whole codebase in context. File contents and descriptions live in their own database models, the virtual file system in core/disk/vfs.py tracks changes with an ignore-path filter from core/disk/ignore.py, and commands run under the process manager in core/proc/process_manager.py with execution logs recorded for the bug hunter to analyze when something fails. This filtered-context pattern predates and parallels what modern coding agents do, and the source shows every trick in plain Python.

The third reason is durable state, the least glamorous and most valuable part. Every model conversation is a row in chat_convo and chat_message tables, every step advances a project_state snapshot in core/db/models/project_state.py, the specification itself is a database row in core/db/models/specification.py, and even raw LLM requests are logged for cost and debugging. Storage defaults to SQLite through async SQLAlchemy with optional PostgreSQL, which means a crashed or interrupted build resumes with python main.py and a project id, and stepping backward to a chosen step is an explicit CLI operation. If you have ever lost an hour-long agent run to a network hiccup, this design needs no selling.

GPT Pilot detail architecture diagram

The detail view: entry and CLI on top, the orchestrator and its agent subclasses in the middle, the LLM layer with provider clients and the response parser beside them, database models underneath, and the file system, process, and UI runtime on the right with the telemetry module isolated.

Walking the detail diagram, the conversation machinery is the heart of the system. Each agent holds a Convo from core/agents/convo.py, which tracks message history through the LLM-side structures in core/llm/convo.py, sends prompts through the provider-agnostic base in core/llm/base.py with concrete clients for OpenAI, Anthropic, and Groq, and parses replies with the structured parser in core/llm/parser.py so agents can return actions, file paths, and code blocks rather than prose. Prompts are not buried in Python strings: the templates directory holds a dot-prompt file for every phase of every agent, with shared partials for things like file listings and coding rules, making the entire behavior of the system auditable text. Human intervention is a first-class concept in core/agents/human_input.py, and review requests pause the orchestrator until the developer answers, which is precisely the five-percent thesis made mechanical.

The runtime layer shows how the tool talked to the world. The console UI in core/ui/console.py renders the interactive CLI, while the IPC client in core/ui/ipc_client.py streams the same events to the VS Code extension, so the extension was a thin shell over the same core. The executor agent and process manager run the app’s dev server, capture its output, and hand logs to the bug hunter, which then decides whether it has found the bug or needs to add more logging and iterate, a loop that reads like an engineer’s actual workflow.

And then there is the cautionary corner: the telemetry package at core/telemetry/init.py is now a stub, but for ten months it contained a hidden loader that downloaded and executed an obfuscated payload harvesting cloud, GitHub, and SSH credentials. The commit hid inside a routine-looking revert, survived because the project had gone quiet, and was caught by an external researcher. The architectural lesson stands, and so does the operational one: pin what you run, watch what telemetry imports, and treat unmaintained repositories with suspicion.

From Install to a Built App

The CLI path is six commands: clone, virtualenv, install requirements, copy the example config, set an OpenAI, Anthropic, or Groq key plus an optional database URL, and run python main.py with your app description. From there the orchestrator asks its clarifying questions, writes the spec for your approval, plans the architecture, generates development tasks one at a time, and edits your workspace folder through reviewed steps, with commands like list, resume, and rollback managing the lifecycle. The same core drove the Pythagora VS Code extension, where all of this happened in panels instead of a terminal.

Honest limits: the project is explicitly unmaintained, so treat it as a study artifact, run it only in a disposable environment, and never on a machine with credentials you care about; the current tree is clean, but the trust is gone. The architecture assumes human availability at review gates, the agent roster carries real token costs over long builds, and the frontend agent was always the roughest edge of the system. But as the clearest published blueprint of a step-wise, stateful, human-in-the-loop code generation system, and as a security cautionary tale, this repository earns its place in the canon.

Watch PyShine on YouTube

Contents