gpt-engineer calls itself “the OG code generation experimentation platform,” and the README is refreshingly honest about what it wants to be: not a polished product but a laboratory where coding-agent builders can specify software in natural language, watch an AI write and execute the code, and then ask for improvements. The project lives at gpt-engineer-org/gpt-engineer, is MIT-licensed, currently ships as version 0.3.1, and is built with Poetry targeting Python 3.10 through 3.12. The README even points you elsewhere on purpose: to the commercial gptengineer.app for a managed service and to aider for a maintained interactive CLI. What remains here is the experimental core, and that core is a beautifully readable answer to one question: what is the smallest honest architecture for turning a prompt file into a running codebase?
The design idea that carries everything is that an agent is a composition of steps, and a step is just a function. The default steps in gpt_engineer/core/default/steps.py are ordinary Python functions with signatures like gen_code, gen_entrypoint, and execute_entrypoint, each taking an AI client, a prompt, a memory, and a preprompt holder and returning a FilesDict. Swap the functions and you have a different agent; wrap them with the variants in gpt_engineer/tools/custom_steps.py and you have lite mode, clarify mode, or self-healing. The preprompts folder is the second half of the idea: the agent’s personality lives in nine plain text files you can override wholesale, which makes behavior a diff instead of a fork.
As always in this series, this is an educational tour of published source code. This particular codebase generates and executes code on your machine, which makes its safety seams worth studying: execute_entrypoint asks you a plain yes/no question before running anything, improve mode refuses to clobber uncommitted work and stages changes through git instead, feedback collection only happens after you have consented in writing to a local consent file, and the repo ships explicit TERMS_OF_USE and DISCLAIMER documents. Read those seams before you copy the patterns; they are what separate an experiment from a liability.
gpt-engineer at a glance: the gpte CLI drives a CliAgent built on the BaseAgent contract, default and custom steps prompt the AI layer and parse its chat into files, DiskMemory logs everything while DiskExecutionEnv runs the generated entrypoint, git helpers protect existing work, and the bench CLI evaluates any BaseAgent against public datasets.
Reading the overview from left to right:
- The front door is the typer application at gpt_engineer/applications/cli/main.py, which the installed
gptecommand points to. - The user-facing agent is CliAgent at gpt_engineer/applications/cli/cli_agent.py, built on the BaseAgent contract in gpt_engineer/core/base_agent.py with SimpleAgent as the default configuration in gpt_engineer/core/default/simple_agent.py.
- Model access is centralized in gpt_engineer/core/ai.py, where the AI class dispatches to LangChain chat clients.
- The generation brain lives in gpt_engineer/core/default/steps.py and its variants in gpt_engineer/tools/custom_steps.py, with the chat parser at gpt_engineer/core/chat_to_files.py turning model output into real files.
- Memory and execution land in gpt_engineer/core/default/disk_memory.py and gpt_engineer/core/default/disk_execution_env.py.
- The safety net for existing projects is gpt_engineer/core/git.py, and the evaluation harness is the bench CLI at gpt_engineer/benchmark/main.py.
Why You Need This
The first reason is architectural humility done right. Many agent frameworks bury their core loop inside classes, callbacks, and framework machinery; gpt-engineer exposes its entire flow as named functions you can read in one sitting. gen_code composes a system prompt from the preprompts, sends the user prompt through the AI layer, and parses the chat into a FilesDict. gen_entrypoint asks the model for a Unix script and writes it to run.sh. execute_entrypoint prints the script, asks whether to execute, and runs it. That is the whole magic trick, spelled out line by line, and once you have seen it you can never again mistake an agent for something more mysterious than a loop over such functions.
The second reason is that the project teaches the two failure modes of code generation and their honest fixes. The first failure is a fresh codebase that never runs, addressed by the entrypoint step plus an execution environment whose run method returns stdout, stderr, and the exit code so the human can see exactly what happened. The second failure is an existing codebase that an improvement request breaks, addressed in improve mode with the machinery of chat_to_files.py: the model returns diffs, parse_diffs applies each hunk with a bounded retry window controlled by the diff timeout, salvage_correct_hunks keeps the hunks that parsed cleanly when others fail, and the git helpers in core/git.py filter and stage only files with uncommitted changes so the human always has a rollback path. Every coding agent since has rediscovered these problems; here they are in their clearest form.
The third reason is benchmark culture. The same install that gives you gpte gives you bench, a harness that loads any agent implementation you point it at and runs it against public datasets including APPS and MBPP, with a template repository for plugging in your own agent. The benchmark types are five small classes, Task and TaskResult among them, and run.py prints results or exports YAML. If you are building a coding agent, this is the discipline that keeps you honest: measure, do not vibe.
The detailed view: the CLI group with its file selector and consent-based learning, the agent core with BaseAgent, SimpleAgent, and the AI layer feeding a TokenUsageLog, data types from Prompt to FilesDict and diffs, the steps and preprompt machinery, disk memory and execution environments, the git and linting safety group, and the benchmark harness.
The detailed diagram rewards a slow pass. In the CLI group, main.py assembles the Prompt, instantiates the AI, memory, and execution environment, and consults FileSelector in gpt_engineer/applications/cli/file_selector.py when improving existing code, while gpt_engineer/applications/cli/learning.py and gpt_engineer/applications/cli/collect.py handle human review and consent-gated feedback. In the agent core, the AI class holds a TokenUsageLog from gpt_engineer/core/token_usage.py so every prompt’s spend is recorded. In the types group, FilesDict from gpt_engineer/core/files_dict.py is a plain dict subclass that can serialize itself back to chat form, and the diff model in gpt_engineer/core/diff.py represents each parsed hunk. The steps group shows PrepromptsHolder at gpt_engineer/core/preprompts_holder.py reading the nine preprompt texts, including generate, improve, clarify, file_format, file_format_diff, file_format_fix, entrypoint, philosophy, and roadmap. The memory group centers on the .gpteng directory defined by the path constants in gpt_engineer/core/default/paths.py, where chats land in all_output.txt and friends, and FileStore in gpt_engineer/core/default/file_store.py adds linting through gpt_engineer/core/linting.py.
From Prompt to Running Code
The README workflow is three moves. Create a folder, write a file named prompt with no extension containing your specification, and run:
gpte projects/my-new-project
Inside the process, load_env_if_needed pulls OPENAI_API_KEY or ANTHROPIC_API_KEY from the environment or a .env file, the prompt is loaded from disk, and CliAgent.init executes the default chain: gen_code asks the model for the full codebase with the system prompt assembled from preprompts, chat_to_files_dict extracts every fenced file into the FilesDict, gen_entrypoint produces run.sh that installs dependencies and runs the code, and execute_entrypoint shows you the script and asks “Do you want to execute this code? (Y/n)” before invoking the execution environment. FilesDict objects carry the content, DiskMemory logs the raw chats into the project’s .gpteng/memory directory, and TokenUsageLog tallies what the session cost.
The interesting override is that every stage is a named step you can replace. The custom variants are wired to CLI flags: lite_gen skips the planning-heavy flow for small tasks behind the lite flag, clarified_gen interrogates you with clarifying questions before generating, and self_heal runs the code, captures the traceback, and asks the model to fix what broke. Pass --use-custom-preprompts and your own preprompts folder replaces the shipped one, which is how the project expects you to encode lessons learned between projects.
Improve Mode and the Diff Dance
Improving existing code is the deeper half of the tool:
gpte projects/my-old-project -i
The flow starts safely: stage_uncommitted_to_git records any dirty state before the agent touches anything, FileSelector lets you choose which files the model may see, and setup_sys_prompt_existing_code primes the assistant with the current codebase instead of a blank slate. The model returns diffs rather than whole files; chat_to_files.py parses them hunk by hunk and apply_diffs patches the FilesDict. When a hunk fails to apply, the retry loop bounded by the diff timeout tries repairs, salvage_correct_hunks commits whichever hunks survived, and handle_improve_mode decides whether to iterate again or hand back. The whole dance ends with the change staged in git, reviewable with ordinary tooling, which is exactly how AI-assisted edits to a real codebase should land.
Models, Clipboard, and Consent
The AI layer is a thin dispatcher over LangChain clients: Azure endpoints get AzureChatOpenAI, Anthropic keys get ChatAnthropic, and everything else routes to ChatOpenAI with a default temperature of 0.1. Local and alternative models are documented paths rather than code forks. Two conveniences stand out for experimentation. ClipboardAI, a subclass of the AI class, replaces API calls with your clipboard, so you can paste model output from any interface and still exercise the full agent machinery. And the caching path wires LangChain’s SQLite cache behind the use-cache flag, so repeated experiments replay instead of re-billing. On the data side, the learning module checks a .gpte_consent file before ever asking for feedback, records your answer, and only then passes reviews to the collector; no consent, no collection.
Try It Yourself
Install stable, export a key, and generate:
python -m pip install gpt-engineer
export OPENAI_API_KEY=[your api key]
gpte projects/my-new-project
Then run the exercise that teaches the architecture: point bench at your own agent through the benchmark template, or write one custom step with the same signature as gen_code and wire it into a run. Because steps are plain functions and memory is just a folder, you can snapshot an entire run, replay it, and diff two agent personalities by editing text files. That is the experimentation platform promise, kept honestly.
gpt-engineer earns its place in this series because it is the hinge between the one-shot code-generation era and the agentic era: small enough to read completely, opinionated about steps and preprompts rather than frameworks, and disciplined about execution consent, git safety, and consent-based telemetry. Before you adopt a heavyweight agent stack, spend an evening inside these files; the vocabulary you learn there transfers everywhere.
Next up in this series: open-interpreter, the project that lets a language model run code on your computer from the terminal. Until then, read the seams, not the slogans. Enjoyed this post? Never miss out on future posts by following us