SWE-agent is the open source project that turned a research question into a benchmark-beating system: what interface should you give a language model so it can actually fix bugs in real repositories? Built by researchers from Princeton and Stanford and described in a NeurIPS 2024 paper on Agent-Computer Interfaces, the tool lets models such as Claude Sonnet or GPT-4o hunt through a codebase, edit files, and run tests inside a sandboxed environment until they can submit a patch for a real GitHub issue. The repository reports state-of-the-art results on SWE-bench among open source projects, and the whole system is governed by one readable YAML file.

The codebase is a compact Python package under the MIT license that separates cleanly into four concerns: run surfaces that launch the agent, an agent core that talks to models through LiteLLM, a tool layer that defines and parses actions, and an environment that owns the sandbox. One honest note up front, straight from the README: most of the team’s current effort has moved to a companion project called mini-SWE-agent, which matches SWE-agent’s performance with far less code. The flagship repository remains the richer system to study, precisely because it shows every layer a production research agent needs.

As with every project in this series, this is an educational tour of published source code. SWE-agent executes model-generated commands inside containers, and its own docs steer you toward Docker sandboxes and cost limits for good reason. If you experiment with it, do so on repositories you are authorized to modify, keep the per-instance cost limits set, and treat every submitted patch the way the built-in review flow does: as a draft that still needs human eyes.

SWE-agent overview architecture diagram

SWE-agent at a glance: CLI entrypoints feed one agent core, whose actions flow through a tool handler into a sandboxed environment and out to trajectory files.

Reading the overview from left to right:

Why You Need This

The first reason is the Agent-Computer Interface idea itself. Most agent frameworks obsess over prompts and orchestration, while SWE-agent’s founding insight is that the shape of the interface, which commands exist, how their outputs are truncated, what the model sees after each action, matters as much as the model. The repository is the reference implementation of that argument. Reading how the windowed file viewer limits observation size, how linting feedback is attached to edits, and how the submit command packages a diff will change how you design any tool-using agent.

The second reason is the configuration discipline. The entire behavior of the system, system prompt, instance template, tool bundle list, environment variables inside the sandbox, history processing, and action parsing, is declared in one YAML file. That single-file governance makes A/B research trivial and makes the codebase a template for experiments: you can swap the editing interface without touching Python, change the model with one flag, and reproduce benchmark runs from a checked-in config. Few agent projects this capable are this reproducible.

The third reason is benchmark infrastructure. SWE-agent grew up alongside SWE-bench, the field’s standard evaluation of real issue fixing, and the batch runner reflects that maturity: dataset slices, deterministic shuffles, automatic instance loading, per-instance cost caps, and optional evaluation submission through the sb-cli tool. Whether or not you chase leaderboards, that harness, loading many problem statements, running them in parallel sandboxes, and collecting predictions and trajectories, is exactly what serious agent evaluation requires, and you can study every line of it.

How It Works

SWE-agent detailed architecture diagram

Inside SWE-agent: the dispatcher's subcommands, agent internals from history processors to retry loops, the tool layer with its bundles, and the sandboxed environment.

Understanding the Architecture

A dispatcher with a dozen surfaces. The main module feeds argv to a parser in run.py whose choices map to subcommands: run and run-batch for solving, run-replay for replaying a saved trajectory, traj-to-demo for converting a trajectory into a demonstration config, merge-preds, extract-pred, compare-runs, remove-unfinished, and quick-stats for working with prediction files, plus an inspector command for the web viewer. Imports are deferred until a subcommand is chosen, so startup stays fast. This one file is the map of everything the project can do.

One agent class, many guard rails. DefaultSWEAgent in agents.py holds the loop: format the instance template, query the model, parse an action, execute it in the environment, append the observation, and repeat until the model calls submit or a limit trips. Around that loop sit explicit guard rails as imported exception types: context-window-exceeded, cost-limit-exceeded, content-policy-violation, and format errors all have dedicated handling, and tenacity-style retries wrap the model call. History processors trim or annotate the message list between turns, and the default config enables the cache-control processor that keeps the last two messages fresh for providers with prompt caching.

Models, sampling, and cost accounting. The models module wraps LiteLLM behind a thin interface, which is how one config line selects Claude, GPT-4o, or an OpenAI-compatible endpoint. An action sampler module handles the retry loop for malformed or empty responses, escalating through strategies before giving up. Every call updates instance statistics, and per-instance and total cost limits are enforced by the agent, not by hope. There is even a human model class, used by the human mode configs so a person can step through the same loop, and a human-thought variant that asks for reasoning before actions.

Tools declared as data, installed as bundles. The tool layer builds the model-facing schema from command definitions: each bundle in the tools directory ships command documentation plus source files, and the tool handler combines them into a function-calling schema, a JSON schema, or a plain text format depending on the configured parser. The parse module implements those parsers, including the function-calling default and a thought-action parser for text models. A filter config blocks interactive commands such as vim and emacs that would hang a sandbox, with a templated error message the model can read and react to.

The windowed editing interface. The default bundle set installs the windowed viewer with linting, the Anthropic-style edit tool, a registry, and a review-on-submit macro. The windowed interface shows the model only a limited view of one file at a time, with commands to open, scroll, search, and edit by line range, and it runs linting after edits so mistakes surface immediately. This is the ACI thesis in miniature: the same model performs dramatically differently when its file access is designed for its context limits instead of mimicking a human shell session.

A sandbox managed by SWE-ReX. The environment module, swe_env.py, owns the computer: it starts a deployment through the companion SWE-ReX package, by default a Docker container, with cloud options such as Modal and AWS Fargate also supported, installs the tool bundles inside it, executes actions, captures observations with truncation, and computes diffs against the base commit. Repository setup and Git operations live in repo.py, and hooks on both the agent and the environment let researchers observe every step without modifying the loop.

Batch runs and trajectories. The batch runner loads instances from a source, with SWE-bench subsets as first-class options, then drives the agent across them with progress reporting via run hooks. Every run writes a trajectory file capturing each step’s messages, actions, and observations, plus prediction files in the SWE-bench format. The inspector serves those trajectories in a browser so you can watch exactly what the model saw and did, and run-replay can re-execute a recorded sequence of actions, which is invaluable for debugging surprising behavior.

End to end. One command deploys a sandbox, installs the tool interface, presents a GitHub issue formatted through the instance template, and lets the model drive a carefully designed shell until it submits a diff. Config comes from one YAML, model calls flow through LiteLLM with enforced cost limits, and every artifact, trajectory, prediction, demonstration, is written to disk for later study. The layering is textbook: run surfaces on top, agent in the middle, tools and environment below, nothing reaching past its neighbor.

Advantages

  • Research-grade results. State-of-the-art SWE-bench performance among open source projects, with the paper and config to reproduce the approach.
  • One-file governance. Prompts, tools, parsers, history handling, and environment variables are all declared in a single YAML config.
  • Engineered interface, not a raw shell. Windowed viewing, linted edits, and observation truncation encode the Agent-Computer Interface thesis directly in the tool layer.
  • Cost and context guard rails. Per-instance and total cost limits, context-window handling, and retry loops are built into the agent rather than bolted on.
  • Complete benchmark harness. Batch mode with dataset slicing, deterministic shuffling, prediction merging, and sb-cli evaluation covers the full evaluation workflow.
  • Inspectable by design. Trajectory files, a web inspector, and replay mode make agent behavior auditable after the fact.

Benefits

  • Learn interface design for agents. The tool bundles show how command shape, output size, and feedback loops change model performance.
  • Copy the evaluation harness. Batch instances, cost caps, and trajectory logging form a ready-made template for benchmarking your own agents.
  • Understand LiteLLM integration. One wrapper handles providers, pricing statistics, and retries, a pattern reusable in any multi-model project.
  • Study defensive execution. Sandboxed deployments, command blocklists, and human review flows demonstrate responsible agent execution end to end.
  • Debug agents like software. Replay, inspector, and demonstration configs turn inscrutable model behavior into reproducible engineering artifacts.
  • A gateway to the ecosystem. The same team ships mini-SWE-agent, SWE-ReX, and SWE-bench, so the concepts here transfer across the whole toolchain.

Usage

Install from source as the docs direct:

git clone https://github.com/SWE-agent/SWE-agent.git
cd SWE-agent
python -m pip install --upgrade pip && pip install --editable .

Set a key in your environment or a .env file:

export ANTHROPIC_API_KEY=your-key-here

Fix a real GitHub issue in one command:

sweagent run \
  --agent.model.name=claude-sonnet-4-20250514 \
  --agent.model.per_instance_cost_limit=2.00 \
  --env.repo.github_url=https://github.com/SWE-agent/test-repo \
  --problem_statement.github_url=https://github.com/SWE-agent/test-repo/issues/1

For benchmark-style batch runs over SWE-bench instances:

sweagent run-batch \
    --config config/default.yaml \
    --agent.model.name gpt-4o \
    --agent.model.per_instance_cost_limit 2.00 \
    --instances.type swe_bench \
    --instances.subset lite \
    --instances.split dev \
    --instances.slice :3 \
    --instances.shuffle=True

Run sweagent --help for the full subcommand list, and open the written trajectory with sweagent inspector to review what the agent did.

Conclusion

SWE-agent demonstrates that an agent is only as good as the interface between the model and the machine. Its YAML-governed configuration, engineered tool bundles, enforced cost limits, and complete trajectory tooling make it both a strong system in its own right and the clearest available case study in Agent-Computer Interface design. Even as the team’s energy shifts to the deliberately minimal mini-SWE-agent, the flagship repository remains worth a long read for anyone building agents that must survive contact with real codebases. Clone it, run the hello-world example inside its sandbox, and study the trajectory it leaves behind.

Links:

Watch PyShine on YouTube

Contents