The README of browser-use/browser-use opens with a dare: navigate the web like a human does. Its demo shows the agent finding an available slot, picking a date, handling the CAPTCHA, and booking a driving test, and its self-description is blunt about scope: the open-source browser agent in Python and TypeScript, with a commercial cloud on the side that rents a browser for about two cents an hour with stealth, CAPTCHA solving, and residential proxies. The library itself is MIT-licensed, version 0.13.11 at the time of writing, needs Python 3.11 through 3.13, and installs with a single uv add browser-use followed by an OpenAI key in a dotenv file. What began as a student project in 2024 became one of the most-starred repositories in the automation space, and the reason is visible in the code: this is not a Playwright wrapper with a prompt sprinkled on top, it is a full control system with watchdogs, event buses, cost accounting, and self-healing sessions.
The package layout under browser_use reads like a layered design document. The agent package holds the reasoning loop; the tools package holds the hands; the browser package holds the car and its safety monitors; the dom package is the sensory system that turns pixels and accessibility trees into something a language model can read; and the llm package adapts a dozen model providers behind one abstract base. Supporting cast includes a token cost service that fetches live OpenRouter pricing, a filesystem sandbox where the agent keeps its own notes and downloads, an MCP server that exposes the agent as callable tools for other assistants, and a skills service that can install new capabilities at runtime. As always in this series, what follows is an educational tour of published source code.
Browser Use at a glance: the CLI and MCP server feed tasks to the Agent loop, which reads the page through the DOM service and its serializer, decides on actions via the LLM base and its provider adapters, and executes them through the tool controller onto the BrowserSession, which runs watchdog monitors and can attach to a cloud browser.
Reading the overview from left to right:
- The front door is browser_use/cli.py, a dispatcher that runs the MCP server, registers skills, and migrates legacy commands.
- Tasks land in the Agent loop at browser_use/agent/service.py, which by default runs up to 500 steps per task.
- The tool controller at browser_use/tools/service.py registers every browser action the model can call, from clicking and typing to extracting structured data.
- Model calls flow through the abstract base at browser_use/llm/base.py, with adapters for OpenAI, Anthropic, Google, and the project’s own BU2 model.
- The BrowserSession core at browser_use/browser/session.py talks raw Chrome DevTools Protocol and hosts the watchdog monitors under browser_use/browser/watchdogs, including the CAPTCHA watcher at browser_use/browser/watchdogs/captcha_watchdog.py.
- Page perception lives in browser_use/dom/service.py, which builds an enhanced DOM from accessibility trees and serializes it for the model.
- Housekeeping runs through the token cost service at browser_use/tokens/service.py and the agent file system at browser_use/filesystem/file_system.py.
Why You Need This
The first reason is the perception layer, because the hardest part of browser automation is not clicking, it is seeing. The DOM service in browser_use/dom/service.py does something clever: instead of parsing raw HTML, it pulls the accessibility tree for every frame, including cross-origin iframes that pass size eligibility checks, merges it with the DOM tree and viewport geometry, and produces an EnhancedDOMTreeNode graph with visibility rules applied per parent. The serializer in browser_use/dom/serializer/serializer.py then turns that graph into indexed, clickable elements with their attributes, so the model receives a compact numbered menu of what it can act on rather than a wall of markup. There is even an browser_use/dom/enhanced_snapshot.py that captures richer CDP snapshots. If you have ever watched a web agent hallucinate a button that is not there, this file tree is the antidote.
The second reason is the BrowserSession, which is engineered like an operating system kernel for Chrome. The session at browser_use/browser/session.py wraps every operation as an event on a resilient event bus, handles navigation, tab switching, cookies, storage-state export, and proxy authentication callbacks, and can attach to a local Chrome it discovers itself, an existing browser over a CDP URL, or a hosted cloud browser through browser_use/browser/cloud/cloud.py. Around it sit fifteen watchdogs, each a small coroutine guarding one failure mode: crashes, downloads, popups, permissions, dialog boxes, blank tabs, and the CAPTCHA watchdog that parks the agent until a challenge is solved. When the WebSocket drops, an auto-reconnect path retries up to three times and replays state. Any team building long-running automation inherits all of this battle-testing for free.
The third reason is extensibility that goes three directions at once. Outward, the MCP server in browser_use/mcp/server.py implements handle_call_tool so Claude or any MCP client can drive a browser session as if it were a native tool. Inward, the skills service at browser_use/skills/service.py registers installed skills as first-class actions on the registry at browser_use/tools/registry/service.py. Downward, the tools controller in browser_use/tools/service.py ships around twenty actions including search, navigate, click by index or coordinate, upload files, dropdown handling, JavaScript evaluation, screenshots, PDF saving, and structured extraction, each with typed parameter models the LLM adapters turn into tool schemas automatically.
The detail view: entry points and config on top, the Agent loop with its message manager and prompt variants in the center, provider adapters on the left, the event-driven browser layer with watchdogs below, and DOM perception with the tool registry on the right.
Walking the detail diagram, the agent loop in browser_use/agent/service.py is the beating heart, and its step method is a textbook perception-action cycle: check for pause or stop signals, prepare the browser state summary through the DOM service, feed the message history through the manager in browser_use/agent/message_manager/service.py which trims and compacts old turns, call the model for the next action, execute the action batch, then post-process results. The loop is defensive in ways that only come from real usage: loop detection nudges the model when it repeats itself, budget warnings fire as token costs climb, failures are counted and escalate to forced done, and a separate judge module in browser_use/agent/judge.py can grade the final answer against the original task. Prompt variants live in browser_use/agent/system_prompts with tuned versions for flash mode, Anthropic, and no-thinking models, and configuration defaults come from browser_use/config.py.
The browser layer rewards the second look. The session constructor is a pydantic model, and nearly every user-facing operation is an event handler: on_NavigateToUrlEvent waits for the page to settle, on_SwitchTabEvent refocuses CDP targets, on_FileDownloadedEvent tracks files into the agent’s filesystem, and a CAPTCHA-aware wait method lets the whole session pause gracefully while the watchdog solves a challenge. The profile object in browser_use/browser/profile.py carries hundreds of launch options from proxy credentials to viewport size, while the chrome helper locates a suitable executable across platforms. The token cost service in browser_use/tokens/service.py deserves a mention too: it registers every LLM involved, including compaction models, and resolves per-token pricing live from OpenRouter when the model name matches, so cost summaries in the history are grounded in real rates rather than hard-coded tables.
From Install to a Booked Appointment
Getting from pip install to a finished task takes four lines of user code: create a ChatOpenAI with your model of choice, create an Agent with a task string, await agent.run, and print history.final_result. Underneath those lines the CLI dispatches, the session starts and attaches fifteen watchdogs, initial actions like a URL navigation run under a step timeout, and the loop begins its perception-action cycle, up to 500 steps or until the model emits done. The same agent is reachable from the command line, from the MCP server with a ten-minute idle session timeout, and from the hosted cloud where a task becomes an API call.
Honest limits: the README itself sells a hosted path, and the open-source loop assumes you bring an API key and accept the token bill of an agent that may reason over screenshots for hundreds of steps. CAPTCHA solving and stealth lean on the commercial cloud or local integrations rather than magic inside the MIT code. And because the DOM perception is snapshot-based, pages that fight accessibility tooling can still confuse the serializer. But as a reference implementation of how an agent should perceive, decide, and act on a live GUI, this repository is the one the others measure themselves against. Enjoyed this post? Never miss out on future posts by following us