Every AI agent eventually hits the same wall: the web was built for browsers, not for language models. Give an agent a URL and it gets JavaScript boilerplate, cookie banners, nav menus, and anti-bot walls — none of which fit in a context window. Firecrawl, the repository behind the firecrawl.dev service with one of the largest star counts on GitHub, attacks exactly this problem. It is an open-source web data API whose whole reason for existing is to turn any URL — or an entire domain — into clean markdown or structured JSON that an LLM can consume directly.

What makes Firecrawl worth more than a quick look is that it is not a thin wrapper around a headless browser. The main repository is a full production system: an Express HTTP API, a durable job queue backed by PostgreSQL, a crawl planner that respects sitemaps and robots-style constraints, a multi-engine scraping waterfall, and a transformer layer that converts raw documents into the formats agents actually request. It ships SDKs for JavaScript, Python, Go, Rust, Java, PHP, Ruby, Elixir, and .NET, and the whole stack can be self-hosted from the bundled Docker Compose file.

That combination — hosted-product maturity plus an AGPL-3.0 codebase you can read — is rare. Most “scraping API” projects stop at a fetch-plus-turndown script. Firecrawl’s source shows how the hard 10 percent is engineered: engine fallbacks when a page resists, retries and timeouts, credit billing, rate limiting, webhooks, and caching. Reading it is a free masterclass in building reliable data infrastructure for AI, which is exactly why we are taking the tour.

Architecture overview of the firecrawl/firecrawl repository

Overview of the Firecrawl architecture: SDKs hit the /v2 REST routes, controllers persist crawl plans and enqueue jobs, NuQ workers run the scrapeURL orchestrator, and the engine waterfall plus transformers produce LLM-ready output.

Reading the overview from left to right: a client SDK (or plain cURL) sends a request to the REST routes in apps/api/src/routes/v2.ts; the controllers under apps/api/src/controllers validate it with zod and, for a crawl, persist a plan via apps/api/src/lib/crawl-redis.ts; jobs flow through apps/api/src/services/queue-jobs.ts into the NuQ queue backed by apps/nuq-postgres; workers in apps/api/src/services/queue-worker.ts pick jobs up and invoke the scrapeURL orchestrator, which drives the engine waterfall — including the Fire-engine client for JS-heavy pages — and finally hands the raw document to the transformers that emit markdown and structured JSON.

Why You Need This

If you have ever built a RAG pipeline or an agent that browses, you already know the failure modes. A naive fetch gets you an empty shell for any single-page application. A headless browser gets you the page, but also nav bars, legal text, and 40 KB of tracking scripts that burn tokens and dilute retrieval quality. Firecrawl’s scrape endpoint exists to absorb all of that: you POST a URL to /v2/scrape and get back a document whose markdown is already the main content, with metadata, screenshots, or structured extraction available as optional formats.

The second problem is breadth, not depth. Agents rarely need one page; they need a whole docs site, a product catalog, or every changelog entry since last year. Writing a crawler that discovers links, stays inside the domain, respects limits, and recovers from failures is a project of its own. Firecrawl’s /v2/crawl endpoint does this: the crawl controller stores a plan, a Crawler object computes allowed links, and the queue fans out scrape jobs in parallel while webhooks and status endpoints report progress.

The third problem is search. An agent that must “find the right sources” needs a search backend whose results arrive pre-scraped, not as ten blue links. Firecrawl’s /v2/search endpoint searches the web — pluggable across backends such as its hosted engine, SearxNG, or DuckDuckGo — and then pushes the top results through the same scrape pipeline, so what lands in your agent’s context is page content, not just URLs.

Finally, there is the operational reality: proxies, rate limits, stale caches, and billing all have to live somewhere. Self-hosters get the same machinery the cloud service runs — rate limiting, credit billing, blocklists, and worker autoscaling knobs — which means the code you deploy is the code you read. For teams that cannot ship customer data through a third-party API, that is the difference between usable and off-limits.

How It Works

At its core, Firecrawl is a pipeline that turns an HTTP request into a queued job, the job into an engine-driven scrape, and the scrape into a formatted document.

Detailed architecture of the firecrawl/firecrawl repository

Detailed view of firecrawl/firecrawl: the Express API and v2 controllers, the NuQ/BullMQ job orchestration layer, the scrapeURL engine waterfall, the crawl and sitemap machinery, the search backends, and the Redis/PostgreSQL infrastructure that holds it together.

Understanding the Architecture

The API layer. The process starts in apps/api/src/index.ts, an Express application with WebSocket support that mounts the v1 and v2 routers, an admin router, and a Bull Board dashboard for queue inspection. Requests are parsed by body-parser with raw-body capture, throttled by the Redis-backed rate limiter in apps/api/src/services/rate-limiter.ts, and handed to controllers such as apps/api/src/controllers/v2/scrape.ts, crawl.ts, map.ts, search.ts, and batch-scrape.ts. Every payload is validated and typed with zod schemas defined in apps/api/src/controllers/v2/types.ts, including the central Document shape every job ultimately produces.

The job orchestration. Synchronous-looking work is deliberately async underneath. The crawl controller in apps/api/src/controllers/v2/crawl.ts mints a UUIDv7 job id, saves a StoredCrawl via apps/api/src/lib/crawl-redis.ts, and enqueues the first job through apps/api/src/services/queue-jobs.ts. Scrapes are queued on NuQ — a queue implemented in apps/api/src/services/worker/nuq.ts with its schema in apps/nuq-postgres (FoundationDB is available as an alternative backend via NUQ_BACKEND) — while side queues for webhooks, billing, and indexing run on BullMQ, orchestrated by apps/api/src/services/queue-worker.ts and the scrape runner in apps/api/src/services/worker/scrape-worker.ts. Workers bill credits through apps/api/src/services/billing/credit_billing.ts and push results to customer endpoints via the webhook service in apps/api/src/services/webhook.

The crawl planner. For site-wide work, the stored crawl is converted into a Crawler object — crawlToCrawler in crawl-redis.ts builds it from the options defined in apps/api/src/scraper/WebScraper/crawler.ts, which handles include/exclude patterns, max depth, limit checks, and link filtering. Sitemap support in apps/api/src/scraper/crawler/sitemap.ts lets /v2/map and crawls discover URLs quickly before individual jobs are enqueued, and the worker loop keeps re-enqueueing discovered links until the crawl finishes.

The engine waterfall. The heart of a single scrape is apps/api/src/scraper/scrapeURL/index.ts. The orchestrator builds a fallback list in apps/api/src/scraper/scrapeURL/engines/index.ts and tries engines in order: the index cache engine first, then the Fire-engine variants — chrome-cdp and tlsclient, with stealth modes — for pages that need a real browser or TLS fingerprint work, then the Playwright engine driving the bundled microservice in apps/playwright-service-ts, and finally plain fetch, with dedicated engines for PDFs and other document types. Each engine declares feature flags (actions, waitFor, screenshots, mobile, location), so the orchestrator only routes a URL to engines that can honor the request, and errors flow through a retry tracker so the next engine gets a clean chance.

The transformers. Once an engine returns raw content, apps/api/src/scraper/scrapeURL/transformers/index.ts applies the requested formats. Markdown conversion goes through parseMarkdown in apps/api/src/lib/html-to-markdown.ts, structured extraction uses the LLM-powered llmExtract.ts transformer to bend page content into a JSON schema, and additional transformers handle screenshots, base64 image removal, PII redaction, and more. The output is the zod-validated Document that the API serializes back to the client — the same shape whether you scraped one URL or ten thousand.

Search and map as first-class citizens. The search controller delegates to apps/api/src/search/v2/index.ts, which picks a backend (the hosted search engine, SearxNG, or DuckDuckGo) and then runs each result through apps/api/src/search/scrape.ts, reusing the very same scrapeURL pipeline. That reuse is the architectural thesis of the whole repo: one battle-tested scrape core, with search, crawl, map, and batch endpoints as thin, reliable fronts over it.

Following one /v2/crawl request end to end: the controller validates the payload and saves a crawl plan in Redis, enqueues a first job on NuQ, a worker dequeues it and runs scrapeURL, the engine waterfall fetches and renders the page, transformers emit markdown, the worker bills credits, fires webhooks, discovers the next allowed links with the Crawler, enqueues them, and repeats — until the limit is reached and the client collects every document through the status endpoint or its SDK.

Advantages

  • LLM-ready output by design. The transformer pipeline exists to produce clean markdown, screenshots, and schema-guided JSON — not raw HTML — so tokens go to content instead of boilerplate.
  • A real engine waterfall. Cached index hits, Fire-engine browser/TLS modes, Playwright, and plain fetch are tried in a principled order with feature-flag routing, which is how the system keeps succeeding where a single engine would fail.
  • Durable, inspectable job orchestration. NuQ on PostgreSQL (with a FoundationDB option) plus BullMQ side queues means crawls survive restarts, and Bull Board exposes queue health to operators.
  • Breadth of endpoints. Search, scrape, crawl, map, batch scrape, extract, and more are all backed by one shared scrape core, so behavior stays consistent across the surface.
  • Serious self-hosting story. The root docker-compose.yaml ships the API, workers, Playwright service, Redis, RabbitMQ, NuQ PostgreSQL, and optional FoundationDB, with the published guide walking through an unauthenticated baseline.
  • SDKs everywhere. JavaScript, Python, Go, Rust, Java, PHP, Ruby, Elixir, and .NET SDKs live in the same monorepo, each handling async polling so callers write three lines of code.

Benefits

  • Faster agent loops. Pre-scraped search results and crawl endpoints mean your agent spends its turns reasoning over content rather than wrestling with pages.
  • Lower token spend. Main-content markdown with unwanted elements removed directly reduces the context budget every downstream LLM call consumes.
  • Operational control. Self-hosting keeps scraped data, logs, and billing inside your own infrastructure — a hard requirement for many compliance-sensitive teams.
  • Extensibility without forking the world. Adding an engine, a transformer, or a search backend means implementing one module and registering it in a list, not rewriting the pipeline.
  • Production lessons for free. Rate limiting, blocklists, credit billing, webhooks, retries, and observability are all in the tree, ready to study or reuse in your own systems.
  • A permissive-path codebase. The core is AGPL-3.0 with SDKs under MIT, so embedding the client libraries in your product is straightforward while the server stays open.

Usage

The quickest way in is an SDK. From the repository README, install the Python package and scrape, crawl, and search in a few lines:

pip install firecrawl-py
from firecrawl import Firecrawl

app = Firecrawl(api_key="fc-YOUR_API_KEY")

# Scrape a single URL
doc = app.scrape("https://firecrawl.dev", formats=["markdown"])
print(doc.markdown)

# Crawl a website (automatically waits for completion)
docs = app.crawl("https://docs.firecrawl.dev", limit=50)
for doc in docs.data:
    print(doc.metadata.source_url, doc.markdown[:100])

# Search the web
results = app.search("best AI data tools 2024", limit=10)
print(results)

The JavaScript SDK mirrors the same surface:

npm install firecrawl
import { Firecrawl } from 'firecrawl';

const app = new Firecrawl({ apiKey: 'fc-YOUR_API_KEY' });

const doc = await app.scrape('https://firecrawl.dev', { formats: ['markdown'] });
console.log(doc.markdown);

Raw HTTP works too — a scrape is a single POST:

curl -X POST 'https://api.firecrawl.dev/v2/scrape' \
  -H 'Authorization: Bearer fc-YOUR_API_KEY' \
  -H 'Content-Type: application/json' \
  -d '{
    "url": "firecrawl.dev"
  }'

And a site-wide crawl with a limit and output format:

curl -X POST 'https://api.firecrawl.dev/v2/crawl' \
  -H 'Authorization: Bearer fc-YOUR_API_KEY' \
  -H 'Content-Type: application/json' \
  -d '{
    "url": "https://docs.firecrawl.dev",
    "limit": 100,
    "scrapeOptions": {
      "formats": ["markdown"]
    }
  }'

To run the stack yourself, clone the repository and use the root Docker Compose file, which brings up the API, workers, Playwright service, Redis, RabbitMQ, and NuQ PostgreSQL (the API is published on port 3002, and USE_DB_AUTHENTICATION=false gives you an unauthenticated baseline to start from):

git clone https://github.com/firecrawl/firecrawl.git
cd firecrawl
docker compose up -d

Conclusion

Firecrawl is what happens when a scraping utility grows into infrastructure: a zod-validated API layer, a durable queue, a crawl planner, a fallback-driven engine waterfall, and a transformer suite whose only output contract is “ready for your LLM.” The source tour shows a codebase that earns its popularity — every hard edge of web data extraction has a named module, a tested fallback, and an operator knob. If your agents touch the web, reading (or self-hosting) this repository is time well spent.

Links:

Watch PyShine on YouTube

Contents