Intermediate 11 minStack

Firecrawl: feeding a local RAG with pages web

Direct response

Firecrawl is an open-source web crawler (AGPL-3.0 core, more than 186,000 GitHub stars in late September 2026) that turns a page or an entire site into clean Markdown or JSON for a model. It can be self-hosted with Docker Compose, but the default stack includes neither screenshots nor page actions nor the managed service's advanced anti-bot bypass: those features remain exclusive to Firecrawl Cloud.

Firecrawl turns a web address into clean text that a model can read directly: the menu, cookie banner, scripts, and footer disappear, while the article remains. It is the first step in a RAG pipeline—the step that determines the quality of everything that follows. The project is open and self-hostable, making it compatible with a fully local setup—provided you know exactly what the self-hosted version does not do compared with the online service of the same name.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#The problem it solves

Giving a local model a web page by pasting in its HTML code wastes most of the context window on tags, menus, and scripts. Worse, the model treats navigation elements as content and gives you an irrelevant answer. Every serious document-retrieval pipeline therefore starts with an extraction step that answers one question: what is the article on this page?

Firecrawl handles this work, along with the crawling logic around it: following internal links, stopping at a limit you set, executing JavaScript on sites that display nothing without it, and returning a structured result for each page. The output is Markdown or JSON, two formats that split cleanly into chunks for indexing. The project, hosted on GitHub under the name firecrawl/firecrawl and described by its authors as “the web data API to search, scrape, and interact at scale,” had more than 186,000 stars by the end of September 2026—a traction signal that matters when choosing a dependency to maintain over the long term.

#Scrape, crawl, map, extract: four different jobs

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
The four operations and their real-world use
OperationWhat it doesWhen to use it
ScrapeAn input address, cleaned content as outputYou already know the exact page
CrawlCrawls a site from a starting address by following internal linksIngest an entire documentation set
MapReturns a site's address list without retrieving its contentDecide what to ingest before spending time on it
ExtractExtracts fields from structured data according to a schema you provideYou want a pricing or specifications table, not prose

Two habits prevent most unpleasant surprises: run a map before a crawl to see the site's structure before starting work on it; and set an explicit page limit for every crawl, because documentation with pagination and language variants is often ten times larger than it appears.

#Host it at home

This is the path that interests us if the goal is to keep the stack on your hardware. It consists of a Docker Compose stack with several services: the API, PostgreSQL as the default queue (an optional FoundationDB engine is available for those who need it), Redis, RabbitMQ, and Playwright as a headless browser service for pages that only exist after JavaScript runs. The official guide pins a specific repository version rather than the main branch—v2.11.162 in late September 2026—and warns that another version may use a different Compose file. The official command to start the stack is docker compose up --build -d, after which the API responds locally on port 3002, using the same request formats as the online version.

#What you need before launching the stack

According to the official documentation, four prerequisites are sufficient: Git to retrieve the repository, Docker Engine or Docker Desktop, Docker Compose v2 (invoked as docker compose, without a hyphen), and curl to check that the API responds once the stack is running. Port 3002 must be available, and the machine must have enough capacity to build and run several containers in parallel—without the project guaranteeing a specific amount of RAM or CPU.

  1. 01
    Read the project's self-hosting page first
    It is the only up-to-date source, and it explicitly states that quick startup on a trusted network disables API authentication and is not a production architecture.
  2. 02
    Settling the browser question
    Without browser rendering, you get static HTML: perfect for documentation, useless for a single-page web app that builds its content in JavaScript.
  3. 03
    Check the license against your use case
    The core project is licensed under AGPL-3.0, a strong copyleft free software license; the client SDKs are distributed under the MIT license. For a commercial product that redistributes modified code, read the license instead of assuming.
  4. 04
    Size the deployment
    The project does not publish a guaranteed minimum machine size for this stack: you are responsible for updates, secrets, storage, monitoring, and disaster recovery, unlike the managed service.
  5. 05
    Add durable storage before relying on it
    Quick start does not add any persistent volume: according to the documentation, it “starts without durable storage, TLS, high availability, or all of Firecrawl Cloud’s capabilities.” Without manually added volumes for PostgreSQL, Redis, and RabbitMQ, the data does not survive container replacement — address this before trusting a corpus ingested only once.
!
Self-hosting does not mean anonymity
The robot fetches pages from your IP address with an identifiable user agent, and remains subject to the terms of use of the target site and its robots.txt file. Hosting the software changes where the text is processed, not your right to collect it.
→
A “healthy” service can still fail
The API's /v0/health/readiness check is a heartbeat, not an end-to-end test: the official documentation states that it “checks neither Redis, PostgreSQL, RabbitMQ, Playwright, the workers, nor outbound network access.” If a call to /v2/scrape fails while this check is green, head to the api and playwright-service container logs rather than the health check.

#Control Firecrawl from an agent: the official MCP server

To connect an MCP-compatible agent (Claude Code, Cursor, or an in-house client) to this self-hosted instance without writing an HTTP integration, the project publishes an official MCP server, firecrawl-mcp-server. With its full profile, it exposes up to 26 tools by default, including feedback. The FIRECRAWL_API_URL environment variable points it to your local deployment instead of the online service, with an optional API key if your instance requires one. This is the most direct way to give a local agent the ability to search, browse, and extract from the web without routing traffic through a third party—provided the agent itself runs on your machine or a trusted network.

#What the local stack does not have

The point many people discover after deploying the stack, not before: self-hosting by default does not cover everything the online service does. According to the official documentation, screenshots and page actions (clicking, scrolling, filling out a form before reading the result) are not available in the default stack. The same applies to the Agent and Browser features, the “interact” option, usage feedback, and certain specialized formats (menu, audio, video): these components remain reserved for Firecrawl Cloud. The advanced anti-bot bypass of the managed service's proprietary engine is not included as-is either.

One other useful nuance before getting started: model-assisted extraction (the “extract” mode driven by a natural-language instruction rather than a schema) requires you to connect a provider compatible with the OpenAI API, or Ollama, and test this path separately. It is not connected automatically in the self-hosted stack. None of this prevents you from using Firecrawl locally to feed a RAG—which is precisely this guide’s central use case—but it is better to know before promising a feature to a team.

#Where it fits in a local RAG

A working local pipeline has four stages, and Firecrawl handles only the first: ingestion transforms URLs into documents; chunking and encoding convert them into vectors using a local embedding model; a vector database stores these vectors and returns the passages closest to a question; finally, a local model writes the answer based on those passages.

For files already on your disk—PDFs, office documents, presentations—you don’t need a crawler but a document converter such as Docling. And if you need a model to consult the web when it responds, rather than use a corpus ingested in advance, a self-hosted metasearch engine such as SearXNG is the right building block: it returns results on demand instead of relying on an index that needs refreshing.

#When not to use it

The right tool for the need
Your needsWhat works better
Convert a page occasionallyA small HTML-to-Markdown library in your script
Ingest PDFs and local documentsA document converter like Docling, not a web crawler
Live web answers in a conversationA self-hosted search engine connected to the model
Screenshots or complex page actionsThe managed Firecrawl Cloud service, absent from the default stack
A complete documentation set, regularly and automaticallySelf-hosted Firecrawl — that's exactly what it's for

#The legal point to watch

Automatically collecting pages makes you responsible, not the tool. The terms of service of the target site, its robots.txt directives, applicable law, and how you use the content matter more than the technique employed. For an internal corporate corpus, this is compounded by the issue of processing any personal data present in the collected pages, and the need to document where each source came from in the processing register.

#FAQ

Is Firecrawl free for self-hosting?+
The core code is open under the AGPL-3.0 license (client SDKs under the MIT license) and runs on your hardware with no licensing cost. You do pay in operating time, however: it is a stack of several containers—API, worker, queue, and headless browser—that you maintain yourself, without managed-service support.
Does the self-hosted version do everything the online service does?+
No. The official documentation explicitly lists what is missing from the default stack: screenshots, page actions, Agent and Browser features, advanced anti-bot bypass, and certain specialized formats. The core—retrieving and converting pages through scrape, crawl, map, and extract—works fully.
Do you need a GPU for Firecrawl?+
No. Web browsing, rendering JavaScript through the Playwright service, and converting HTML to Markdown are CPU, memory, and network tasks, not GPU compute. Instead, rely on available CPU and RAM to run multiple containers in parallel. The GPU only comes into play later in the pipeline, for embedding and response generation by the local model that then uses the extracted text.
Can it read sites that require JavaScript?+
Yes: the default stack includes a dedicated Playwright service for browser rendering, which executes the page’s JavaScript before extraction. Without this service running, Firecrawl retrieves only the static HTML delivered by the server, which fails on single-page applications that build their content client-side after the initial load. If the page is blank, first check the playwright-service container logs.
What's the difference between crawling and mapping?+
Map lists the addresses found on a site without retrieving their contents: it’s fast and lightweight, making it ideal for assessing a site’s size before committing. Crawl actually retrieves and converts each followed page. Mapping first, then crawling a filtered subset, prevents you from ingesting hundreds of pages you didn’t want.
Does AI extraction (extract mode) work directly locally?+
Not without configuration: instruction-driven extraction in natural language requires you to connect a provider compatible with the OpenAI API or Ollama, then test this path separately. It is not connected by default in the self-hosted stack, unlike the online service.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.