Firecrawl: feeding a local RAG with pages web
Firecrawl is an open-source web crawler (AGPL-3.0 core, more than 186,000 GitHub stars in late September 2026) that turns a page or an entire site into clean Markdown or JSON for a model. It can be self-hosted with Docker Compose, but the default stack includes neither screenshots nor page actions nor the managed service's advanced anti-bot bypass: those features remain exclusive to Firecrawl Cloud.
Firecrawl turns a web address into clean text that a model can read directly: the menu, cookie banner, scripts, and footer disappear, while the article remains. It is the first step in a RAG pipeline—the step that determines the quality of everything that follows. The project is open and self-hostable, making it compatible with a fully local setup—provided you know exactly what the self-hosted version does not do compared with the online service of the same name.
#The problem it solves
Giving a local model a web page by pasting in its HTML code wastes most of the context window on tags, menus, and scripts. Worse, the model treats navigation elements as content and gives you an irrelevant answer. Every serious document-retrieval pipeline therefore starts with an extraction step that answers one question: what is the article on this page?
Firecrawl handles this work, along with the crawling logic around it: following internal links, stopping at a limit you set, executing JavaScript on sites that display nothing without it, and returning a structured result for each page. The output is Markdown or JSON, two formats that split cleanly into chunks for indexing. The project, hosted on GitHub under the name firecrawl/firecrawl and described by its authors as “the web data API to search, scrape, and interact at scale,” had more than 186,000 stars by the end of September 2026—a traction signal that matters when choosing a dependency to maintain over the long term.
#Scrape, crawl, map, extract: four different jobs
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
| Operation | What it does | When to use it |
|---|---|---|
| Scrape | An input address, cleaned content as output | You already know the exact page |
| Crawl | Crawls a site from a starting address by following internal links | Ingest an entire documentation set |
| Map | Returns a site's address list without retrieving its content | Decide what to ingest before spending time on it |
| Extract | Extracts fields from structured data according to a schema you provide | You want a pricing or specifications table, not prose |
Two habits prevent most unpleasant surprises: run a map before a crawl to see the site's structure before starting work on it; and set an explicit page limit for every crawl, because documentation with pagination and language variants is often ten times larger than it appears.
#Host it at home
This is the path that interests us if the goal is to keep the stack on your hardware. It consists of a Docker Compose stack with several services: the API, PostgreSQL as the default queue (an optional FoundationDB engine is available for those who need it), Redis, RabbitMQ, and Playwright as a headless browser service for pages that only exist after JavaScript runs. The official guide pins a specific repository version rather than the main branch—v2.11.162 in late September 2026—and warns that another version may use a different Compose file. The official command to start the stack is docker compose up --build -d, after which the API responds locally on port 3002, using the same request formats as the online version.
#What you need before launching the stack
According to the official documentation, four prerequisites are sufficient: Git to retrieve the repository, Docker Engine or Docker Desktop, Docker Compose v2 (invoked as docker compose, without a hyphen), and curl to check that the API responds once the stack is running. Port 3002 must be available, and the machine must have enough capacity to build and run several containers in parallel—without the project guaranteeing a specific amount of RAM or CPU.
- 01Read the project's self-hosting page firstIt is the only up-to-date source, and it explicitly states that quick startup on a trusted network disables API authentication and is not a production architecture.
- 02Settling the browser questionWithout browser rendering, you get static HTML: perfect for documentation, useless for a single-page web app that builds its content in JavaScript.
- 03Check the license against your use caseThe core project is licensed under AGPL-3.0, a strong copyleft free software license; the client SDKs are distributed under the MIT license. For a commercial product that redistributes modified code, read the license instead of assuming.
- 04Size the deploymentThe project does not publish a guaranteed minimum machine size for this stack: you are responsible for updates, secrets, storage, monitoring, and disaster recovery, unlike the managed service.
- 05Add durable storage before relying on itQuick start does not add any persistent volume: according to the documentation, it “starts without durable storage, TLS, high availability, or all of Firecrawl Cloud’s capabilities.” Without manually added volumes for PostgreSQL, Redis, and RabbitMQ, the data does not survive container replacement — address this before trusting a corpus ingested only once.
#Control Firecrawl from an agent: the official MCP server
To connect an MCP-compatible agent (Claude Code, Cursor, or an in-house client) to this self-hosted instance without writing an HTTP integration, the project publishes an official MCP server, firecrawl-mcp-server. With its full profile, it exposes up to 26 tools by default, including feedback. The FIRECRAWL_API_URL environment variable points it to your local deployment instead of the online service, with an optional API key if your instance requires one. This is the most direct way to give a local agent the ability to search, browse, and extract from the web without routing traffic through a third party—provided the agent itself runs on your machine or a trusted network.
#What the local stack does not have
The point many people discover after deploying the stack, not before: self-hosting by default does not cover everything the online service does. According to the official documentation, screenshots and page actions (clicking, scrolling, filling out a form before reading the result) are not available in the default stack. The same applies to the Agent and Browser features, the “interact” option, usage feedback, and certain specialized formats (menu, audio, video): these components remain reserved for Firecrawl Cloud. The advanced anti-bot bypass of the managed service's proprietary engine is not included as-is either.
One other useful nuance before getting started: model-assisted extraction (the “extract” mode driven by a natural-language instruction rather than a schema) requires you to connect a provider compatible with the OpenAI API, or Ollama, and test this path separately. It is not connected automatically in the self-hosted stack. None of this prevents you from using Firecrawl locally to feed a RAG—which is precisely this guide’s central use case—but it is better to know before promising a feature to a team.
#Where it fits in a local RAG
A working local pipeline has four stages, and Firecrawl handles only the first: ingestion transforms URLs into documents; chunking and encoding convert them into vectors using a local embedding model; a vector database stores these vectors and returns the passages closest to a question; finally, a local model writes the answer based on those passages.
For files already on your disk—PDFs, office documents, presentations—you don’t need a crawler but a document converter such as Docling. And if you need a model to consult the web when it responds, rather than use a corpus ingested in advance, a self-hosted metasearch engine such as SearXNG is the right building block: it returns results on demand instead of relying on an index that needs refreshing.
- Build an end-to-end local RAG with Python
- Which embedding model for French
- Give a local model web search with SearXNG
- For your PDFs and local files: Docling
- Store vectors in Qdrant
- The QuelLLM local RAG kit: all components on one page
- Source: the official Firecrawl repository on GitHub
- Source: Firecrawl self-hosting documentation
- Source: Firecrawl official MCP server
#When not to use it
| Your needs | What works better |
|---|---|
| Convert a page occasionally | A small HTML-to-Markdown library in your script |
| Ingest PDFs and local documents | A document converter like Docling, not a web crawler |
| Live web answers in a conversation | A self-hosted search engine connected to the model |
| Screenshots or complex page actions | The managed Firecrawl Cloud service, absent from the default stack |
| A complete documentation set, regularly and automatically | Self-hosted Firecrawl — that's exactly what it's for |
#The legal point to watch
Automatically collecting pages makes you responsible, not the tool. The terms of service of the target site, its robots.txt directives, applicable law, and how you use the content matter more than the technique employed. For an internal corporate corpus, this is compounded by the issue of processing any personal data present in the collected pages, and the need to document where each source came from in the processing register.
#FAQ
Is Firecrawl free for self-hosting?+
Does the self-hosted version do everything the online service does?+
Do you need a GPU for Firecrawl?+
Can it read sites that require JavaScript?+
What's the difference between crawling and mapping?+
Does AI extraction (extract mode) work directly locally?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.