BestLLMfor Your hardware. Your LLM. Your call.
◆ The kits◆ Kits APIOpen data Find my LLM
Guide · 2026-09-28

Firecrawl, Self-Hosted: Turning Websites Into Clean Text for a Local LLM

◆ Local RAG — Ask questions to your own documents with a local AI, no cloud · $24 · or all kits $49 →

Firecrawl converts a URL into clean markdown a model can read. Running it yourself keeps the pipeline local. What it does, how the self-hosted build differs from the cloud, and where it fits in a RAG stack.

By Mohamed Meguedmi·Last updated 2026-09-28·11 min read·Tested on Windows, macOS, Linux

Key takeaways

  • Firecrawl takes a URL and returns clean markdown or JSON instead of raw HTML: navigation, cookie banners, scripts and footers stripped out, leaving the text a language model should actually read.
  • The core is open source under AGPL-3.0 (client SDKs under MIT) and the project counted more than 186,000 GitHub stars in late September 2026 — a large, actively maintained dependency to build on.
  • Self-hosting runs on Docker Compose with five services: the API, PostgreSQL as the default queue backend, Redis, RabbitMQ, and Playwright for JavaScript rendering. The official quickstart pins a specific release, v2.11.162 as of late September 2026.
  • The quickstart stores nothing durably by default. Firecrawl's own docs say it "starts without durable storage, TLS, high availability" — add volumes yourself or lose a crawled corpus on the next container rebuild.
  • An official MCP server lets an agent drive a self-hosted instance directly through the FIRECRAWL_API_URL environment variable, with no custom integration code.

The problem it solves

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • 30-day refund

Pasting a web page's raw HTML into a local model wastes most of the context window on markup, navigation and scripts, and it confuses the model into treating menu links as content it should answer from. Every retrieval pipeline that reads the live web needs an extraction step first, one that answers a single question for each page: what here is actually the article? Firecrawl exists to answer that question at scale, not just for one page but for a whole site, and to hand back a result that chunks cleanly for indexing rather than a wall of tags.

The project is hosted on GitHub as firecrawl/firecrawl, describes itself as "the web data API to search, scrape, and interact at scale," and had close to 186,000 stargazers at the end of September 2026 — the kind of traction that matters when you are picking a dependency you intend to keep running for years rather than months. The core is licensed AGPL-3.0, a strong-copyleft open license; the official client SDKs ship separately under MIT, which matters if you are embedding either piece in a commercial product.

Scrape, crawl, map, extract: four different jobs

OperationWhat it doesUse it when
ScrapeOne URL in, cleaned content outYou already know the exact page you want
CrawlWalks a site from a starting URL, following internal links up to a limit you setIngesting a whole documentation site or knowledge base
MapReturns the list of URLs on a site without fetching the contentDeciding what to ingest before spending time and bandwidth
ExtractPulls structured fields out of a page against a schema you provideYou want a table of prices or specs, not prose

Two habits save most of the trouble in practice. Run map before crawl so you see the shape of a site before committing to a job against it, and set an explicit page limit on every crawl — a documentation site with pagination and language variants is routinely an order of magnitude bigger than it looks from the front page.

Self-hosting it: what actually runs

Self-hosting is the route that matters if the goal is to keep the whole pipeline on your own hardware. It is not a single container: the official Docker Compose stack brings up the API, PostgreSQL as the default queue backend (an optional FoundationDB backend exists for teams that need it), Redis, RabbitMQ, and a Playwright browser service for pages that only render once JavaScript runs. The self-hosting guide pins a specific release — v2.11.162 as of late September 2026 — and warns that a different tag can ship a different Compose file, so treat the version as a decision, not a default. Once the stack is up, the API answers locally on port 3002 with the same request shapes as the hosted service, started with a single docker compose up --build -d.

Before running that command, the project lists four prerequisites: Git to pull the repository, Docker Engine or Docker Desktop, Docker Compose v2 (invoked as docker compose), and curl to verify the API once it is up. Port 3002 needs to be free, and the docs are candid that they publish no verified minimum host size — you find out by trying.

  • Read the self-hosting page first, not a blog post about it. It is the only source that stays current, and it says plainly that the trusted-network quickstart disables API authentication and "is not a production architecture."
  • Decide about the browser service deliberately. Without Playwright rendering you get static HTML only — fine for most documentation, wrong for single-page applications that build their content client-side.
  • Do not trust the quickstart with data you cannot lose. It "starts without durable storage, TLS, high availability, or every Firecrawl Cloud capability." Add persistent volumes for PostgreSQL, Redis and RabbitMQ yourself, and test a restore, before you point a production crawl at it.
  • Do not trust the health check either. The /v0/health/readiness endpoint is a heartbeat, not an end-to-end test: Firecrawl's own troubleshooting guide says it "does not check Redis, PostgreSQL, RabbitMQ, Playwright, workers, or outbound network access." A green health check with a failing /v2/scrape call means the next stop is the API and Playwright container logs, not the readiness probe.

Self-hosting is not the same as anonymous. The crawler still fetches pages from your IP address with an identifiable user agent, and it is still bound by the target site's terms of service and robots.txt directives. Running the software on your own machine changes where the text is processed, not whether you are allowed to collect it.

Driving it from an agent: the official MCP server

Firecrawl also ships an official Model Context Protocol server, firecrawl-mcp-server, which is how a coding agent such as Claude Code or Cursor gets scrape, search and crawl abilities without any custom HTTP glue. With its full profile and default settings — feedback tools included — the server lists 26 tools; a hosted keyless endpoint trims that to three core tools for clients with a small tool-slot budget. For a self-hosted deployment, the relevant setting is the FIRECRAWL_API_URL environment variable, documented as the "custom API endpoint for self-hosted instances": point it at your own stack instead of the cloud, add an API key only if your instance requires one, and the agent talks to hardware you control end to end. It is the most direct way to give a local agent live web access without routing that traffic through a third party — as long as the agent itself also runs on your machine or a trusted network.

Where it sits in a local RAG stack

A working local pipeline has four stages, and Firecrawl only covers the first one:

  1. Ingest — Firecrawl turns URLs into markdown or JSON documents.
  2. Chunk and embed — documents are split and converted into vectors by an embedding model running locally.
  3. Store and retrieve — the vectors go into a vector database, which returns the passages closest to a question. See our guides on pgvector and FAISS for two very different ways to do that step.
  4. Generate — a local model answers using those passages, through Ollama, vLLM, or llama.cpp.

For files that already live on disk — PDFs, Office documents, slide decks — the ingestion stage is a different tool entirely: a document converter such as Docling or MarkItDown, not a web crawler. And if the goal is a model that checks the live web while it answers, rather than a corpus ingested ahead of time, a self-hosted metasearch engine is the better building block: it returns results on demand instead of an index you have to keep refreshing. Readers assembling a full local RAG stack from these pieces can start from our local RAG toolkit, which lines up the ingestion, storage and generation choices side by side.

When not to use it

Your needBetter fit
Convert one page, occasionallyA single-purpose HTML-to-markdown library in your script
Ingest local PDFs, DOCX, slidesA document conversion tool, not a web crawler
Live web answers inside a chatA self-hosted search engine wired to your model
Screenshots or complex page actionsFirecrawl Cloud — not part of the default self-hosted stack
A full documentation site, repeatedly, on a scheduleFirecrawl, self-hosted — this is exactly its case

Automated collection is your responsibility, not the tool's. The target site's terms of service, its robots.txt directives, applicable law, and what you do with the content afterward matter far more than which crawler you ran. For an internal company corpus, add the question of personal data that might appear inside collected pages, and keep a record of where each source came from — self-hosting moves where the processing happens, not who is accountable for it.

Verdict

Firecrawl earns its place when web content is a real, recurring input to your local AI stack rather than an occasional copy-paste: it handles the unglamorous parts — link discovery, JavaScript rendering, a consistent clean output format — that a homemade scraper gets wrong on the second site you try. Self-host it with clear eyes: the default stack ships without durable storage, without the hosted service's screenshot and anti-bot features, and without anyone to call when a container falls over. Add persistent volumes before you rely on it, pin your version, and it is a solid, free ingestion layer for a pipeline that stays entirely on your own hardware.

Sources

Frequently asked questions

Is Firecrawl free if I self-host it?

The core is open source under AGPL-3.0 and runs on your own hardware at no license cost; the client SDKs are MIT-licensed separately. You still pay in operating effort: it is a five-service Docker Compose stack — API, PostgreSQL, Redis, RabbitMQ, Playwright — that you patch, monitor and back up yourself, with none of the hosted service's support.

Does the self-hosted version do everything the cloud one does?

No. Firecrawl's own documentation lists what the default stack skips: screenshots, page actions like clicking or scrolling, the Agent and Browser features, Fire-engine's advanced anti-bot handling, and a few specialized formats. The core operations — scrape, crawl, map and extract — work fully; the managed extras stay on Firecrawl Cloud.

Do I need a GPU to run Firecrawl?

No. Crawling the web, rendering JavaScript through the Playwright service, and converting HTML to markdown are CPU, memory and network work, not GPU work. Plan for enough CPU and RAM to run five containers at once. The GPU only becomes relevant further down the pipeline, for the embedding model and the local model that generates the final answer.

Can Firecrawl scrape pages that need JavaScript?

Yes, because the default stack already includes a dedicated Playwright service for browser rendering, which executes a page's JavaScript before extraction runs. Without that service active, Firecrawl only fetches the static HTML the server sends first, which fails on single-page applications that build their content client-side after load.

Is scraping a website with Firecrawl legal?

That depends on the target site's terms of service, its robots.txt directives, the jurisdiction you operate in, and what you do with the content afterward — not on which tool did the fetching. Self-hosting changes where the processing happens; it changes nothing about whether you were permitted to collect the pages in the first place.

What is the difference between crawl and map?

Map lists the URLs it finds on a site without fetching their content, so it is fast and cheap — ideal for sizing a site before committing. Crawl actually retrieves and converts every page it follows. Mapping first, then crawling a filtered subset of the results, is how you avoid ingesting hundreds of pages you never wanted.

Can an AI agent use my self-hosted Firecrawl directly?

Yes, through the official firecrawl-mcp-server. Set the FIRECRAWL_API_URL environment variable to point at your own deployment instead of the cloud endpoint, add an API key only if your instance requires one, and an MCP-compatible agent such as Claude Code or Cursor gets scrape, search, crawl and map tools backed entirely by your own containers rather than a third-party service.

Recommended hardware

A current option for local AI: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395). Match memory to your model and software. A mini PC is a complete PC alternative; Mac/MLX and CUDA instructions require compatible hardware.

Amazon Check GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) price →

As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.

Did this guide help?

Found an error or have feedback? Let us know — it helps everyone who reads this guide.