Firecrawl reviewed: turning the live web into LLM-ready markdown for agents
Firecrawl is a web-data API that turns websites into machine-readable content for AI apps and agents — search, scrape pages into clean markdown/JSON/structured data, and interact with pages (click, navigate) — handling JavaScript rendering and dynamic content automatically.
What does Firecrawl do?
Firecrawl is a web-data API that turns live web pages into LLM-ready output — markdown, HTML, JSON, links, screenshots and summaries — through v2 endpoints named Search, Scrape, Interact, Agent, Crawl, Map and Parse (Firecrawl docs, 2026). The documentation says it “handles the hard stuff: proxies, anti-bot, JavaScript rendering, and dynamic content”, and the homepage states that Firecrawl renders JavaScript automatically, so full page content comes back even from single-page applications. SDKs cover Python and Node.js plus a CLI. The core is open source under AGPL-3.0 (SDKs and UI components under MIT) and can be self-hosted; on 7 September 2026 the GitHub repository showed 177.5k stars (GitHub, 2026).
Where does Firecrawl fit in the agentic web?
Unlike the measurement and analytics tools on this hub, Firecrawl is agent-side data infrastructure: what an agent, or an app building one, uses to consume the web. It produces on demand the same thing a site can serve natively, a clean markdown twin of a page. That symmetry is the point of readiness: a site that ships its own markdown twins lowers the Cost of Retrieval so tools like this do less work to read it. The gap is measured: Cloudflare's scan of the 200,000 most visited domains found markdown content negotiation on only 3.9 % of sites (Cloudflare, April 2026), so on the remaining 96.1 % a tool such as Firecrawl has to derive the markdown itself. Every page on agentswelcome.dev ships its twin (this one at /tools/firecrawl.md) plus a full-text /llms-full.txt.
Our take
The category leader for “give me this site as markdown”, and a useful mirror for site owners: it demonstrates exactly what an agent wants from your pages. If Firecrawl has to strip your layout to read you, so does every other agent; the fix is to serve the clean version yourself. Included here as the agent-side counterpart to the readiness this site advocates.
Does Firecrawl respect robots.txt?
Standard Firecrawl crawls honor robots.txt: the crawl documentation offers an option to “ignore the website's robots.txt rules” and marks it Enterprise only, and a robotsUserAgent parameter, also enterprise-only, sets the User-Agent under which robots.txt is evaluated (Firecrawl docs, 2026). The crawler reads sitemaps by default (sitemap: "include", with skip and only as alternatives) and exposes delay and maxConcurrency parameters to respect a site's rate limits. Two caveats: that documentation names no default User-Agent token, and the Almanac's AI Crawler & Agent Registry (41 records) has no Firecrawl entry, so this review found no documented token to write a robots.txt rule against.
Firecrawl — frequently asked questions
Is Firecrawl open source?
Yes. The core repository is licensed under AGPL-3.0, with SDKs and UI components under MIT, and a self-hosting guide is published; the cloud version adds features beyond the open-source offering (GitHub, 2026).
Which output formats does Firecrawl return?
Markdown, HTML, JSON, links, screenshots and summaries, according to the v2 documentation (Firecrawl docs, 2026).
Does Firecrawl render JavaScript?
Yes. Firecrawl states that it renders JavaScript automatically, so it returns full page content from single-page applications and dynamically loaded sites (Firecrawl, 2026).
Does Firecrawl respect robots.txt?
Standard crawls do: the option to ignore a website's robots.txt rules is marked Enterprise only in the crawl documentation (Firecrawl docs, 2026).
Is there a free Firecrawl plan?
Yes. As of 7 September 2026 the pricing page lists Free, Hobby, Standard, Growth, Scale and Enterprise tiers; Free includes 1,000 credits per month, described as 500 searches or 1,000 pages scraped, and Hobby starts at $16 per month billed yearly (Firecrawl, 2026).
Sources
- Firecrawl: Homepage (Search, Scrape, Interact; automatic JavaScript rendering), accessed 2026. firecrawl.dev
- Firecrawl: Documentation (v2 endpoints, output formats, SDKs), accessed 2026. docs.firecrawl.dev
- Firecrawl: Crawl documentation (robots.txt, sitemap, delay and maxConcurrency parameters), accessed 2026. docs.firecrawl.dev
- Firecrawl: Pricing (verified 2026-09-07), 2026. firecrawl.dev
- GitHub: firecrawl/firecrawl repository (AGPL-3.0; 177.5k stars on 2026-09-07), 2026. github.com
- Cloudflare: Introducing the Agent Readiness score, 2026. blog.cloudflare.com