# www.firecrawl.dev > AI-optimized mirror of www.firecrawl.dev containing 50 pages totalling 48,463 words of clean markdown content, structured data, and semantic HTML. Original source: https://www.firecrawl.dev/. Last updated: 2026-04-30T23:43:59.676Z. Each page is available as HTML (with JSON-LD structured data) and Markdown (text-only, ideal for LLMs and RAG). ## Homepage - [Power AI agents with clean web data](/site-root.html): The API to search, scrape, and interact with the web at scale. Power AI agents with clean web data. Firecrawl delivers the entire internet to AI agents and builders. (2,368 words) ## Articles & Blog Posts - [Introducing /parse](/blog/index.html): The API to search, scrape, and interact with the web at scale. Power AI agents with clean web data. Firecrawl delivers the entire internet to AI agents and builders. (2,114 words) - [Web Data for any AI-powered solution](/use-cases/index.html): The API to search, scrape, and interact with the web at scale. Power AI agents with clean web data. Firecrawl delivers the entire internet to AI agents and builders. (1,015 words) - [Live Web Data with MCP](/use-cases/ai-mcps/index.html): Give Claude Code, Cursor, and Windsurf access to the live web with Firecrawl's official MCP server. Built for AI agents and developers. Get clean, LLM-ready data in under 3 minutes. Works with any MCP-compatible code editor. (1,659 words) - [Real-Time Agents & Deep Research](/use-cases/deep-research/index.html): Use Firecrawl’s search, scrape, and interact APIs to power deep research agents. Collect sources across thousands of sites, keep citations attached, and refresh datasets on a schedule for competitive analysis, market research, and technical investigations. (1,520 words) - [Flexible pricing](/pricing/index.html): The API to search, scrape, and interact with the web at scale. Power AI agents with clean web data. Firecrawl delivers the entire internet to AI agents and builders. (992 words) - [Introducing Agent](/agent/index.html): Firecrawl /agent is a magic API that searches, navigates, and gathers data from even the most complex websites. Describe what data you want and agent handles the rest. Find information in hard-to-reach places, return single datapoints or entire datasets at scale. (884 words) - [Free web extraction tools by Firecrawl](/tools/index.html): Free online tools for developers and marketers: extract URLs from any website, summarize articles with AI, and more. No signup required. Powered by Firecrawl. (909 words) - [blog/modern-rag-tech-stack/index.html](/blog/modern-rag-tech-stack/index.html) (1 words) - [blog/agent-skills/index.html](/blog/agent-skills/index.html) (1 words) - [blog/deploy-web-scrapers/index.html](/blog/deploy-web-scrapers/index.html) (1 words) - [blog/fine-tuning-deepseek/index.html](/blog/fine-tuning-deepseek/index.html) (1 words) - [blog/15-python-projects-2025/index.html](/blog/15-python-projects-2025/index.html) (1 words) - [blog/google-adk-multi-agent-tutorial/index.html](/blog/google-adk-multi-agent-tutorial/index.html) (1 words) - [blog/best-web-search-mcp/index.html](/blog/best-web-search-mcp/index.html) (1 words) - [blog/category/use-cases-and-examples/index.html](/blog/category/use-cases-and-examples/index.html) (1 words) - [blog/langflow-tutorial-visual-ai-workflows/index.html](/blog/langflow-tutorial-visual-ai-workflows/index.html) (1 words) - [blog/ai-resume-parser-job-matcher-python/index.html](/blog/ai-resume-parser-job-matcher-python/index.html) (1 words) - [blog/claude-managed-agents/index.html](/blog/claude-managed-agents/index.html) (1 words) - [What is browser fingerprinting evasion in web scraping?](/glossary/web-scraping-apis/what-is-browser-fingerprinting-evasion-web-scraping.html): Browser fingerprinting describes how websites identify browsers through their characteristics, and how scraping tools manage browser configurations for reliable data collection. (987 words) - [Best PDF Parsers for AI and RAG Workflows in 2026](/blog/best-pdf-parsers/index.html): A hands-on comparison of the best PDF parsers for AI and RAG pipelines in 2026, covering speed, output quality, table handling, and LLM-readiness for each tool. (3,837 words) - [Best Chunking Strategies for RAG (and LLMs) in 2026](/blog/best-chunking-strategies-rag/index.html): Compare seven chunking strategies for RAG systems using real benchmark data from NVIDIA and Chroma. Learn when to use recursive splitting, semantic chunking, page-level chunking, late chunking, and LLM-based approaches with practical code examples and honest trade-offs. (8,277 words) - [Introducing /parse](/blog/how-to-create-an-llms-txt-file-for-any-website/index.html): Learn how to generate an llms.txt file for any website using the llms.txt Generator and Firecrawl. (1,803 words) - [What is redirect handling in crawling?](/glossary/web-crawling-apis/what-is-redirect-handling-in-crawling/index.html): Redirect handling determines how web crawlers follow HTTP redirects like 301 and 302 to reach content at different URLs. (1,056 words) - [Ready to Build?](/blog/launch-week-iii-day-7-integrations/index.html): Firecrawl now connects with over 20 platforms including Discord, Make, Langflow, and more. Discover what's new on Integration Day. (630 words) - [Browser Sandbox: Secure Environments for Agents to Interact with the Web](/blog/introducing-browser-sandbox/index.html): Firecrawl Browser Sandbox gives AI agents a fully managed, isolated browser environment - zero config, pre-loaded with tools, and works alongside Firecrawl's scrape and search endpoints. (1,293 words) - [AI Agent Sandbox: How to Safely Run Autonomous Agents in 2026](/blog/ai-agent-sandbox/index.html): An AI agent sandbox is an isolated execution environment where an agent can take actions without those actions affecting the host system. Learn how it works. (3,929 words) - [How do I create a fact-checking agent skill?](/glossary/web-search-apis/fact-checking-agent-skill/index.html): A fact-checking agent skill is a callable tool registered with an AI agent that takes a claim, searches the web for evidence, and returns a verdict with source citations — built by combining a search API with an LLM reasoning step. (626 words) - [Top 7 AI-Powered Web Scraping Solutions in 2026](/blog/ai-powered-web-scraping-solutions/index.html): Discover the most advanced AI web scraping tools that are revolutionizing data extraction with natural language processing and machine learning capabilities. (2,922 words) - [What is a web scraping API?](/glossary/web-extraction-apis/what-is-web-scraping-api/index.html): A web scraping API handles the technical complexity of web scraping so developers can extract data with simple API calls instead of managing proxies, browsers, and anti-bot systems. (995 words) - [What is structured data vs unstructured data when extracting web data?](/glossary/web-extraction-apis/what-is-structured-vs-unstructured-data-web-extraction.html): Structured data comes organized in predefined formats like JSON or CSV, while unstructured data lacks organization and requires parsing before use. (1,102 words) - [Launch Week I / Day 5: Real-Time Crawling with WebSockets](/blog/launch-week-i-day-5-real-time-crawling-websockets/index.html): Our new WebSocket-based method for real-time data extraction and monitoring. (611 words) - [What is an index in the context of a web search API?](/glossary/web-search-apis/what-is-index-web-search-api/index.html): A search index is a structured database that stores organized, searchable content from websites, enabling web search APIs to return relevant results in milliseconds. (928 words) - [How do you get all links from a webpage?](/glossary/web-scraping-apis/get-all-links-from-webpage/index.html): Getting all links from a webpage means collecting every outbound URL the page contains, either by parsing raw HTML for anchor tags or by running a browser render first to capture links injected by JavaScript. (557 words) - [Introducing /parse](/playground/sitemap-xml.html): The API to search, scrape, and interact with the web at scale. Power AI agents with clean web data. Firecrawl delivers the entire internet to AI agents and builders. (142 words) - [What is OCR (optical character recognition) in web scraping?](/glossary/web-extraction-apis/what-is-ocr-in-web-scraping/index.html): OCR converts text trapped inside images into machine-readable data that scrapers can extract when traditional HTML parsing cannot access visual content. (928 words) - [What's the best tool for extracting content from pages that frequently redesign?](/glossary/web-extraction-apis/best-tool-extracting-content-redesigning-pages/index.html): Use LLM-powered extraction APIs like Firecrawl that understand content semantically rather than relying on CSS selectors that break when page layouts change. (492 words) - [How do you convert PDFs to RAG-ready data?](/glossary/web-extraction-apis/pdf-to-rag-ready-data/index.html): Converting PDFs to RAG-ready data means extracting text into clean, structured chunks that a vector store can index. The extraction step must handle scanned pages, preserve document structure, and produce consistent output regardless of PDF formatting. (557 words) - [What is natural language browser automation?](/glossary/web-scraping-apis/natural-language-browser-automation/index.html): Natural language browser automation lets you control a browser with plain English prompts instead of code, describing what to do and letting an AI agent handle the clicks, typing, and navigation. (543 words) - [How do I crawl an entire website and get content for every page?](/glossary/web-crawling-apis/crawl-entire-website-get-content/index.html): Point a crawl API at a root URL and it discovers every page automatically, returning the content of each as clean markdown. (417 words) - [Build an agent that checks for website contradictions](/blog/contradiction-agent/index.html): Using Firecrawl and Claude to scrape your website's data and look for contradictions. (889 words) - [What is a remote browser for web scraping?](/glossary/web-scraping-apis/remote-browser-web-scraping/index.html): A remote browser is a cloud-hosted browser instance that runs scraping tasks on a remote server, eliminating the need to manage local browser infrastructure. (428 words) - [What is a deep research API?](/glossary/web-search-apis/deep-research-api/index.html): A deep research API automates multi-step research by issuing queries, reading sources, and synthesizing findings with citations, producing comprehensive reports without requiring orchestration code from the caller. (536 words) - [Firecrawl - Search, Scrape, and Interact with the Web for AI](/signin/index.html): The API to search, scrape, and interact with the web at scale. Power AI agents with clean web data. Firecrawl delivers the entire internet to AI agents and builders. (94 words) - [Become a Firestarter](/firestarters/index.html): Join the Firestarters — an exclusive community for Firecrawl power users. Free Standard Plan, early access to new features, and a direct line to the engineering team. (407 words) - [What is the Chrome DevTools Protocol (CDP) in web scraping?](/glossary/web-scraping-apis/chrome-devtools-protocol-web-scraping/index.html): The Chrome DevTools Protocol (CDP) is a set of APIs that lets external tools control a Chromium browser over a WebSocket connection, enabling network interception, JavaScript execution, and screenshot capture for web scraping. (526 words) - [Playground](/playground/index.html): The API to search, scrape, and interact with the web at scale. Power AI agents with clean web data. Firecrawl delivers the entire internet to AI agents and builders. (258 words) - [blog/pdf-rag-system-langflow-firecrawl/index.html](/blog/pdf-rag-system-langflow-firecrawl/index.html) (1 words) - [Firecrawl Editor Theme: Launch Week III - Day 0](/blog/launch-week-iii-day-0-firecrawl-editor-theme/index.html): Our official Firecrawl Editor Theme provides a clean, focused coding experience optimized for everyone. (617 words) - [Launch Week I / Day 1: Introducing Teams](/blog/launch-week-i-day-1-introducing-teams/index.html): Our new Teams feature, enabling seamless collaboration on web scraping projects. (603 words) ## Resources - [Full Page Index](/index.html): Browse all cached pages with rich metadata - [About This Cache](/content/about.html): Methodology, technical details, and usage guidelines - [XML Sitemap](/content/sitemap.xml): Machine-readable sitemap for crawler discovery - [Robots.txt](/content/robots.txt): Crawler directives