How to Crawl an Entire Website with a Web Crawler API
Learn how to crawl a whole site with sitemap and link discovery, scope controls, async jobs, pagination, audits, retries, and practical API code.
Direct answer: Start with a root URL, define the host, paths, subdomains, depth, and page limit you will allow, then let the crawler discover URLs from the sitemap and internal links. Submit the crawl as an asynchronous job, poll its status, follow every paginated result, and audit returned, skipped, duplicate, and failed URLs against your intended inventory. “Entire website” is a scope you configure, not a guarantee that every URL will be found.
This guide uses the documented Firecrawl v2 workflow as a concrete example. Its API is POST https://api.firecrawl.dev/v2/crawl; the service returns a job ID and you retrieve status and results afterward. Check the Firecrawl documentation before publishing or running production code because defaults, limits, and pricing can change.
1. Define what “entire website” means
Write the boundary before writing code. Decide whether you mean:
- One path, such as
https://example.com/docs/. - The complete host, including every path on
www.example.com. - The registrable domain and its subdomains, such as
docs.example.comandblog.example.com. - Pages linked from the site only, or also URLs listed in XML sitemaps.
- HTML content only, or rendered JavaScript output, screenshots, links, images, and metadata.
Set an explicit maximum page count and discovery depth even when you expect a small site. Firecrawl documents a default crawl limit of 10,000 pages when limit is omitted. A limit protects you from calendars, faceted navigation, session URLs, and accidentally including another host.
2. Understand URL discovery
Most hosted crawlers combine two discovery routes:
| Route | Finds | Can miss |
|---|---|---|
| Sitemap | URLs the publisher deliberately listed in XML sitemap files | Pages omitted from the sitemap, stale entries, or blocked sitemap files |
| Internal links | Pages reachable by following links in fetched pages | Orphan pages, JavaScript-only links, and pages requiring a form or login |
Firecrawl supports sitemap modes include, skip, and only. Use only for a sitemap inventory, skip when link discovery is your source of truth, and include when you want both. A combined crawl is usually the best starting point, followed by an audit.
Path filters are evaluated against the starting URL too. If your include pattern excludes the root URL, the job can return zero pages. Test the root against every allow rule before launching a large crawl.
3. Configure scope and safety controls
Important Firecrawl crawl options include:
| Option | Purpose | Practical guidance |
|---|---|---|
limit |
Maximum pages to crawl | Set a number based on your inventory, with headroom for retries and redirects. |
maxDiscoveryDepth |
Maximum link depth from the seed | Use a low value for a section; raise it for a full site. |
crawlEntireDomain |
Allow crawling beyond the starting path | Its documented default is false; enable deliberately. |
allowSubdomains |
Permit subdomains | Enable only when those hosts belong in your corpus. |
allowExternalLinks |
Follow links to other domains | Usually leave false for a site crawl. |
include/skip/only |
Path and URL filters | Keep include rules narrow and verify the seed matches. |
delay |
Delay between requests | Use it for a polite crawl or a fragile origin; Firecrawl says setting it forces concurrency to one. |
| similar-URL deduplication | Merge URLs judged equivalent | Keep the default unless near-duplicate URLs are expected. |
| ignore query parameters | Treat query variants as one URL | Enable only if query strings do not change content; otherwise distinct pages can be merged. |
Respect the target’s access rules. Firecrawl documents reading robots.txt rules for FirecrawlAgent and *. Apify’s Website Crawler listing also documents robots.txt compliance by default. Confirm current provider behavior and the site owner’s policy before running a production crawl.
4. Choose the representation you need
For search or documentation pipelines, Markdown is convenient. Structured JSON is better when you need stable fields such as title, canonical URL, headings, and links. HTML preserves markup. Some hosted products also return screenshots, images, and metadata. Firecrawl’s crawl API allows scrape options per page, so choose the smallest representation that satisfies your downstream job.
JavaScript-heavy sites require browser rendering. HTTP-only fetching may return an application shell without the content a user sees. Rendering improves coverage of client-side pages but costs more time and resources. Treat rendering as a provider setting to verify, not a universal crawler guarantee.
5. Submit a crawl job with cURL
The following request starts a bounded crawl. Replace the token, URL, and options with values appropriate for your site. Consult the current Firecrawl schema for authentication and scrape-option details.
curl -X POST "https://api.firecrawl.dev/v2/crawl" \
-H "Authorization: Bearer $FIRECRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com",
"limit": 1000,
"maxDiscoveryDepth": 10,
"crawlEntireDomain": true,
"allowSubdomains": false,
"allowExternalLinks": false,
"sitemap": "include",
"scrapeOptions": {
"formats": ["markdown", "links", "metadata"]
}
}'
Save the returned job ID. Do not assume the POST response contains pages; the documented workflow is asynchronous.
6. Poll status and consume every result page
Python
import os
import time
import requests
API = "https://api.firecrawl.dev/v2"
HEADERS = {
"Authorization": f"Bearer {os.environ['FIRECRAWL_API_KEY']}",
"Content-Type": "application/json",
}
start = requests.post(
f"{API}/crawl",
headers=HEADERS,
json={
"url": "https://example.com",
"limit": 1000,
"maxDiscoveryDepth": 10,
"crawlEntireDomain": True,
"allowSubdomains": False,
"allowExternalLinks": False,
"sitemap": "include",
"scrapeOptions": {"formats": ["markdown", "links", "metadata"]},
},
timeout=30,
)
start.raise_for_status()
job_id = start.json()["id"]
next_url = f"{API}/crawl/{job_id}"
all_pages = []
while next_url:
response = requests.get(next_url, headers=HEADERS, timeout=60)
response.raise_for_status()
payload = response.json()
status = payload.get("status")
if status in {"failed", "cancelled"}:
raise RuntimeError(payload)
all_pages.extend(payload.get("data", []))
next_url = payload.get("next")
if status not in {"completed", "failed", "cancelled"} and not payload.get("next"):
time.sleep(5)
next_url = f"{API}/crawl/{job_id}"
print(f"received {len(all_pages)} pages")
for page in all_pages:
print(page.get("metadata", {}).get("sourceURL"))
Node.js
const api = 'https://api.firecrawl.dev/v2';
const headers = {
'Authorization': `Bearer ${process.env.FIRECRAWL_API_KEY}`,
'Content-Type': 'application/json'
};
const created = await fetch(`${api}/crawl`, {
method: 'POST',
headers,
body: JSON.stringify({
url: 'https://example.com',
limit: 1000,
maxDiscoveryDepth: 10,
crawlEntireDomain: true,
allowSubdomains: false,
allowExternalLinks: false,
sitemap: 'include',
scrapeOptions: { formats: ['markdown', 'links', 'metadata'] }
})
});
if (!created.ok) throw new Error(await created.text());
const { id } = await created.json();
let url = `${api}/crawl/${id}`;
const pages = [];
while (url) {
const response = await fetch(url, { headers });
if (!response.ok) throw new Error(await response.text());
const payload = await response.json();
if (['failed', 'cancelled'].includes(payload.status)) throw new Error(JSON.stringify(payload));
pages.push(...(payload.data || []));
if (payload.next) {
url = payload.next;
} else if (payload.status === 'completed') {
url = null;
} else {
await new Promise(resolve => setTimeout(resolve, 5000));
url = `${api}/crawl/${id}`;
}
}
console.log(`received ${pages.length} pages`);
7. Handle pagination and the 10 MB response boundary
Firecrawl documents a next URL when a crawl is still running or when the content exceeds 10 MB. Always follow that URL until it is absent and the job is complete. Store pages incrementally instead of holding a very large response in memory. De-duplicate by canonical URL or the provider’s source URL field, while preserving redirects and error records for your audit.
8. Audit whether the crawl covered the intended site
- Export the URLs returned by the crawler.
- Fetch the site’s sitemap index and child sitemaps independently.
- Normalize URLs consistently: scheme, host casing, fragments, trailing slash policy, and approved query parameters.
- Compare sitemap URLs, discovered URLs, returned URLs, skipped URLs, and errors.
- Inspect a sample from each path and depth boundary, including pages rendered by JavaScript.
- Run a second, deliberately scoped crawl for missing sections instead of silently calling the first run exhaustive.
A missing URL can mean it was absent from the sitemap, unreachable by links, blocked by robots rules, outside your filters, duplicate under the provider’s rules, or failed during rendering. Keep those categories separate in your report.
9. Reliability, performance, and cost
- Reliability: Treat jobs as retryable. Persist the job ID, poll with backoff, record provider errors, and make downstream writes idempotent by URL and crawl version.
- Performance: Narrow path filters and a realistic depth reduce work. Browser rendering is slower than raw HTTP. A delay can make the crawl polite but reduces throughput; Firecrawl documents that delay forces concurrency to one.
- Completeness: No provider can infer orphan pages that are neither in a sitemap nor reachable through discovered links. Authentication, paywalls, CAPTCHAs, and robots rules also limit coverage.
- Cost: Firecrawl’s product page states a price of one credit per page crawled. Its displayed plans and prices can change, so verify them before budgeting. Large limits, rendered pages, retries, and repeated recrawls all increase usage.
- Freshness: Save crawl timestamps and source URLs. Recrawl only changed sections when your application can identify them.
10. Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Zero pages | The root URL fails an include rule, or the sitemap is empty | Test the seed against filters; temporarily use sitemap include and inspect discovery logs. |
| Only the home page | Depth is too low, links are client-rendered, or the site has no reachable navigation | Raise maxDiscoveryDepth, enable browser rendering where supported, and add sitemap discovery. |
| Subdomains missing | allowSubdomains is disabled |
Enable it only for approved hosts and include those hosts in your audit. |
| External pages appear | External-link crawling is enabled | Set allowExternalLinks to false and add host filters. |
| Important query variants disappear | Query parameters were ignored or similar-URL deduplication merged them | Preserve meaningful parameters and disable or adjust deduplication for that site. |
| Job appears incomplete | Results are paginated or the job is still running | Follow next until completion; do not stop after the first response. |
| HTML contains no article text | The page requires JavaScript rendering | Use the provider’s browser-rendering option and compare the rendered output with a real browser. |
| Many denied or failed pages | Robots rules, authentication, rate limits, bot checks, or origin errors | Review the site policy and provider error details; slow the crawl, supply authorized access when permitted, or record the pages as unavailable. |
| Unexpectedly high usage | Limit too high, duplicate URLs, retries, or broad subdomain scope | Set a lower limit, tighten filters, inspect normalized URLs, and avoid unnecessary recrawls. |
11. When to use Crawl, Scrape, or Map
Use a crawl when you need many pages and link or sitemap discovery. Use a single-page scrape when you already know the URL and need its current content. Use a map-style operation when you need a URL inventory before deciding which pages to fetch. Firecrawl’s product FAQ frames these as separate workflows; choose based on whether discovery or extraction is the immediate task.
12. Or skip the browser setup
If your workflow only needs images or PDFs of selected pages discovered by the crawler, ScreenshotNeo provides a single screenshot API request. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. An MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
13. FAQ
Can an API really crawl every URL?
No. Coverage is limited to discoverable, permitted, in-scope URLs and the configured page and depth limits. Verify the result against a sitemap or URL inventory.
Should I use sitemap-only mode?
Use it when the sitemap is your authoritative inventory. It will miss pages absent from that sitemap, so combine it with link discovery when you need broader coverage.
How do I process pages while a crawl runs?
Poll the asynchronous job, consume each response page as it becomes available, and write records idempotently. Follow the provider’s next URL until the job finishes.
What output should I store?
Store the source URL, crawl timestamp, status, error details, and the representation needed by your application. Markdown is useful for search and documentation; structured fields are better for indexing and validation.
How often should I recrawl?
Base the schedule on how often the site changes and the cost of stale data. Keep crawl versions so you can identify additions, removals, and changed content.
Sources
- Firecrawl Advanced Scraping Guide — crawl configuration, discovery, pagination, deduplication, and robots.txt behavior.
- Firecrawl Web Crawling API — output formats, FAQ, and provider pricing claims.
- Apify Website Crawler — one hosted crawler’s page, depth, rendering, and robots.txt options.


