ScreenshotNeo

BlogHow-to

How to Build a Web Crawler with Headless Chrome

Build a JavaScript crawler with Puppeteer, robots.txt checks, a bounded queue, and reliable extraction from rendered pages.

By the ScreenshotNeo team4 October 202610 min read

To crawl JavaScript-rendered pages with headless Chrome, create a bounded URL queue, check each URL against your scope and the site’s robots.txt rules, then use Puppeteer to load pages, wait for the content you need, and extract the rendered DOM. Use a regular HTTP client when the response already contains the content; browser rendering costs more resources and is needed only when JavaScript or browser interaction changes the content or links you want.

Headless Chrome runs without a visible UI. Current Chrome Headless uses the same browser implementation as headful Chrome. Since Chrome 132.0.6793.0, the older Headless implementation is available separately as chrome-headless-shell. [Chrome Headless documentation]

1. Decide whether the page needs a browser

Inspect a representative page before building a browser crawler. Fetch its HTML with an HTTP client and check whether the text and links you need are already present. If they are, ordinary HTTP retrieval is simpler. If the application has a prerendering option, consider enabling that. Use Chrome when you need JavaScript execution, browser interaction, or the resulting rendered DOM. [Chrome Headless article]

Puppeteer is a JavaScript library for controlling Chrome or Firefox. Playwright is another option, with configurable Chromium and headless modes. Choose based on your runtime, browser management, interaction needs, deployment footprint, and browser coverage. Their available documentation does not establish a universal crawler throughput or memory winner, so benchmark your own pages if those are deciding factors. [Puppeteer getting started] [Playwright browser documentation]

2. Set a responsible crawl scope

  1. Write down the purpose of the crawl and the hosts you are authorized to access. Exclude authenticated or private content unless you have authorization.
  2. Allow only http: and https: URLs and reject hosts outside your configured allowlist before opening a browser.
  3. Fetch the host’s top-level /robots.txt, identify the crawler with a descriptive user agent, and apply the matching parseable rules.
  4. Check the site’s terms and applicable rules. Robots.txt is crawler guidance, not access authorization.

RFC 9309 says successful robots.txt retrievals require following parseable rules and recommends following at least five consecutive redirects. A 4xx unavailable response may permit crawling under the protocol; a 5xx response or network-unreachable file requires assuming complete disallow. Cache robots.txt for no more than 24 hours unless it is unreachable. These protocol rules do not grant permission to access a site. [RFC 9309]

Robots.txt also does not secure information. Google notes that a disallowed URL can still appear in search results if linked elsewhere. Use access controls to protect private material, and an appropriate indexing control when the goal is to prevent search-result inclusion. [Google Search Central: robots.txt]

3. Create the Puppeteer project

Install Puppeteer and its compatible browser in the project. Pin the dependency version in your lockfile and make browser installation part of the deployment or CI setup. Puppeteer’s installation script downloads a compatible browser; if package installation scripts are disabled, the browser may be missing. [Puppeteer installation]

mkdir headless-crawler
cd headless-crawler
npm init -y
npm install puppeteer

Save the following as crawler.mjs. It takes seed URLs as command-line arguments, restricts crawling to their exact hostnames, checks robots.txt with the robots-parser package, limits simultaneous pages, and extracts page titles, text, and links. Its robots helper is intentionally simple: it follows redirects through the HTTP client, and treats robots fetch errors as disallow. For a production crawler, review your chosen robots library’s parsing and redirect behavior against RFC 9309, including redirect chains and caching.

npm install robots-parser
import puppeteer from 'puppeteer';
import robotsParser from 'robots-parser';

const seeds = process.argv.slice(2);
if (seeds.length === 0) {
  console.error('Usage: node crawler.mjs https://example.com/');
  process.exit(2);
}

const allowedHosts = new Set(seeds.map((raw) => new URL(raw).hostname));
const userAgent = 'ExampleResearchCrawler/1.0 (+https://example.com/crawler-info)';
const maxPages = Number(process.env.MAX_PAGES ?? 30);
const concurrency = Number(process.env.CONCURRENCY ?? 2);
const timeoutMs = Number(process.env.TIMEOUT_MS ?? 30000);
const perHostDelayMs = Number(process.env.PER_HOST_DELAY_MS ?? 1000);
const queue = [...new Set(seeds.map((raw) => new URL(raw).href))];
const queued = new Set(queue);
const visited = new Set();
const robotsCache = new Map();
const lastStartedByHost = new Map();

function inScope(raw) {
  try {
    const u = new URL(raw);
    return ['http:', 'https:'].includes(u.protocol) && allowedHosts.has(u.hostname);
  } catch {
    return false;
  }
}

async function canFetch(url) {
  const origin = new URL(url).origin;
  if (!robotsCache.has(origin)) {
    const robotsUrl = `${origin}/robots.txt`;
    try {
      const response = await fetch(robotsUrl, {
        headers: { 'User-Agent': userAgent },
        signal: AbortSignal.timeout(timeoutMs),
        redirect: 'follow'
      });
      if (response.status >= 500) {
        robotsCache.set(origin, { parser: null, disallowAll: true });
      } else if (response.status >= 400) {
        robotsCache.set(origin, { parser: null, disallowAll: false });
      } else {
        robotsCache.set(origin, {
          parser: robotsParser(response.url || robotsUrl, await response.text()),
          disallowAll: false
        });
      }
    } catch (error) {
      console.error(`robots.txt unavailable for ${origin}: ${error.message}`);
      robotsCache.set(origin, { parser: null, disallowAll: true });
    }
  }
  const rules = robotsCache.get(origin);
  return !rules.disallowAll && (!rules.parser || rules.parser.isAllowed(url, userAgent) !== false);
}

async function pace(url) {
  const host = new URL(url).host;
  const previous = lastStartedByHost.get(host) ?? 0;
  const waitMs = Math.max(0, perHostDelayMs - (Date.now() - previous));
  if (waitMs) await new Promise((resolve) => setTimeout(resolve, waitMs));
  lastStartedByHost.set(host, Date.now());
}

const browser = await puppeteer.launch({ headless: true });
let active = 0;
let stopped = false;
let processed = 0;
let resolveWorkers;
const workersDone = new Promise((resolve) => { resolveWorkers = resolve; });

async function worker() {
  while (!stopped) {
    if (processed >= maxPages) {
      stopped = true;
      break;
    }
    const url = queue.shift();
    if (!url) {
      if (active === 0) break;
      await new Promise((resolve) => setTimeout(resolve, 50));
      continue;
    }
    if (visited.has(url)) continue;
    visited.add(url);
    active++;
    processed++;
    try {
      if (!inScope(url) || !(await canFetch(url))) {
        console.log(JSON.stringify({ url, outcome: 'skipped_by_scope_or_robots' }));
        continue;
      }
      await pace(url);
      const page = await browser.newPage();
      try {
        page.setDefaultNavigationTimeout(timeoutMs);
        await page.setUserAgent(userAgent);
        const response = await page.goto(url, { waitUntil: 'domcontentloaded', timeout: timeoutMs });
        // Replace this with a selector meaningful to the target site when possible.
        // A bounded delay is a fallback, not proof that every app has finished rendering.
        await page.waitForTimeout(500);
        const result = await page.evaluate(() => ({
          title: document.title,
          text: document.body?.innerText ?? '',
          links: [...document.querySelectorAll('a[href]')].map((a) => a.href)
        }));
        console.log(JSON.stringify({
          requestedUrl: url,
          finalUrl: page.url(),
          fetchedAt: new Date().toISOString(),
          status: response?.status() ?? null,
          title: result.title,
          text: result.text,
          linkCount: result.links.length,
          outcome: 'extracted'
        }));
        for (const link of result.links) {
          if (inScope(link)) {
            const normalized = new URL(link).href;
            if (!visited.has(normalized) && !queued.has(normalized)) {
              queue.push(normalized);
              queued.add(normalized);
            }
          }
        }
      } catch (error) {
        console.error(JSON.stringify({ url, outcome: 'navigation_or_extraction_error', error: error.message }));
      } finally {
        await page.close();
      }
    } catch (error) {
      console.error(JSON.stringify({ url, outcome: 'crawler_error', error: error.message }));
    } finally {
      active--;
      if (queue.length === 0 && active === 0) resolveWorkers();
    }
  }
  if (queue.length === 0 && active === 0) resolveWorkers();
}

try {
  await Promise.all(Array.from({ length: Math.max(1, concurrency) }, () => worker()));
  await workersDone;
} finally {
  await browser.close();
}

Run it with one or more seed URLs:

node crawler.mjs https://example.com/

Useful environment settings are MAX_PAGES to cap pages per run, CONCURRENCY to bound simultaneous workers, TIMEOUT_MS to bound navigation and robots requests, and PER_HOST_DELAY_MS to set a minimum interval between starts on each host. They are conservative example controls, not universally safe values; tune them to site policy and observed server behavior.

4. Wait for the right page state

The example waits for domcontentloaded and then a short bounded delay. For a real target, prefer a condition tied to the content you need, such as a results container or a known article selector:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: timeoutMs });
await page.waitForSelector('main article', { timeout: 10000 });

Choose a wait condition based on the application. A fixed delay can waste time on fast pages and capture incomplete pages on slow ones. Waiting for every network connection to become idle can also be unsuitable for pages with analytics, long polling, or streaming. Use bounded timeouts and record when a readiness condition fails. Chrome’s crawler example demonstrates navigation and reading serialized page content, but its older networkidle0 example should not be treated as a universal rule. [Chrome Headless article]

5. Extract, normalize, and persist results

The example extracts title, visible body text, and resolved anchor URLs. In a production pipeline, define the fields you need and keep extraction separate from crawling policy. Normalize URLs consistently before de-duplication: at minimum, resolve relative links, allow only supported schemes, remove fragments, and decide explicitly how to handle query parameters, trailing slashes, and duplicate tracking parameters. Do not discard query parameters blindly because they may identify distinct content.

Persist crawl state outside the browser process. Record the original URL, final URL after redirects, fetch time, response status when available, extraction outcome, and an error summary. For large crawls, persist the frontier and visited set in a database or durable queue so a restart does not lose progress. Track queue depth, successes, errors, render time, and duplicate rate to understand coverage and resource use.

6. Reduce browser work safely

  • Keep a bounded number of browser workers and pages. Do not launch an unbounded process per URL.
  • Close each page in a finally block and close the browser when the crawl ends.
  • Use a host-specific delay, bounded navigation timeout, and capped retries with backoff. Stop retrying persistent failures.
  • Consider request interception to skip unnecessary images, fonts, or media only after checking that the page still renders the required text and links. Blocking scripts, styles, or XHR/fetch can break content extraction.
  • Use HTTP retrieval for content already present in the response; reserve browser rendering for pages that need it.

Puppeteer supports request interception. Resource filtering is a tradeoff: it may reduce work, but sites can depend on resources you block. Compare extracted output before and after changing filters. [Chrome Headless article]

7. Troubleshooting

Symptom Likely cause What to do
Browser executable not found Browser download did not run, or the expected browser is not installed in the deployment image. Check Puppeteer’s installation output and deployment setup. Install the compatible browser explicitly and pin the Puppeteer version. [Puppeteer installation]
Page has a title but no expected content The app has not rendered the target content yet, or extraction ran before the right state. Wait for a target-specific selector with a timeout. Check whether the content is inside a frame or requires an authorized interaction.
Navigation timeout Slow response, long-running page requests, or an unsuitable readiness condition. Use a bounded timeout and a readiness condition appropriate to the target. Log the final URL and failure stage; retry only transient failures with a cap.
Robots rules appear to block everything Robots fetch failed, returned a server error, or the parser did not handle the response as expected. Inspect the status and fetched file. This example fails closed on network errors and 5xx responses. Verify parser behavior and redirects against RFC 9309 before production use.
Many duplicate pages enter the queue URLs differ only by fragments, tracking parameters, or slash conventions. Apply a consistent URL normalization policy before inserting into the queue; preserve query parameters that may change page content.
Content disappears after blocking resources The site needs a blocked script, stylesheet, or API request to render the page. Restore the needed resource types and validate the output. Filter only resources shown to be unnecessary for your extraction.
Crawl consumes too much memory or browser capacity Too many pages are open or pages are not closed after errors. Lower concurrency, ensure page closure in finally, and keep only extracted data in the crawler process.

8. Performance, reliability, and cost

There is no universal safe request rate or published Puppeteer-versus-Playwright crawler benchmark established by the sources for this guide. Start with low concurrency, respect robots rules and site terms, and tune pacing based on the target’s policy and observed response behavior. A browser uses more resources than direct HTTP retrieval because it runs page code and manages browser state; avoid paying that cost for pages that do not need rendering.

For reliability, separate frontier storage, robots policy, browser navigation, extraction, and persistence. Retry only transient failures, use capped backoff, and keep a crawl record so failures can be resumed or audited. A timeout, CAPTCHA, or access restriction is a reason to stop or review authorization, not to evade the site’s controls.

Or skip the browser setup

If your task is to capture a page as an image or PDF rather than crawl its links, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF, with options including full-page capture, CSS selector capture, waits, custom CSS and JavaScript, headers, cookies, and viewport settings. See the ScreenshotNeo API docs.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
  • Cookie banners, newsletter popups, and chat widgets are removed before the shot.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify the page verdict and billing status.
  • An MCP server lets AI agents using Claude, Cursor, or another MCP client take screenshots, get page information, and capture PDFs.
  • 1,000 screenshots per month are free with no card. Paid plans start at $5 for 3,000 screenshots; every feature is on every plan.

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

FAQ

Can a headless crawler access pages that require a login?

Only crawl authenticated content when you have authorization and the site’s rules allow it. Keep credentials secure and avoid collecting private data outside the stated purpose.

Does robots.txt mean a page is private?

No. It is crawler guidance, not access control. Protect private pages with authentication or other access controls.

Should I choose Puppeteer or Playwright?

Use the library that fits your language, browser needs, deployment, and target interactions. Make the browser mode and binary explicit, then validate behavior on the pages you intend to crawl.

Can I use Chrome’s command line instead of Puppeteer?

Chrome supports headless operation from the command line. For a crawler with a queue, selectors, retries, and stored state, an automation library provides a more practical control surface. [Chrome Headless documentation]