ScreenshotNeo

BlogGuides

How to Scrape Websites in Stealth Mode with Puppeteer and Playwright

Build reliable, authorized Puppeteer and Playwright scrapers with coherent browser sessions, useful waits, clear diagnostics, and a safe response to access blocks.

By the ScreenshotNeo team4 October 202613 min read

Short answer: Use Puppeteer or Playwright with the smallest configuration that reliably renders the content you are authorized to access. Keep session settings consistent with the task, wait for the content you need, and log what happened. “Stealth” is not a guarantee that a site will accept automation. If the site presents a CAPTCHA, explicit denial, or repeated blocks, stop and use an official API, request permission, or find another permitted source.

This guide builds that workflow in both libraries. It covers setup, runnable extraction examples, session handling, waits, diagnostics, operational limits, and when a screenshot API is a better fit than writing browser automation.

1. Check access before automating

First look for an official API, export, feed, or documented integration. Review the site’s terms and its crawler instructions. The Robots Exclusion Protocol is standardized in RFC 9309; robots.txt provides crawler guidance. It is not access authorization and does not replace legal review.

Puppeteer’s project guidance says that the calling code is responsible for using its automation capabilities safely and as intended. That is project guidance, not a legal conclusion. See the Puppeteer security policy.

Use browser automation only for content and workflows you are allowed to access. Keep request rates modest, honor applicable limits, and avoid collecting personal or restricted information without authorization.

2. What “stealth mode” can and cannot do

Browserless’s vendor-authored stealth scraping guide describes detection signals that can include IP and network reputation, HTTP or TLS hints, browser fingerprint consistency, interaction timing, and challenges. That is one vendor’s description of possible signal categories, not a universal or complete model used identically by every site.

The practical engineering goal is coherence and fewer avoidable failures: use a browser session whose locale, timezone, viewport, user-agent-related settings, permissions, and stored state fit the workflow. Randomly changing many values can make a session inconsistent or break page behavior. A plugin or hosted route cannot guarantee acceptance or invisibility.

Browserless documents managed stealth routes and BrowserQL integrations for Puppeteer and Playwright. Those are vendor-described product options, not independent evidence that a route will work for a particular site. Its documentation also warns that stealth routes can have unexpected effects on automation. Check current availability and behavior in the Browserless documentation before designing around them.

3. Install and run a normal browser workflow first

Use a supported Node.js release and install the library in a new project directory. Puppeteer downloads a compatible browser during installation by default. Playwright requires installing its browser binaries after adding the package.

# Choose one library for this example
npm init -y
npm install puppeteer
# or:
npm install playwright
npx playwright install chromium

Start with an ordinary browser session. The examples below fetch a public demonstration page, wait for a meaningful selector, extract visible quote text, and close the browser even if navigation or extraction fails. Change the URL and selectors only for pages you are authorized to access.

4. Scrape with Puppeteer

Save this as scrape-puppeteer.mjs and run node scrape-puppeteer.mjs.

import puppeteer from 'puppeteer';

const url = 'https://quotes.toscrape.com/';
const browser = await puppeteer.launch({ headless: true });

try {
  const page = await browser.newPage();
  page.setDefaultNavigationTimeout(30_000);
  page.setDefaultTimeout(10_000);

  page.on('requestfailed', request => {
    console.error('Request failed:', request.url(), request.failure()?.errorText);
  });

  const response = await page.goto(url, { waitUntil: 'domcontentloaded' });
  if (!response) throw new Error('Navigation returned no main resource response');
  console.log('HTTP status:', response.status());
  if (!response.ok()) throw new Error(`Navigation returned HTTP ${response.status()}`);

  await page.waitForSelector('.quote .text', { visible: true });
  const quotes = await page.$$eval('.quote', nodes => nodes.map(node => ({
    text: node.querySelector('.text')?.textContent?.trim() ?? '',
    author: node.querySelector('.author')?.textContent?.trim() ?? ''
  })));
  console.log(JSON.stringify(quotes, null, 2));
} finally {
  await browser.close();
}

domcontentloaded avoids waiting for every image, ad, and long-lived connection. The selector wait then gates extraction on the content that matters. Use networkidle0 or networkidle2 only if network quiet is actually a useful readiness signal for the target; analytics, polling, and streaming can prevent idle indefinitely or make it misleading.

Puppeteer configuration choices

  • Headless mode: headless: true is appropriate for unattended jobs. Use visible mode during local diagnosis if a display is available.
  • Viewport: set page.setViewport({ width, height, deviceScaleFactor }) when layout depends on screen size. Choose dimensions that match the intended task.
  • Locale and timezone: set them deliberately with supported page or browser emulation APIs when the target legitimately varies by locale or time zone. Keep them aligned with the workflow rather than cycling values.
  • User agent: set it only when the task requires a particular browser profile. A user-agent string alone does not make the rest of the browser or network behave like that profile.
  • Cookies and authentication: load only state that belongs to the authorized account and task. Keep account sessions separated.
  • Permissions and geolocation: grant only permissions needed for the workflow. Avoid requesting or spoofing location without a task requirement.
  • Request interception: blocking images or fonts can save resources when extracting text, but can also alter layout or prevent application code from running. Measure the effect on the actual page.

5. Scrape with Playwright

Save this as scrape-playwright.mjs and run node scrape-playwright.mjs. A new browser context gives this run its own cookies and local and session storage. Playwright describes this isolation in its browser context documentation.

import { chromium } from 'playwright';

const url = 'https://quotes.toscrape.com/';
const browser = await chromium.launch({ headless: true });

try {
  const context = await browser.newContext({
    viewport: { width: 1365, height: 900 },
    locale: 'en-US',
    timezoneId: 'UTC'
  });
  const page = await context.newPage();
  page.setDefaultNavigationTimeout(30_000);
  page.setDefaultTimeout(10_000);

  page.on('requestfailed', request => {
    console.error('Request failed:', request.url(), request.failure()?.errorText);
  });
  page.on('response', response => {
    if (response.status() >= 400) {
      console.error('HTTP response:', response.status(), response.url());
    }
  });

  const response = await page.goto(url, { waitUntil: 'domcontentloaded' });
  if (!response) throw new Error('Navigation returned no main resource response');
  if (!response.ok()) throw new Error(`Navigation returned HTTP ${response.status()}`);

  const quote = page.locator('.quote');
  await quote.first().waitFor({ state: 'visible' });
  const quotes = await quote.evaluateAll(nodes => nodes.map(node => ({
    text: node.querySelector('.text')?.textContent?.trim() ?? '',
    author: node.querySelector('.author')?.textContent?.trim() ?? ''
  })));
  console.log(JSON.stringify(quotes, null, 2));
} finally {
  await browser.close();
}

Playwright configuration choices

  • Browser engine: Playwright offers Chromium, Firefox, and WebKit through its API. Choose the engine needed for the task and install its browser binary.
  • Context isolation: create a new context for a separate account, run, or identity. Contexts keep cookies and storage isolated. Do not share saved state across unrelated accounts or sites.
  • Persistent state: for an authorized workflow that needs a login across runs, save and restore storage state deliberately, restrict access to the state file, and expire it when no longer needed. Treat it like a credential.
  • CDP connections: chromium.connectOverCDP() attaches to an existing Chromium browser. Playwright documents this as Chromium-only and lower fidelity than its Playwright protocol connection. See the BrowserType API reference. Use CDP when integration with an existing Chromium instance requires it, not as a stealth setting.
  • Tracing: Playwright tracing can record actions, DOM snapshots, and network activity for debugging. Store trace files carefully because they may contain page content or session data. See Trace Viewer documentation.

6. Make sessions consistent and reproducible

  1. Use one context per independent session or account. In Playwright, contexts isolate browser storage. In Puppeteer, use separate browser contexts where supported by the installed version, or separate browser processes when stronger isolation is needed.
  2. Set viewport, locale, and timezone only when the workflow benefits from fixed values. Record those settings with each run so failures are reproducible.
  3. Keep cookies and storage for the intended task only. Do not casually reuse one authenticated profile across unrelated targets.
  4. Avoid contradictory identity settings across browser APIs and outbound requests. Setting a header does not change every browser-level or network-level signal.
  5. Do not add fingerprint modifications as the first response to a failed run. First determine whether the failure was navigation, a JavaScript error, missing content, expired session state, rate limiting, or a direct access denial.

These are practical recommendations, not a method for defeating a site’s controls. Browserless’s article recommends minimal changes and warns that excessive signal changes can increase risk or break layouts; treat that as vendor guidance rather than a guarantee.

7. Wait for the page state you need

Arbitrary sleeps are easy to write but often unreliable. A fixed delay can be too short on a slow run and waste time on a fast one. Prefer a readiness condition tied to the task:

  • Wait for a specific content selector to become visible.
  • Wait for a URL change or navigation after a permitted interaction.
  • Wait for a known application event when the page exposes one.
  • For an authorized page with a known API response, wait for that response and verify the content rendered as expected.
  • Use a bounded timeout and report what condition timed out.

Browserless’s BrowserQL documentation notes that empty extraction can occur when a page requires JavaScript and recommends waiting for a selector or event. That diagnosis applies broadly: a successful initial navigation does not mean a single-page application has finished rendering the data.

8. Diagnose failures before changing settings

Capture enough evidence to distinguish failure classes. Log the target URL, run ID, navigation duration, main document status, final URL, selector wait result, and failed requests. Avoid logging secrets, full cookies, authorization headers, or unnecessary personal data.

Signal Likely class Next step
Navigation timeout; document never responds Network, DNS, overloaded site, or timeout budget Check connectivity and the target’s availability. Increase the timeout only if the task’s legitimate load time requires it.
Navigation succeeds but expected selector is absent Render timing, changed markup, wrong route, or client-side error Inspect the final URL, page title, console errors, and DOM. Wait for the relevant content state and update the selector if the page changed.
HTTP 401 or 403, login redirect, or explicit access-denial page Authentication or authorization denial Confirm the account and access permission. Use the site’s supported API or contact the owner. Do not escalate evasion.
HTTP 429 or a rate-limit notice Request rate exceeded or policy limit reached Stop or reduce requests, use a permitted retry interval, and seek an approved rate limit or data feed.
CAPTCHA or challenge interstitial Site is asking to verify or restrict access Stop automated retries. Seek permission, an official access path, or a different source.
Only some resources fail Blocked, unavailable, or nonessential subresource Identify whether the failed resource is needed for the requested content. Do not treat a successful document status as proof the page is complete.
Works once, fails on later runs Expired session, changed page, transient network failure, or site policy change Compare run artifacts and session age. Retry transient failures sparingly; reassess repeated denial.

9. Puppeteer or Playwright?

There is no evidenced universal stealth winner. Choose based on the job’s browser support, existing code, session needs, debugging tools, and deployment model. The frameworks do not independently guarantee that a site will accept automation.

Constraint Practical choice
Existing Puppeteer scripts and Chromium-only task Keep Puppeteer if it meets the maintenance and observability needs.
Need Chromium, Firefox, and WebKit through one API Evaluate Playwright and confirm the exact browser support needed.
Need isolated cookies and storage for parallel accounts or runs Playwright contexts directly provide isolated browser state; design context lifetimes explicitly.
Need to attach to an existing Chromium over CDP Both ecosystems can work with Chromium integrations, but Playwright notes its CDP connection has lower fidelity than its native protocol connection.
Repeated problems at high concurrency Investigate browser lifecycle, memory, queueing, and per-domain limits. Switching libraries alone may not address infrastructure constraints.

10. Reliability, performance, and cost

Reliability

  • Use bounded navigation and selector timeouts, and close pages and browser processes in cleanup paths.
  • Retry only transient errors, with a small retry budget and backoff. Do not repeatedly retry 401, 403, 429, CAPTCHA, or explicit denial responses.
  • Set per-domain concurrency limits and stop conditions. A queue and a circuit breaker prevent one failing target from consuming the whole worker pool.
  • Keep a small canary run and compare status, final URL, and extracted-field presence before increasing batch size.
  • Pin and update library and browser versions deliberately. A browser update can change rendering or automation behavior.

Performance

  • Do not wait for every resource if the task only needs the initial document and a known content selector.
  • Reuse a browser process for a bounded batch where appropriate, while isolating work with separate contexts. Monitor memory and recycle workers when needed.
  • Block images, fonts, or other resource types only after confirming they are not necessary for layout or content.
  • Limit per-host concurrency and add pacing appropriate to the target’s terms and published limits.
  • Measure end-to-end time, browser startup time, page wait time, failures, and worker memory on your workload. There is no benchmark here that predicts another site or deployment.

Cost

Self-hosting costs include compute, memory, browser downloads, storage for artifacts, network egress, engineering time, and the operational work of keeping workers healthy. Hosted browser infrastructure trades some operational control for a provider’s integration and management features. Compare actual workload volume, concurrency, session requirements, data handling, and provider-specific limits before choosing; the available sources do not establish a neutral performance or price comparison.

11. When a managed browser service may help

A managed browser can be useful when browser installation, remote execution, or worker lifecycle is a larger burden than the scraping logic. Browserless documents managed routes and BrowserQL integrations for Puppeteer and Playwright, but verify current feature availability and limits in its docs. Compare operational control, debugging, session handling, data processing, and cost for your workload.

If the task is simply to capture a page as an image or PDF, full browser-control code may be unnecessary. ScreenshotNeo is a website screenshot API and MCP server for developers. It returns a screenshot or PDF from one GET request, and its options include full-page capture, selector capture, waits, custom CSS and JavaScript, caching, bulk capture, and async jobs. See ScreenshotNeo and its API documentation.

12. Or skip the browser setup

For a screenshot or PDF task, call ScreenshotNeo directly. The following cURL, Python, and Node.js examples use the documented endpoint and a placeholder key. See the ScreenshotNeo docs for supported options and response behavior.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));

In Node.js environments without Bun, write the response bytes with the built-in node:fs/promises module:

import { writeFile } from 'node:fs/promises';

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo accepts cookie and consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response reports the page verdict and billing state in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card.

13. Troubleshooting checklist

  • Browser binary missing: install the browser required by your library; for Playwright, run npx playwright install chromium (or the engine you use).
  • Navigation returns no response: inspect network errors and redirects; check whether the page closed or changed protocol before the main response.
  • Selector wait times out: inspect the final URL and rendered DOM; verify the selector and wait for the actual content state, not just navigation.
  • Text is empty: the content may load client-side, be inside a frame, or use a different selector. Inspect the DOM and console errors; do not assume a 200 response means data is present.
  • Script hangs on network idle: switch to a relevant selector, event, or bounded delay if the page keeps polling or streaming.
  • 401, 403, 429, CAPTCHA, or block page: verify permission and account access, reduce or stop requests, and use an official or approved route. Do not respond by endlessly varying fingerprints.
  • Context state unexpectedly disappears: a fresh context is isolated by design. Persist only the authorized storage state needed for continuity, protect it as a credential, and verify its expiry.
  • Runs become slow or workers crash: reduce concurrency, close contexts and pages, inspect memory and browser process counts, and use bounded worker lifetimes.
  • Different layout between runs: record viewport, locale, timezone, browser version, final URL, and session state to find which task-relevant input changed.

14. Frequently asked questions

Can Puppeteer or Playwright be made undetectable?

No. Neither library nor a stealth plugin guarantees acceptance. Sites can use different controls, and those controls can change.

Does Playwright avoid detection better than Puppeteer?

The supplied evidence does not establish a universal winner. Choose based on browser engines, existing stack, context and session needs, debugging, and deployment.

Should I rotate user agents or other browser values?

Only configure values for a real task requirement and keep them consistent. Random rotation is not a reliable fix for a denial and can create contradictory signals.

Is robots.txt permission to scrape?

No. RFC 9309 standardizes crawler instructions; it does not grant access authorization or settle legal obligations.

When should I stop a run?

Stop on explicit denial, CAPTCHA, repeated 403 or 429 responses, or a site’s stated restriction. Seek permission or another permitted data source.