How to Scrape Websites Stealthily with Puppeteer and Playwright
Learn reliable, authorized browser automation with Puppeteer and Playwright, including runnable examples, access checks, troubleshooting, and safer alternatives.
Direct answer: Use Puppeteer or Playwright to automate a real browser only when the site permits your access. There is no dependable way to make browser automation universally “stealthy”: sites can detect mismatches in browser fingerprints, network hints, or behavior, and no plugin guarantees an undetected session. Check the site’s current terms and robots.txt, prefer an API or export, identify your automation honestly, use conservative request rates, and stop if access is denied or challenged.
In this guide, “stealth” means operating reliably and responsibly, not concealing automation or defeating access controls. Technical ability to automate a browser is not permission to collect or copy a site’s content. Puppeteer’s project says that the calling code is responsible for using its capabilities safely and as intended (Puppeteer security policy).
1. Check access before you automate
- Look for an official API, data export, or documented integration. It is often a more stable and appropriate source than rendered pages.
- Read the target site’s current terms and access instructions. Check what content and use are permitted, and whether you need authorization.
- Review
robots.txt. RFC 9309 describes crawler instructions that crawlers are requested to honor. A robots file is one input to your decision; it does not settle every legal or contractual question. See RFC 9309. - Limit collection and request rate. Retrieve only what the task needs, avoid unnecessary repeat visits, and choose a conservative pace based on the site’s published guidance and your authorization. There is no universal safe numeric rate.
- Stop on a denial or challenge. Do not try to work around a CAPTCHA, bot check, login restriction, or other access control. Seek permission or an approved access method.
Terms and applicable requirements depend on the site and circumstances. Cloudflare’s sample terms on automated scraping for AI development are written for site operators, are not universal rules for scrapers, and explicitly are not legal advice (Cloudflare sample terms).
2. Choose Puppeteer or Playwright
Both tools automate browsers and can read rendered page content. Pick one that fits your existing application and browser requirements. This example uses Node.js with each framework separately; install and run only the version you intend to use.
| Question | What to consider |
|---|---|
| Can an API or export provide the data? | Prefer that documented access path when it meets the need. |
| Does the page require browser rendering? | Use a browser framework only for the parts that need it. |
| Do you need local or hosted browsers? | Consider authorization, data handling, session needs, concurrency, reliability, observability, and cost. |
| Does the site require a login or present a challenge? | Proceed only through an approved, authorized flow; stop at denials and bot challenges. |
Browserless documents connections for Puppeteer and Playwright, browser sessions, content scraping, and crawl APIs. It is one hosted infrastructure example, not an endorsement; the cited documentation does not establish comparative performance or pricing (Browserless API overview).
3. Set up a small, permitted collection
Use a project directory and a current Node.js installation. Keep the target URL and selectors specific to a site you are authorized to access. These examples visit one page, extract a heading and links, and print JSON. They do not include fingerprint spoofing, proxy rotation, CAPTCHA solving, or techniques for hiding automation.
Puppeteer
npm init -y
npm install puppeteer
Save as scrape-puppeteer.mjs. Replace the example URL and selectors with ones appropriate to your permitted target.
import puppeteer from 'puppeteer';
const url = 'https://example.com/';
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
page.setDefaultNavigationTimeout(30_000);
const response = await page.goto(url, { waitUntil: 'domcontentloaded' });
if (!response || !response.ok()) {
throw new Error(`Navigation failed: ${response?.status() ?? 'no response'}`);
}
const data = await page.evaluate(() => ({
title: document.title,
heading: document.querySelector('h1')?.textContent?.trim() ?? null,
links: [...document.querySelectorAll('a[href]')]
.map((a) => ({ text: a.textContent.trim(), href: a.href }))
.slice(0, 20),
}));
console.log(JSON.stringify({ url, ...data }, null, 2));
} finally {
await browser.close();
}
Run it with node scrape-puppeteer.mjs. The finally block closes the browser after success or error. A domcontentloaded navigation wait avoids waiting for every subresource; if the required content is rendered later, wait for a specific selector with a bounded timeout instead of adding an unbounded sleep.
Playwright
npm init -y
npm install playwright
npx playwright install chromium
Save as scrape-playwright.mjs.
import { chromium } from 'playwright';
const url = 'https://example.com/';
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage();
page.setDefaultNavigationTimeout(30_000);
const response = await page.goto(url, { waitUntil: 'domcontentloaded' });
if (!response || !response.ok()) {
throw new Error(`Navigation failed: ${response?.status() ?? 'no response'}`);
}
const data = await page.evaluate(() => ({
title: document.title,
heading: document.querySelector('h1')?.textContent?.trim() ?? null,
links: [...document.querySelectorAll('a[href]')]
.map((a) => ({ text: a.textContent.trim(), href: a.href }))
.slice(0, 20),
}));
console.log(JSON.stringify({ url, ...data }, null, 2));
} finally {
await browser.close();
}
Run with node scrape-playwright.mjs. Playwright’s browser installation is a separate setup step; install only the browser you plan to use. See the official Playwright documentation and Puppeteer documentation for API details.
cURL, Python, and Node.js for a screenshot
If the job is to capture a page as an image or PDF rather than extract structured fields, a screenshot API can avoid maintaining a local browser. ScreenshotNeo accepts a URL in one GET request. See the ScreenshotNeo API documentation for parameters and formats.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
4. Extract data carefully
Selectors are specific to each page and can change when a site updates its markup. Start with the smallest useful set of fields, handle missing elements as shown above, and validate output before storing or using it. For multiple pages, build a bounded queue and a modest concurrency limit; do not launch an unbounded number of navigations.
- Normalize and validate URLs before following links; restrict traversal to the authorized domain and path scope.
- Deduplicate URLs to avoid repeated work and accidental loops.
- Set navigation and selector timeouts. Treat timeout as a failed page, not as a reason to retry indefinitely.
- Record status and errors with enough context to diagnose the run, while avoiding storage of unnecessary personal or sensitive data.
- Retry only transient failures, with a small bounded retry count and backoff. Never retry a denial or challenge in an attempt to get around it.
5. “Stealth” limits and responsible operation
Automation detection can use multiple signals. Browserless’s January 23, 2026 article describes mismatches across browser fingerprints, network hints, and behavior, and cautions against assuming a plugin defeats advanced detection. That is a vendor-authored characterization, not an independent benchmark (Browserless on stealth scraping).
There is no universal setting that makes Puppeteer or Playwright invisible. Avoid instructions or tools aimed at concealing automation or defeating a site’s access controls. If your legitimate workload is blocked, stop and ask the site owner for authorization, an allowlisted integration, or an API.
6. Or skip the browser setup
If your goal is a screenshot rather than custom data extraction, ScreenshotNeo takes a URL and returns an image or PDF. One GET request is enough:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan. Read the API docs and sign up for 1,000 free screenshots a month, with no card.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Browser executable not found | Playwright’s browser binary is not installed, or Puppeteer’s browser setup is incomplete. | Run Playwright’s browser install command from setup. For either tool, follow its official installation instructions and confirm the browser is available in the runtime environment. |
| Navigation timeout | The page is slow, waiting for all network activity, or stalled. | Use a bounded timeout and an appropriate readiness condition such as domcontentloaded; wait for the specific required selector. Do not keep extending waits or retrying without a limit. |
| Expected text is missing | The selector is wrong, content is rendered later, or the page structure changed. | Inspect the page you are authorized to access, verify the selector, and wait for the needed element with a finite timeout. |
| HTTP error or empty response | The server returned an error, redirected, or served an empty page. | Check the final URL and response status. Confirm the access method is permitted; stop if the site denies the request. |
| CAPTCHA, bot check, or access denied | The site is challenging or refusing automated access. | Stop. Seek permission or use a documented, approved route; do not attempt to defeat the control. |
| Process hangs after extraction | A browser or page remains open after an exception. | Keep browser shutdown in a finally block, close pages you create, and ensure your worker exits cleanly. |
| Results repeat or crawl expands unexpectedly | Links are duplicated, query strings create variants, or traversal has no scope boundary. | Deduplicate canonicalized URLs, enforce an explicit domain/path scope, and cap the work queue. |
8. Performance, reliability, and cost
- Performance: Browser startup and page rendering have overhead. Reuse a browser process for a bounded job when appropriate, but isolate pages and close resources. Collect only necessary fields, and avoid waiting for full network idle unless the target content requires it.
- Reliability: Page markup, network conditions, and site behavior change. Use explicit selectors, finite timeouts, validation, bounded retries for transient errors, and logs. A successful browser navigation does not guarantee that the extracted data is complete or current.
- Concurrency: More simultaneous pages consume more resources and create more requests to the target. Keep concurrency conservative and within the site’s published terms or your authorization.
- Cost: Local browser automation has infrastructure and maintenance costs, including compute, browser updates, and operational time. Hosted browser services may simplify infrastructure but have service-specific costs and data-handling implications; check current provider documentation and terms. No comparative price or performance figures are established here.
- Screenshot cost: ScreenshotNeo’s stated plans are Free: 1,000 shots/month; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; Business: $249 for 1,000,000. Yearly billing gives two months free.
9. Frequently asked questions
Is website scraping legal?
There is no universal answer. It depends on the site, the data, your authorization, and applicable terms and law. Review the actual terms and seek permission or legal advice where appropriate.
Do I need a stealth plugin?
No plugin can guarantee that a browser session will go undetected. This guide does not recommend concealment or access-control evasion.
Can I use these examples on any public page?
A page being publicly reachable does not by itself grant permission to automate access or reuse its contents. Check the site’s rules and use approved access.
When should I choose a screenshot API instead?
Choose one when the required output is a page image or PDF and you do not need to inspect custom page data in your own browser code. For fields or workflows beyond capture, use an authorized API or browser automation approach that fits the task.


