ScreenshotNeo

BlogGuides

Web Scraping APIs With Puppeteer and Playwright

Choose between local browser automation, a managed browser connection, and a stateless scraping API. Get runnable Puppeteer and Playwright examples, cleanup guidance, and a decision checklist.

By the ScreenshotNeo team1 October 202610 min read

Short answer: Use Puppeteer or Playwright when you need browser-level control to render a page, interact with it, or extract data after JavaScript runs. Run the browser locally for full infrastructure ownership, connect your script to a managed browser over WebSocket when you want to keep browser automation code, or use a stateless HTTP endpoint for a one-off render or extraction. A browser is unnecessary for many pages: try a direct HTTP request and parse the returned HTML first.

A “scraping API” can mean either a remote browser your automation script controls or an HTTP service that performs a specific task and returns a result. Those approaches have different control, session, and operational tradeoffs.

1. When do you need a browser?

Start with the least complex approach that returns the data you need:

  • Fetch and parse HTML when the content is already in the server response and you do not need browser interaction.
  • Use a browser when JavaScript creates the content, the page requires navigation or interaction, or you need browser-rendered output such as a screenshot.
  • Use an extraction or rendering API when a service’s documented inputs and output match your task and you do not need a persistent interactive session.

Browser automation consumes more resources than plain HTML work in some managed crawling environments. For example, Apify says its Actors using Puppeteer or Playwright for real browser rendering require at least 1024 MB of memory; that is an Apify platform requirement, not a universal minimum. See Apify’s Actor usage and resources documentation.

2. Choose an execution model

Approach What runs where Consider
Local browser automation Your application launches and controls the browser. Browser installation and versions, compute, control, and closing each task’s context or browser.
Managed browser over WebSocket Your existing automation code connects to a provider-managed browser. Connection compatibility, session limits and lifecycle, latency, geography, data handling, and provider terms.
Stateless HTTP endpoint A single request asks a service to render or extract and return a result. Supported inputs and output, interaction limits, retries, and behavior on dynamic or inaccessible pages.

Puppeteer and Playwright are browser automation libraries, not scraping APIs by themselves. Puppeteer controls Chrome or Firefox through DevTools Protocol or WebDriver BiDi and runs headless by default. Playwright’s browser API documents Chromium, Firefox, and WebKit. See the Puppeteer documentation and Playwright Browser API.

3. Scrape a JavaScript-rendered page with Puppeteer

This local example launches a browser, navigates to a page, waits for a selector, reads text, and closes the browser even if navigation or extraction fails. Use a page and selector you are permitted to access.

npm install puppeteer
// scrape-puppeteer.js
const puppeteer = require('puppeteer');

async function main() {
  const browser = await puppeteer.launch({ headless: true });
  try {
    const page = await browser.newPage();
    await page.goto('https://example.com', {
      waitUntil: 'domcontentloaded',
      timeout: 30000,
    });
    await page.waitForSelector('h1', { timeout: 10000 });
    const result = await page.$eval('h1', el => el.textContent.trim());
    console.log(result);
  } finally {
    await browser.close();
  }
}

main().catch(error => {
  console.error(error);
  process.exitCode = 1;
});

Replace the URL and selector with the target page and the element containing the data. Puppeteer’s Page API documents navigation, evaluation, screenshots, and related page operations.

Extract several records

const records = await page.$$eval('.product-card', cards =>
  cards.map(card => ({
    name: card.querySelector('.name')?.textContent?.trim() ?? null,
    price: card.querySelector('.price')?.textContent?.trim() ?? null,
  }))
);

Keep evaluation self-contained: code passed into page evaluation runs in the page context, so it cannot use variables from the Node.js process unless you pass them explicitly. Optional chaining handles missing sub-elements; validate required fields before persisting records.

4. Scrape a JavaScript-rendered page with Playwright

This example uses Playwright’s Chromium browser locally. The finally block closes the browser on both successful and failed runs.

npm init -y
npm install playwright
npx playwright install chromium
// scrape-playwright.js
const { chromium } = require('playwright');

async function main() {
  const browser = await chromium.launch({ headless: true });
  try {
    const page = await browser.newPage();
    await page.goto('https://example.com', {
      waitUntil: 'domcontentloaded',
      timeout: 30000,
    });
    await page.locator('h1').waitFor({ state: 'visible', timeout: 10000 });
    const result = (await page.locator('h1').innerText()).trim();
    console.log(result);
  } finally {
    await browser.close();
  }
}

main().catch(error => {
  console.error(error);
  process.exitCode = 1;
});

Playwright also supports Firefox and WebKit. Install the browser you intend to use and select the matching browser launcher. See the official Browser API and network documentation.

Wait for the page condition you need

There is no single wait condition that fits every site. domcontentloaded waits for initial document parsing, while a selector wait expresses the actual condition your extraction depends on. A fixed delay can be useful for a known delayed widget, but it may waste time or still finish too early. Avoid assuming that network idleness means an application has finished rendering.

await page.goto('https://example.com/catalog', { waitUntil: 'domcontentloaded' });
await page.waitForSelector('.product-card');

5. Connect an existing script to a managed browser

A managed browser can move browser installation and execution to a provider while retaining much of your automation code. Connection APIs and supported protocols are provider-specific. Browserless documents Puppeteer connect() and Playwright connectOverCDP() integration. Its Playwright guide uses CDP because the connection speaks the Chrome DevTools Protocol, rather than Playwright’s own browser server protocol. Follow the provider’s current docs for its endpoint and authentication format.

The following shows the connection shape. Set BROWSER_WS_ENDPOINT to the WebSocket endpoint and credentials supplied by your chosen provider; the placeholder is not a real endpoint.

// Puppeteer with a provider-supplied WebSocket endpoint
const puppeteer = require('puppeteer-core');

async function main() {
  const endpoint = process.env.BROWSER_WS_ENDPOINT;
  if (!endpoint) throw new Error('Set BROWSER_WS_ENDPOINT');
  const browser = await puppeteer.connect({ browserWSEndpoint: endpoint });
  try {
    const page = await browser.newPage();
    await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
    console.log(await page.title());
  } finally {
    await browser.close();
  }
}
main().catch(error => { console.error(error); process.exitCode = 1; });
// Playwright connecting over CDP to a provider-supplied endpoint
const { chromium } = require('playwright');

async function main() {
  const endpoint = process.env.BROWSER_WS_ENDPOINT;
  if (!endpoint) throw new Error('Set BROWSER_WS_ENDPOINT');
  const browser = await chromium.connectOverCDP(endpoint);
  try {
    const context = browser.contexts()[0] ?? await browser.newContext();
    const page = await context.newPage();
    await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
    console.log(await page.title());
  } finally {
    await browser.close();
  }
}
main().catch(error => { console.error(error); process.exitCode = 1; });

Use the exact connection method supported by your provider and browser. A provider’s session may remain active until a timeout if it is left open and can consume usage units; Browserless’s guide demonstrates closing the browser in a finally block. Billing behavior depends on the service. Do not assume that closing a local object, disconnecting a client, or letting a connection expire has the same effect across providers. Consult the provider’s lifecycle and billing documentation.

6. Use a stateless HTTP scraping or rendering API

A REST endpoint is a good fit when you can describe the task in one request and accept the endpoint’s supported controls and response format. Browserless documents separate REST routes for smart scraping, rendered content, CSS-selector extraction, and screenshots, rather than one universal scraping interface. Choose a route by required task and output. See its REST API documentation and managed browser overview.

Endpoint paths, authentication, request schemas, limits, and billing are provider-specific. Do not copy a route or parameter from one service into another. Before adopting an endpoint, check how it represents extraction failures, timeouts, blocked pages, retries, and partial output.

7. Configure sessions, network access, and extraction

Keep session state deliberate

For independent jobs, use a fresh browser context or equivalent isolation where supported. If the workflow requires cookies or a logged-in session, provide only the state needed for that task, handle it as a secret, and close the context when finished. Do not accidentally reuse one user’s authenticated context for another job.

Configure proxy routing only when needed

Playwright documents HTTP(S) and SOCKSv5 proxy configuration at the browser or context level. A proxy changes network routing; it does not establish permission to collect a site’s content or guarantee that a request will succeed. Follow the target site’s terms and applicable law.

const browser = await chromium.launch({
  proxy: { server: 'http://proxy.example:8080' },
});

Proxy credentials and configuration options vary by proxy and library. Avoid logging secrets in URLs or error output. See Playwright’s official HTTP proxy documentation.

Extract data defensively

  • Prefer stable, meaningful selectors over brittle positional selectors.
  • Wait for the specific element or state that contains the data.
  • Handle missing values and duplicate elements explicitly.
  • Normalize whitespace and types, and validate records before storing them.
  • Record the URL and a useful error category for failed records without logging credentials or sensitive page content.

8. Reliability, performance, and cost

  • Resource use: A real browser includes a browser process and page resources. Limit concurrent browsers to fit available memory and CPU, and close pages, contexts, and browsers when work ends. Platform resource figures are not universal; use the requirements of your own runtime or provider.
  • Wait strategy: Waiting for one relevant selector is often more targeted than waiting for every network request to stop. Set navigation and selector timeouts; treat them as bounds, not proof the page is complete.
  • Retries: Retry transient network or provider errors with a small bounded policy and backoff. Avoid retrying indefinitely or treating a selector mismatch as a transient network failure.
  • Concurrency: Bound parallel jobs. More pages at once can raise memory and CPU use and may change request patterns seen by the target.
  • Remote sessions: Account for connection latency, service limits, session duration, and provider billing rules. Close sessions promptly and check usage after implementing the lifecycle.
  • Stateless APIs: Compare the endpoint’s documented per-request billing and limits with the cost of operating a browser. The sources reviewed do not establish neutral cross-provider prices, success rates, or performance rankings.

9. Troubleshooting

Symptom Likely cause What to do
Browser executable missing The browser binary was not installed in the environment or the installed package expects a separate browser installation. Install the browser required by your library in the deployment image; for Playwright, run the documented browser install command during setup.
Navigation timeout The page is slow, the network is unavailable, or the chosen wait condition never occurs. Check connectivity and the target URL, use a condition appropriate to the task, and set explicit navigation and selector timeouts. Do not simply remove all bounds.
Selector timeout or empty result The selector is wrong, content is not rendered yet, or the page structure changed. Inspect the rendered page and confirm the selector; wait for the relevant element and handle absent or optional fields.
Works locally but fails in deployment Missing browser dependencies, different browser versions, memory limits, or deployment network restrictions. Install browser dependencies in the runtime image, align versions, check outbound access, and size resources for the workload.
WebSocket connection rejected Incorrect endpoint or credentials, unsupported protocol, or provider session limits. Verify the provider’s current endpoint and authentication instructions, use its documented library connection method, and check session limits.
Usage continues after a job ends A remote browser session may remain active until timeout. Close the browser in a finally block and check the provider’s billing and session lifecycle rules.
Duplicate or incomplete records Repeated elements, lazy loading, pagination, or an early extraction point. Wait for the needed state, explicitly traverse pages or load-more controls when permitted, deduplicate by a stable field, and validate record completeness.

10. Which approach should you use?

  1. Try a direct HTTP request first if the required content is present in the returned HTML.
  2. Choose local Puppeteer or Playwright if you need detailed browser interaction and want to own the runtime.
  3. Choose a managed WebSocket browser if you want to keep a browser automation script while moving browser hosting to a provider.
  4. Choose a stateless REST endpoint if the task is a supported one-shot render or extraction and its limits fit your requirements.
  5. Check site terms and applicable law, then account for browser resources, session lifecycle, and provider-specific costs.

11. Or skip the browser setup

If your task is to capture a page as an image or PDF, ScreenshotNeo provides a website screenshot API and MCP server. Its one-call API returns a screenshot or PDF; see the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners are accepted and removed before capture, along with known newsletter popups and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers indicate the page verdict and billing status. The MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for 1,000 free screenshots a month, with no card required.

FAQ

Are Puppeteer and Playwright scraping APIs?

No. They are libraries for controlling browsers. A managed browser endpoint or stateless HTTP service is a separate provider service.

Can a scraping API run my existing Playwright script?

A managed browser may support connecting an existing script, but compatibility depends on its connection protocol and provider instructions. A stateless HTTP endpoint generally accepts a task request rather than an arbitrary Playwright program.

Do I need a browser for every scraped page?

No. If the response HTML already contains the data and no interaction is needed, a direct HTTP client and HTML parser may be simpler.

Does using a proxy make scraping permitted?

No. A proxy changes routing, not the permission status of collection. Check applicable law and the target site’s terms.