ScreenshotNeo

BlogHow-to

Using the Puppeteer Node.js SDK for Remote Browser Automation

Connect Puppeteer to a hosted browser with a WebSocket endpoint, then automate pages remotely while handling sessions, files, latency and concurrency.

By the ScreenshotNeo team1 October 20267 min read

Use puppeteer.connect() with the remote browser’s WebSocket endpoint in browserWSEndpoint. For the Browserless managed-browser flow documented here, install puppeteer-core, connect to the provider-issued wss:// URL, automate pages as usual, and close the remote session in a finally block.

Puppeteer is a JavaScript library with a high-level API for automating Chrome and Firefox through Chrome DevTools Protocol and WebDriver BiDi. Navigation, selectors, waits, DOM evaluation, screenshots and PDFs continue to use familiar page APIs when the browser runs remotely. The differences are the connection endpoint, session lifecycle, filesystem, environment defaults, network latency and concurrency accounting.

1. Install Puppeteer for a remote browser

npm install puppeteer-core

puppeteer-core does not download a local Chromium binary, which suits a remote-only workflow. The full puppeteer package can also call connect(), but it downloads a browser binary that this workflow does not need.

2. Connect to Browserless with Node.js

Browserless is a provider-specific example. Copy the secure WebSocket endpoint and token format from the provider’s current documentation; endpoint paths and query parameters differ between hosting services. Keep the credential out of source control and do not log the complete URL.

import puppeteer from 'puppeteer-core';

const endpoint = process.env.BROWSER_WS_ENDPOINT;
if (!endpoint || !endpoint.startsWith('wss://')) {
  throw new Error('BROWSER_WS_ENDPOINT must be a provider-issued wss:// URL');
}

const browser = await puppeteer.connect({
  browserWSEndpoint: endpoint,
});

try {
  const page = await browser.newPage();
  await page.setViewport({ width: 1440, height: 900, deviceScaleFactor: 1 });
  await page.goto('https://example.com', {
    waitUntil: 'networkidle2',
    timeout: 60_000,
  });

  console.log('Title:', await page.title());
  await page.screenshot({ path: 'example.png', fullPage: true });
} finally {
  await browser.close();
}

Set BROWSER_WS_ENDPOINT to the provider-issued wss:// URL. Browserless documents authentication in the endpoint query string, commonly with a token parameter. Treat that format as provider-specific and follow the selected service’s current instructions. See the Browserless documentation and Chrome for Developers Puppeteer documentation.

3. What remains the same after connecting

Most page-level automation does not change:

const page = await browser.newPage();
await page.goto('https://news.ycombinator.com', { waitUntil: 'domcontentloaded' });
await page.waitForSelector('a');
const links = await page.$$eval('a', nodes =>
  nodes.map(node => ({ text: node.textContent?.trim(), href: node.href }))
);
const heading = await page.$eval('h1', node => node.textContent?.trim());
await page.pdf({ path: 'page.pdf', format: 'A4', printBackground: true });

Selectors, navigation, explicit waits, evaluation, cookies, request interception and page events use the same Puppeteer APIs. Remote execution changes where those operations run, not their basic syntax.

4. Remote execution differences to design for

Concern What changes Practical approach
Connection Use a WebSocket endpoint with connect() instead of starting a local process with launch(). Validate that the endpoint is wss:// and provider-authenticated.
Session lifecycle The connection consumes a remote browser session. Always call browser.close() in finally, including error paths.
Files The browser host cannot see paths on your Node.js machine. Use the provider’s upload/download mechanism, or transfer generated bytes through your application.
Environment Viewport, user agent, timezone and locale may differ from local Chrome. Set them deliberately when visual or behavioral parity matters.
Latency Requests cross the network, while browser-to-target-site traffic occurs from the remote region. Choose a browser region near the sites you automate and avoid unnecessary round trips.
Concurrency Each connection is a session and counts toward the provider’s concurrency limit. Reuse one browser connection for pages in one job; open separate connections for genuinely parallel jobs.
Browser flags The browser may start before your client connects. Provider-specific launch settings may need endpoint query parameters; array values can require encoded JSON.

5. Set the remote browser environment explicitly

const page = await browser.newPage();
await page.setViewport({ width: 1366, height: 768, deviceScaleFactor: 1 });
await page.setUserAgent(
  'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 Chrome/131 Safari/537.36'
);
await page.emulateTimezone('UTC');
await page.setExtraHTTPHeaders({ 'Accept-Language': 'en-US,en;q=0.9' });

await page.goto('https://example.com', {
  waitUntil: 'networkidle2',
  timeout: 60_000,
});

Use the same values in local and remote runs when comparing screenshots or test results. A changed locale, timezone, font set or user agent can alter layout, dates and content even when your selectors are unchanged.

6. Handle files and downloads

A path such as /tmp/report.pdf belongs to the machine running Chromium, not automatically to the Node.js process. For downloads, ask the provider how to retrieve files, or stream content through your application when the response is available in memory.

const response = await page.goto('https://example.com/report.pdf', {
  waitUntil: 'networkidle0',
  timeout: 60_000,
});
if (!response || !response.ok()) {
  throw new Error(`Download failed: ${response?.status() ?? 'no response'}`);
}
const bytes = await response.buffer();
await import('node:fs/promises').then(fs => fs.writeFile('report.pdf', bytes));

For browser-generated downloads, configure the provider’s file-transfer API rather than assuming the remote filesystem is mounted locally.

7. Reuse connections and control concurrency

const browser = await puppeteer.connect({ browserWSEndpoint: endpoint });
try {
  const urls = ['https://example.com', 'https://developer.chrome.com'];
  for (const url of urls) {
    const page = await browser.newPage();
    try {
      await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 60_000 });
      console.log(url, await page.title());
    } finally {
      await page.close();
    }
  }
} finally {
  await browser.close();
}

One connection with several short-lived pages is usually the simplest model for a single job. Separate connections make sense for independent parallel jobs, but each one consumes another provider session. Limit worker count to the documented concurrency for your plan and add backoff for transient connection failures.

8. Reliability checklist

  • Keep the endpoint and token in environment variables or a secret manager.
  • Use explicit navigation and selector timeouts instead of relying on defaults.
  • Close pages after each unit of work and close the browser in finally.
  • Record URL, provider region, viewport, user agent and failure stage for reproducibility.
  • Retry only idempotent jobs, with bounded exponential backoff; do not duplicate purchases or other side effects.
  • Check provider session, timeout and concurrency limits before increasing parallelism.

9. Troubleshooting remote Puppeteer

Symptom Likely cause Fix
Invalid URL or connection refused An HTTPS page URL was supplied instead of a WebSocket endpoint. Use the provider-issued wss:// URL for browserWSEndpoint.
Authentication failure Missing, expired or incorrectly encoded provider token. Regenerate credentials, follow the provider’s query-parameter format and keep the full URL out of logs.
Works locally, different layout remotely Viewport, user agent, timezone, locale, fonts or browser version differ. Set those values explicitly and compare the remote environment.
Local file not found The path exists only on the Node.js host. Use provider upload/download APIs or transfer bytes through your application.
Jobs remain active or usage keeps increasing A code path did not close the browser. Put browser.close() in an unconditional finally block.
Parallel jobs are rejected The provider concurrency limit was reached. Reduce worker count, reuse connections within jobs and inspect plan limits.
Navigation timeout Target site is slow, blocked from the remote region or waiting for a never-idle resource. Choose an appropriate waitUntil, increase timeout carefully, wait for a meaningful selector and verify regional access.
Browser flags have no effect The browser launched before the client connected. Use the provider’s documented endpoint parameters; encode array options as required.

10. Performance and cost considerations

Remote automation adds network setup and command latency. Keep related actions in one page, avoid repeated connect and disconnect cycles, and place the browser region close to the target sites. Reuse a connection within a job, while balancing that against session duration and provider limits.

Provider billing and limits are service-specific. An unclosed Browserless session remains active until timeout and may accrue charges, so cleanup is both a reliability and cost requirement. Do not publish provider prices or concurrency numbers without checking the current plan.

11. When a screenshot API is simpler

If your job only needs a rendered image or PDF, a browser SDK can be more infrastructure than necessary. ScreenshotNeo provides a GET endpoint for screenshots and PDFs, plus an MCP server for AI agents. It removes known cookie and consent banners, newsletter popups and chat widgets before capture; bot checks, blank pages, failed loads, timeouts and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

12. Or skip the browser setup

Use the ScreenshotNeo API when you want a clean capture without managing Chromium, WebSocket sessions or remote files. The API supports full-page and element captures, device presets or custom viewports, dark mode, retina scale, PDF options, custom CSS and JavaScript, clicks, waits, blocking rules, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous jobs, bulk capture and usage reporting. Every feature is available on every plan.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for parameters and response headers. Cookie banners, popups and chat widgets are removed before the shot; bot checks, blank pages and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account to get started.

13. FAQ

Can I use the puppeteer package instead of puppeteer-core?

Yes. Both can call puppeteer.connect(). For a remote-only workflow, puppeteer-core avoids downloading a local browser binary.

Does remote Puppeteer require rewriting selectors?

No. Page APIs remain the same. Differences usually come from environment settings, network access or session handling.

Should concurrent scripts share one connection?

Reuse a browser object for pages in one job. Independent parallel jobs should use separate connections and stay within the provider’s concurrency limit.

Can the remote browser read my local uploads?

No. Transfer files through the provider’s upload and download mechanisms.

What must be protected?

The WebSocket endpoint often contains a token. Store it as a secret, avoid committing it, and redact it from logs and error reports.