ScreenshotNeo

BlogHow-to

How to Download a Web Page With JavaScript

Use Playwright or Puppeteer to save JavaScript-rendered downloads, HTML, and PDFs with reliable waits and complete runnable examples.

By the ScreenshotNeo team1 October 20267 min read

Direct answer

A normal HTTP client downloads the server response but does not execute page JavaScript. To save what a user sees after scripts run, open the URL in a real browser with Playwright or Puppeteer, wait for a page-specific readiness signal, then save a download, the rendered DOM, or a PDF. The artifact you need determines the API.

Choose the artifact first

Need Use Why
File generated by a button Playwright download event Captures the exact attachment and filename.
HTML after JavaScript Puppeteer page.content() or the Playwright equivalent Returns the current DOM, including the DOCTYPE.
Shareable visual document Puppeteer or Playwright PDF Renders print CSS, or screen CSS when selected.
Several browser engines Playwright One API with its documented browser installation workflow.

Option 1: download a file after a click with Playwright

Start waiting for the event before clicking. Playwright emits the download event once the download starts; waiting first prevents a race. See the Playwright Download API.

import { chromium } from 'playwright';

const browser = await chromium.launch();
const context = await browser.newContext({ acceptDownloads: true });
const page = await context.newPage();
await page.goto('https://example.com/reports', { waitUntil: 'domcontentloaded' });
await page.getByRole('button', { name: 'Download file' }).waitFor();
const downloadPromise = page.waitForEvent('download');
await page.getByRole('button', { name: 'Download file' }).click();
const download = await downloadPromise;
await download.saveAs('/tmp/' + download.suggestedFilename());
await browser.close();

Install with npm i -D playwright, then install the browser binaries and operating-system dependencies using Playwright’s documented browser installation command. Browser-context downloads are temporary and are deleted when the context closes, so call saveAs() or copy the file before closing the context.

Download edge cases

  • If the click opens a new tab, wait for the popup and download as appropriate.
  • If a form submits, create the download wait around the submit action rather than after it.
  • Authenticated files require a context with the required cookies, headers, or login flow.
  • A browser cannot bypass a paywall, CAPTCHA, bot defense, or access control. Use only credentials and interactions you are authorized to use.

Option 2: save rendered HTML with Puppeteer

page.content() returns the full current HTML, including the DOCTYPE. Navigate, wait for the application’s content, then write the string. See the Puppeteer Page API.

import puppeteer from 'puppeteer';
import { writeFile } from 'node:fs/promises';

const browser = await puppeteer.launch();
const page = await browser.newPage();
await page.goto('https://example.com/app', { waitUntil: 'networkidle2' });
await page.waitForSelector('main');
const html = await page.content();
await writeFile('rendered.html', html, 'utf8');
await browser.close();

Install with npm i puppeteer. Puppeteer downloads a compatible Chrome during installation; if package install scripts are blocked, install the required browser explicitly.

The file contains the current markup only. External stylesheets, images, fonts, and API responses remain separate resources, so this is not a self-contained archive. To make an offline copy, download those resources and rewrite URLs, or save a PDF instead.

Option 3: save a rendered PDF

Puppeteer’s page.pdf() generates a PDF with print CSS media. Select screen media when the page’s screen layout is what you need.

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch();
const page = await browser.newPage();
await page.goto('https://example.com/app', { waitUntil: 'networkidle2' });
await page.waitForSelector('main');
await page.emulateMediaType('screen');
await page.pdf({
  path: 'page.pdf',
  format: 'A4',
  printBackground: true,
  preferCSSPageSize: true
});
await browser.close();

Playwright exposes the same workflow with page.pdf({ path: 'page.pdf' }). Choose paper size, margins, landscape mode, and page ranges deliberately. See Puppeteer Page.pdf(), the PDF guide, and the Playwright PDF API.

Waiting for JavaScript correctly

Readiness is page-specific. Prefer a condition tied to the content you need:

await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.waitForSelector('[data-ready=true]');
// or: await page.waitForFunction(() => window.app?.loaded === true);
// or, when appropriate: await page.waitForLoadState('networkidle');
  • domcontentloaded means the initial document is parsed; client rendering may still be running.
  • networkidle2 waits for a low number of active connections, but analytics or streams can keep a page busy.
  • A selector or application flag is usually more reliable than an arbitrary sleep.
  • Lazy-loaded sections may need scrolling into view before capture.

Configuration checklist

  1. Set realistic navigation and action timeouts; catch timeout errors and save diagnostics.
  2. Use a deterministic viewport, locale, timezone, and color scheme when output must be reproducible.
  3. Handle cookie-consent dialogs before waiting for the final selector.
  4. Supply authorized authentication state through cookies, storage state, or headers.
  5. Close pages and browsers in a finally block so failed jobs do not leak processes.
  6. Persist downloaded files before closing the browser context.

Troubleshooting

Symptom Likely cause Fix
Saved HTML is only a shell Captured before client rendering finished. Wait for a content selector or application-ready flag after navigation.
TimeoutError waiting for network idle Analytics, WebSockets, or polling never become idle. Use a specific selector or function with a bounded timeout.
Download event never arrives Listener started after the click, or the click did not trigger a download. Create waitForEvent('download') first; verify the locator and inspect whether a popup or navigation occurred.
Download disappears Context closed before the temporary file was copied. Call download.saveAs() before teardown.
PDF misses images or fonts Resources are still loading or print CSS differs. Wait for the relevant selector, select screen media when needed, and use printBackground: true.
Browser executable missing Browser binaries or OS dependencies were not installed. Run the framework’s browser install command; for Puppeteer, install Chrome explicitly when install scripts are disabled.
403, CAPTCHA, or login wall Site access control or bot defense. Use authorized authentication and follow site rules; automation does not guarantee bypass.

More failure cases

If the download event never resolves, confirm that the listener was registered before the click and that no consent or authentication dialog blocks the control. If rendered content is missing, replace fixed sleeps with a page-specific selector and trigger lazy loading by scrolling when necessary. Redirects and consent dialogs should be handled in the browser context before waiting for the final readiness condition. A timeout can mean the selector never appears, the page keeps connections open, or access is denied; increasing the timeout alone does not prove that content loaded.

Performance, reliability, and cost

Launching a browser is heavier than an HTTP request. Reuse one browser process across jobs, create isolated contexts per user or task, and avoid waiting on global network idle when a stable selector is available. Set bounded timeouts, retry only transient navigation failures, and record the URL, readiness condition, browser version, and error details for reproducibility. Cache static inputs where your use case allows it. Browser runtime and infrastructure costs depend on hosting, concurrency, and page behavior; the official Playwright and Puppeteer documentation does not publish a universal benchmark.

Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF, while the service handles the browser work. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

ScreenshotNeo supports full-page capture with lazy images loaded, element selectors, dark mode, device presets and custom viewports, retina scale, PDFs with paper size, margins, landscape and page ranges, HTML/CSS to image, custom CSS and JavaScript, clicks, hide selectors, waits, ad and tracker blocking, custom headers, cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, resizing, caching with a chosen TTL, signed links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work.

Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; each cleanup step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Every response reports the result through X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is on every plan.

Start with 1,000 free screenshots a month—no card required.

FAQ

Can wget, curl, or requests download a JavaScript page?

They can download the server response, but they do not execute browser JavaScript. Use browser automation or a rendering API.

Should I save HTML or PDF?

Save HTML when you need the post-JavaScript DOM for processing. Save PDF when you need a stable visual document to share or archive.

Why does a page look different in the PDF?

PDF generation uses print CSS by default. Call emulateMediaType('screen') when you need screen styling, and wait for fonts and images before generating it.

Can browser automation defeat a CAPTCHA?

No guarantee exists. Access controls and bot defenses require authorized handling and may prevent capture.

How do I tell whether a ScreenshotNeo response was billable?

Read the X-Page-Verdict and X-Billed headers on every response.