ScreenshotNeo

BlogHow-to

How to Capture Web Pages After They Finish Rendering

Wait for the content your page needs, verify it exists, then capture. Learn reliable Playwright, Selenium, Chrome and API workflows.

By the ScreenshotNeo team29 September 202610 min read

How to Capture Web Pages After They Finish Rendering

Direct answer: open the page in a real browser, wait for a signal tied to the content you need, verify that content is present and stable, then capture the screenshot or PDF. A browser lifecycle event such as load is not a universal definition of “finished.” For reliable results, prefer a visible locator, an application readiness flag, or a bounded combination of both.

This guide shows complete workflows for Playwright, Selenium, Chrome Headless, Puppeteer-style browser automation, and low-level Chrome DevTools Protocol (CDP). It covers dynamic data, lazy loading, iframes, shadow DOM, virtualized lists, fonts, animation, authentication, PDFs, performance, reliability, and operating cost.

1. What “finished rendering” means

Rendering completion is page-specific. A server-rendered article may be ready at domcontentloaded. A dashboard can fire load while API requests are still filling charts. A single-page app may continue changing after the network becomes quiet.

Wait for the content signal that proves the rendered page is ready before capturing.
Wait for the content signal that proves the rendered page is ready before capturing.

Choose a readiness condition that represents the output you require:

Readiness signal Use it when Limit
Visible element or assertion You know the heading, result count, chart, table, or article body that must appear Requires a stable selector and expected state
Application flag or event The application exposes a hydration or data-ready signal Requires cooperation from the page
domcontentloaded Markup is the main output and JavaScript is not required Dynamic content may not exist yet
load Subresources such as images are important and the page has a predictable load lifecycle API-driven UI can still be incomplete
networkidle You need a diagnostic fallback or a page with finite network activity Long polling, analytics, ads, WebSockets, and post-idle layout work make it early or unreachable
Bounded delay Fonts, image decoding, or a short transition needs time after a stronger signal A fixed sleep alone is fragile

Playwright defines networkidle as no network connections for at least 500 ms and marks it discouraged for tests. Use a content assertion first, then add a short, documented stabilization delay only when you have a known reason. See the Playwright wait-for-load-state documentation.

Playwright provides browser control, explicit waits, screenshots, and PDF generation in one API. Install it with npm install playwright, then install a browser with npx playwright install chromium.

import { chromium } from 'playwright';

const browser = await chromium.launch();
const page = await browser.newPage({
  viewport: { width: 1440, height: 900 },
  deviceScaleFactor: 1
});

await page.goto('https://example.com/article', {
  waitUntil: 'domcontentloaded',
  timeout: 30_000
});

// Replace this with a selector that proves the content you need exists.
const heading = page.locator('article h1');
await heading.waitFor({ state: 'visible', timeout: 30_000 });

// Optional: assert that the element contains meaningful content.
await expect(heading).toContainText('Example');

// If the application exposes a stronger readiness flag, wait for that instead.
await page.waitForFunction(() => window.__CONTENT_READY__ === true, null, {
  timeout: 30_000
}).catch(() => {});

// Let fonts and image decoding settle when geometry matters.
await page.evaluate(() => document.fonts.ready);
await page.screenshot({ path: 'article.png', fullPage: true });
await page.pdf({ path: 'article.pdf', format: 'A4', printBackground: true });

await browser.close();

If your page does not define __CONTENT_READY__, remove that block. The important wait is the locator or assertion tied to the content. Playwright’s full-page screenshot captures the page’s scrollable area. Its PDF output uses print CSS media by default, so a PDF can look different from a screen capture; see the PDF documentation.

Wait for a chart, table, or result count

await page.locator('[data-testid="sales-chart"]').waitFor({ state: 'visible' });
await expect(page.locator('[data-testid="result-count"]')).toHaveText(/\d+/);
await expect(page.locator('table tbody tr')).toHaveCount(25);

Stabilize layout before capture

await page.evaluate(async () => {
  await document.fonts.ready;
  await new Promise(requestAnimationFrame);
  await new Promise(requestAnimationFrame);
});

const boxBefore = await page.locator('article').boundingBox();
await page.waitForTimeout(150);
const boxAfter = await page.locator('article').boundingBox();
if (!boxBefore || !boxAfter || Math.abs(boxBefore.height - boxAfter.height) > 1) {
  throw new Error('Layout is still changing');
}

3. Make lazy content and virtualized pages complete

A full-page screenshot does not automatically guarantee that every lazy image has loaded. Many applications mount list rows only while they are near the viewport. Scroll through the document or use the application’s own “load more” interaction before capture.

Scrolling can materialize lazy images and virtualized content before a full-page capture.
Scrolling can materialize lazy images and virtualized content before a full-page capture.
await page.evaluate(async () => {
  const distance = Math.max(document.body.scrollHeight, document.documentElement.scrollHeight);
  for (let y = 0; y < distance; y += 700) {
    window.scrollTo(0, y);
    await new Promise(resolve => setTimeout(resolve, 100));
  }
  window.scrollTo(0, 0);
});

await page.waitForFunction(() =>
  [...document.images].every(img => img.complete || img.loading === 'lazy')
);
await page.screenshot({ path: 'long-page.png', fullPage: true });

For a virtualized list, scrolling alone may not preserve all rows in the DOM. Capture each viewport and stitch the result, or use the product’s export endpoint if one exists. For an iframe, select the correct frame and wait inside it:

const frame = page.frameLocator('iframe[data-report]');
await frame.locator('[data-testid="report-ready"]').waitFor({ state: 'visible' });

Shadow DOM content must be accessed through component-supported locators or evaluated inside the shadow root. Cross-origin frames cannot be inspected as if they belonged to the top-level document; wait for a visible frame result or capture through the browser context.

4. Selenium, Chrome Headless, Puppeteer, and CDP

Selenium with Python

Selenium fits teams already standardized on WebDriver. Use an explicit wait for the target content rather than a fixed sleep.

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

options = webdriver.ChromeOptions()
options.add_argument('--headless=new')
driver = webdriver.Chrome(options=options)
try:
    driver.get('https://example.com/article')
    WebDriverWait(driver, 30).until(
        EC.visibility_of_element_located((By.CSS_SELECTOR, 'article h1'))
    )
    WebDriverWait(driver, 30).until(
        lambda d: d.execute_script('return document.fonts.status') == 'loaded'
    )
    driver.save_screenshot('article.png')
    text = driver.find_element(By.CSS_SELECTOR, 'article').text
    print(text)
finally:
    driver.quit()

Selenium can execute JavaScript, capture an element or viewport, and print a page to PDF through browser-specific support. Its explicit waits are useful when network and rendering time vary.

Chrome Headless CLI

chrome --headless --screenshot=page.png --timeout=5000 https://example.com/article
chrome --headless --print-to-pdf=article.pdf --timeout=5000 https://example.com/article

Chrome’s --timeout is a maximum wait in milliseconds before capture, even if the page is still loading. It is a ceiling, not proof that the target content is ready. For important captures, pair it with a page-specific readiness mechanism. See Chrome Headless documentation.

Puppeteer and CDP

Puppeteer is a JavaScript library for automating Chrome and Firefox through Chrome DevTools Protocol and WebDriver BiDi. It supports screenshots, PDFs, navigation, interaction, and performance analysis. CDP gives lower-level control through methods such as Page.navigate, lifecycle events, layout metrics, Page.captureScreenshot, and Page.printToPDF. Choose CDP when you are building a capture service that needs custom orchestration; choose Playwright or Puppeteer for a higher-level API.

  • Authentication: create the session before the readiness wait. Reuse a browser context with saved cookies or set an authorization header before navigation.
  • Consent dialogs: accept or dismiss them before waiting for the target element; an overlay can hide content or prevent interaction.
  • Animations: disable transitions with injected CSS or wait for them to finish when pixel-stable output matters.
  • Fonts: wait for document.fonts.ready when text wrapping or geometry affects the result.
  • Viewport and scale: fix viewport dimensions, device scale factor, timezone, locale, and user agent when comparing captures.
  • Print output: PDFs use print media rules, margins, paper size, and page breaks. Validate those separately from screen screenshots.
await page.addStyleTag({ content: `
  *, *::before, *::after {
    animation-duration: 0s !important;
    animation-delay: 0s !important;
    transition: none !important;
    caret-color: transparent !important;
  }
` });

6. Troubleshooting incomplete or blank captures

Symptom Likely cause Fix
Blank shell or missing API data Captured at domcontentloaded or load before hydration Wait for the result element, expected text, or application readiness flag
Screenshot times out on networkidle Analytics, polling, ads, or WebSockets never stop Use a content locator and a bounded timeout; treat network idle as a fallback
Images are missing Lazy loading or decoding has not completed Scroll the page, wait for image completion, then capture
Iframe is empty The top-level page is ready while the frame is not Wait inside the correct frame; account for cross-origin limits
Rows disappear in a long table Virtualization keeps only visible rows mounted Scroll and capture segments, or use an export endpoint
Text wraps differently between runs Fonts or viewport are changing Pin viewport and scale; await document.fonts.ready
Two captures differ Animation, rotating content, ads, or layout shift Disable motion, block unstable resources, and verify bounding-box stability
Login page is captured Session cookies or headers were not applied Establish authentication before navigation and assert the signed-in marker
PDF differs from screenshot Print CSS media rules change layout Review print styles, paper size, margins, and page breaks separately

7. Performance, reliability, and cost decisions

Browser startup is often the largest fixed cost. Reuse a browser process and create isolated contexts for concurrent jobs. Keep each job bounded with navigation and readiness timeouts. Capture only after the required signal, because unnecessary delays increase latency without improving correctness.

Parallelism improves throughput until CPU, memory, bandwidth, or the target site becomes the bottleneck. Limit concurrent pages, record navigation and readiness durations, and preserve failure diagnostics such as the URL, selector, timeout stage, console errors, and a small trace or HTML snapshot when policy allows.

Retries should be selective. Retry transient navigation failures and overloaded upstreams with exponential backoff. Do not blindly retry deterministic selector failures; fix the readiness condition or page-specific handling. Cache stable pages when freshness permits, but include the cache key inputs that affect output: URL, viewport, user agent, cookies, headers, timezone, geolocation, and custom CSS or JavaScript.

DOM extraction, screenshots, PDFs, and HTML snapshots preserve different parts of the rendered state. DOM and text are best for semantic checks, screenshots preserve canvas and visual layout, PDFs preserve printable records, and HTML snapshots can omit runtime state or shadow-DOM internals unless serialized after rendering.

8. Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response reports the result through X-Page-Verdict and X-Billed headers.

Use the same readiness controls without managing a browser process. The API supports full-page capture with lazy images loaded, CSS element capture, dark mode, 12 device presets or any viewport, retina scale, PDF paper size and margins, page ranges, custom CSS and JavaScript, clicks, selector waits, delay or network-idle waits, blocking ads, trackers, requests, or resource types, custom headers, cookies, user agent and Authorization, timezone, geolocation, transparent backgrounds, image resizing, selectable TTL caching, signed links, asynchronous jobs with signed webhooks, bulk capture of 100 URLs per call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which helps when switching.

See the ScreenshotNeo API documentation for the complete option list.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', data));

An MCP server also exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

9. Practical checklist

  1. Define what “complete” means for this page.
  2. Navigate with a real browser engine.
  3. Wait for a content locator, assertion, or application-ready signal.
  4. Handle authentication and consent before the readiness wait.
  5. Materialize lazy and virtualized content.
  6. Wait for fonts and stabilize layout when geometry matters.
  7. Disable or wait for animations.
  8. Capture the required format and validate the expected element or pixels.
  9. Record timing, failures, and the readiness stage.
  10. Use bounded retries and cache only when the inputs and freshness policy are explicit.

10. FAQ

Is networkidle the best wait?

No. It describes network activity, not whether the UI contains the content you need. Prefer a locator or assertion tied to the output.

How long should I wait after JavaScript finishes?

There is no universal delay. Wait for the target content, then add a short bounded delay only for a known font, image, or transition stabilization issue.

Should I capture the DOM or pixels?

Use DOM or text for semantic checks and archives. Use screenshots for visual output, including canvas. Use PDFs for printable records and validate print CSS independently.

Why is a full-page screenshot missing lower images?

Those images may be lazy-loaded or the page may virtualize content. Scroll through the page or trigger its loading behavior before capture.

When should I use Chrome Headless instead of Playwright?

Chrome Headless is convenient for bounded one-off captures. Playwright is better when you need selectors, assertions, authentication, interactions, multiple formats, and reusable orchestration.

Can an API replace browser automation?

For routine URL-to-image or URL-to-PDF work, yes. ScreenshotNeo handles browser setup, readiness options, consent cleanup, caching, bulk jobs, and signed delivery while retaining verdict and billing headers.