Puppeteer Page API: A Guide to Browser Page Automation
Learn how Puppeteer’s Page API handles navigation, interaction, waiting, evaluation, screenshots, and PDFs, with runnable Node.js examples and troubleshooting.
Puppeteer’s Page class is the automation surface for one browser tab. Use it to navigate, find and interact with elements, run JavaScript in the page, wait for meaningful conditions, and capture screenshots or PDFs. A browser can have multiple Page instances. The examples below target Puppeteer 25.12.0; check the matching Page API reference and interaction guide when using another version.
1. Install Puppeteer and open a page
Puppeteer’s full package downloads a compatible Chrome for Testing browser during installation. If you already manage a browser executable, see Puppeteer’s installation documentation for the browser-management options. The following CommonJS script opens a page, prints its title, and closes the browser even if an operation fails.
npm install puppeteer@25.12.0
const puppeteer = require('puppeteer');
(async () => {
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
const response = await page.goto('https://example.com', {
waitUntil: 'domcontentloaded',
timeout: 30_000,
});
console.log({
status: response ? response.status() : null,
title: await page.title(),
url: page.url(),
});
} finally {
await browser.close();
}
})();
Save as page.js and run node page.js. browser.newPage() creates a page in a new browser context; use browserContext.newPage() when you need several pages to share that context’s session state. A Page is for one tab. Browser-wide launch settings and context-wide storage or permissions belong to their corresponding browser APIs.
2. Navigate and understand readiness
The Page API includes goto(), goBack(), goForward(), and reload(). Navigation methods return a response when a response exists; for example, a navigation that changes the URL through a client-side route may not have a new main-resource response. Check the response before reading its status.
const response = await page.goto('https://example.com/products', {
waitUntil: 'domcontentloaded',
timeout: 30_000,
});
if (response && !response.ok()) {
throw new Error(`Navigation returned HTTP ${response.status()}`);
}
await page.goBack({ waitUntil: 'domcontentloaded' });
await page.goForward({ waitUntil: 'domcontentloaded' });
await page.reload({ waitUntil: 'domcontentloaded' });
waitUntil accepts lifecycle events such as load, domcontentloaded, networkidle0, and networkidle2; navigation defaults to load. A lifecycle event is not necessarily the same as application readiness: a page may fetch data or render content afterward. Wait for the element, response, or page state that means your task can proceed. Network-idle waits can be a poor fit for applications that keep connections open or continuously poll.
3. Find elements and interact
For new interaction code, consider Puppeteer’s Locator API. Locators express an interaction and provide synchronization behavior. Use lower-level selector methods or ElementHandle when you need a capability a Locator does not expose. Check the current interaction guide for the Locator methods and supported selector syntax in your installed version.
// Locator interaction: waits for the target to be available for the action.
await page.locator('input[name="email"]').fill('dev@example.com');
await page.locator('button[type="submit"]').click();
// Selector APIs when you need direct DOM data.
const heading = await page.$eval('h1', element => element.textContent.trim());
const links = await page.$$eval('a', elements =>
elements.map(element => ({ text: element.textContent.trim(), href: element.href }))
);
console.log({ heading, links });
page.$(selector) returns the first matching element handle or null; page.$$(selector) returns handles for all matches. $eval() finds the first match and passes it to the callback, throwing if none exists. $$eval() passes all matches to the callback, including an empty array when none match. Keep handle lifetimes short: handles point to objects in a particular document and can become stale after navigation or DOM replacement.
4. Wait for the condition your task needs
waitForSelector() resolves immediately if the selector already matches. It can wait for visibility or hidden state, and it works across navigations. Its documented default timeout is 30,000 ms; configure a page default with page.setDefaultTimeout() or override a specific wait.
await page.setDefaultTimeout(15_000);
// Wait until the element exists and is visible, then click it.
await page.waitForSelector('[data-testid="continue"]', {
visible: true,
timeout: 10_000,
});
await page.click('[data-testid="continue"]');
// Wait for a page-side condition.
await page.waitForFunction(() => {
const status = document.querySelector('[data-testid="status"]');
return status && status.textContent.includes('Complete');
}, { timeout: 20_000 });
Other options include waiting for a request, response, or network idle. Select the wait based on the observable outcome: a selector for an element, waitForFunction() for a page condition, or a request/response wait for network behavior. Avoid using a fixed sleep as a substitute for readiness when a specific condition can be observed.
If a click can trigger navigation, start the navigation wait and click together. Starting the wait afterward can miss a fast navigation.
const [response] = await Promise.all([
page.waitForNavigation({ waitUntil: 'domcontentloaded', timeout: 30_000 }),
page.click('a.some-link'),
]);
console.log('Navigation response:', response ? response.status() : 'none');
This pattern is appropriate only when the action is expected to navigate. For a single-page application route change, wait for the resulting URL or page content instead. Navigation can also result from a form submission or script; synchronize with the actual action you perform.
5. Run JavaScript in the page
page.evaluate() executes a function in the browser page’s JavaScript context. Node.js variables are not automatically available there, so pass values as arguments. If the function returns a Promise, Puppeteer waits for it and returns the resolved, serializable value. Use evaluateHandle() when you need a handle to an in-page object rather than a serialized result.
const selector = 'main h1';
const pageData = await page.evaluate((selector) => {
const heading = document.querySelector(selector);
return {
title: document.title,
heading: heading ? heading.textContent.trim() : null,
canonical: document.querySelector('link[rel="canonical"]')?.href ?? null,
};
}, selector);
console.log(pageData);
Return plain data such as strings, numbers, arrays, and objects for predictable serialization. DOM nodes and many browser objects should be handled with evaluateHandle() or used inside the evaluation callback rather than returned as ordinary data.
6. Capture a screenshot or PDF
page.screenshot() returns image data by default. Choose a format and path explicitly when the output is an artifact. Full-page capture includes the full scrollable page; otherwise the screenshot covers the current viewport. page.pdf() generates a PDF using print CSS media by default. To use screen media for PDF rendering, call page.emulateMediaType('screen') first.
// Viewport screenshot
await page.screenshot({ path: 'page.png', type: 'png' });
// Full-page screenshot
await page.screenshot({ path: 'page-full.png', fullPage: true });
// PDF with screen styles
await page.emulateMediaType('screen');
await page.pdf({ path: 'page.pdf', format: 'A4', printBackground: true });
Capture is a representation of the rendered page, not verification that the page’s underlying data is correct. Before capturing dynamic content, wait for its meaningful ready condition. If the page uses lazy-loaded images, scroll or otherwise trigger the content before capture and verify that the needed content has rendered.
7. Set page timeouts and configure the viewport
Timeouts should match the slowest legitimate operation in your environment. The documented selector wait default is 30 seconds; navigation also documents a 30-second default. Set an explicit value for a specific operation when that wait has a different budget, or use Page default timeout settings for consistent selector and general waits. Navigation timeouts can be configured separately.
await page.setViewport({ width: 1440, height: 900, deviceScaleFactor: 1 });
page.setDefaultTimeout(15_000);
page.setDefaultNavigationTimeout(45_000);
await page.goto('https://example.com', {
waitUntil: 'domcontentloaded',
timeout: 45_000,
});
Set the viewport before navigation when responsive layout or viewport-triggered behavior matters. A new page’s viewport and device emulation affect layout and can affect what is visible in a capture. Configure only the dimensions and emulation your task requires.
8. Handle errors and make automation reliable
Automation commonly fails because a page did not reach the assumed state, a selector changed, navigation raced with an action, or the browser process could not start. Catch errors at the job boundary, record the URL and operation, and always close the browser.
| Symptom | Likely cause | Fix |
|---|---|---|
TimeoutError from waitForSelector() |
The selector is wrong, the element never appears, it is hidden when visibility was required, or the page is delayed. | Confirm the selector against the rendered page, wait for the actual state, and set an intentional timeout. Do not just raise every timeout indefinitely. |
Execution context was destroyed |
Navigation replaced the document while evaluation or a handle operation was running. | Coordinate the action with navigation, then query or evaluate on the new page after navigation completes. |
| Navigation wait times out after a click | The click did not navigate, or the page used a client-side route change. | Wait for the resulting selector or URL condition if the action is an SPA transition; use waitForNavigation() only when navigation is expected. |
$eval() throws |
No element matched the selector. | Wait for the element first or use page.$() and handle a null result. |
| Screenshot is missing content | Capture happened before rendering or lazy content completed. | Wait for a meaningful page condition and trigger lazy content loading before capture. |
| Chrome fails to launch | The executable is missing, installation is incomplete, or the runtime lacks required system libraries. | Install the package’s supported browser, check the launch error and deployment environment, and configure an executable path only when managing the browser yourself. |
| Page hangs on network idle | The page maintains long-lived connections or continual requests. | Wait for the specific selector, response, or page condition instead of network idle. |
For production jobs, use bounded timeouts, close pages and browsers in finally blocks, and make retries conditional. Retrying a deterministic selector error wastes time; retry a transient navigation or infrastructure failure only when repeating the operation is safe.
9. Performance, reliability, and cost
Launching a browser is usually a larger unit of work than reusing an already-running browser for another page. For repeated jobs, manage browser and context lifetimes deliberately, isolate user sessions in separate contexts, and close each page when finished. Keep concurrent work within the memory and CPU capacity available to the process; rendering cost varies with page complexity, image size, scripts, and viewport.
Choose the lightest wait that proves readiness. Waiting for the full load event or network idle can extend a job when the task needs only a particular element. Conversely, capturing too early creates unreliable artifacts and repeat work. Block resources only when the task does not need them, since scripts, stylesheets, or images can change layout and application behavior.
Puppeteer itself is an open-source browser automation library; infrastructure costs depend on where Chrome runs and how much CPU, memory, storage, and concurrency the workload consumes. Budget for browser processes and generated artifacts, and clean up temporary files. No universal per-screenshot runtime or cost follows from the Page API.
10. Or skip the browser setup
For a one-off website screenshot or a screenshot pipeline where you do not need to control a browser session, ScreenshotNeo provides a website screenshot API and MCP server. Its options include full-page shots, element capture, PDF output, custom wait conditions, and device presets. See the API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
require('node:fs').writeFileSync('shot.webp', bytes);
- Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers report the page verdict and billing status.
- An MCP server offers
take_screenshot,get_page_info, andcapture_pdffor AI agents and MCP clients. - The Free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 screenshots.
Sign up free for 1,000 screenshots a month, with no card required.
11. Frequently asked questions
Does each Puppeteer Page create a separate browser?
No. A Page represents one tab within a browser. One browser can manage multiple pages, and browser contexts control session separation.
Should I use Locator or waitForSelector()?
Use a Locator for an interaction when its built-in synchronization fits the job. Use waitForSelector() or an element handle when you need lower-level control or a capability the Locator does not provide.
Can evaluate() access variables from my Node.js script?
Not through lexical scope. Pass values as arguments to the function, and return serializable data when you want an ordinary Node.js result.
Why does a PDF look different from my screenshot?
PDF generation uses print media by default, which can apply different CSS. Emulate screen media before generating the PDF if screen styles are required.
When is Puppeteer the right choice?
Use it when you need browser-level control over a tab, interactions, page JavaScript, or a custom capture workflow. For a direct website screenshot without managing Chrome, a screenshot API can reduce setup.


