ScreenshotNeo

BlogHow-to

How to Get a Page’s HTML with Puppeteer

Use Puppeteer’s page.content() for the full current HTML, or extract one or more elements after the page reaches the state you need.

By the ScreenshotNeo team4 October 20267 min read

Use await page.content() to get a page’s current full HTML with Puppeteer, including the DOCTYPE. Navigate to the page first, then wait for any client-rendered content you need before reading it. This returns the browser’s current DOM serialization; it is not necessarily the exact bytes originally sent in the HTTP response.

Get the full page HTML

Install Puppeteer, save this as get-html.mjs, and run it with Node.js. The finally block closes the browser even if navigation or extraction fails.

import puppeteer from 'puppeteer';

const url = process.argv[2] ?? 'https://example.com';
const browser = await puppeteer.launch();

try {
  const page = await browser.newPage();
  await page.goto(url, { waitUntil: 'domcontentloaded' });

  const html = await page.content();
  console.log(html);
} finally {
  await browser.close();
}

Install the package with npm install puppeteer, then run node get-html.mjs https://example.com. Puppeteer’s Page.content() API reference defines the result as the full HTML contents, including the DOCTYPE.

Choose the right extraction method

Need Method Behavior
Full document await page.content() Returns the current page HTML, including the DOCTYPE.
Custom DOM value await page.evaluate(() => ...) Runs your function in the page’s JavaScript context.
One element await page.$eval(selector, fn) Runs against the first match; throws if nothing matches.
All matching elements await page.$$eval(selector, fn) Runs against all matches and returns your mapped result.

For example, retrieve the full document element through page-context JavaScript:

const html = await page.evaluate(() => document.documentElement.outerHTML);

Or extract one element’s markup:

const mainHtml = await page.$eval('main', element => element.outerHTML);

To collect several matching elements:

const cards = await page.$$eval('.card', elements =>
  elements.map(element => element.outerHTML)
);

Use outerHTML when you need the selected element itself and its descendants. Use innerHTML when you only need its children. The whole page form, page.content(), is usually the simplest choice for a complete document.

Wait for the content you need

page.goto() completing does not guarantee that every application has finished its client-side rendering. Choose a readiness condition that matches the content you will extract. For a known element, a locator waits for it to be present before the action:

await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
await page.locator('main article').wait();
const articleHtml = await page.$eval('main article', element => element.outerHTML);

You can also wait for a page condition directly:

await page.waitForFunction(() => {
  const article = document.querySelector('main article');
  return article && article.textContent.trim().length > 0;
});
const articleHtml = await page.$eval('main article', element => element.outerHTML);

Prefer a meaningful condition, such as a specific selector or text becoming available, over an arbitrary delay. A fixed delay can be useful when a known third-party widget has no observable ready state, but it makes capture slower and remains sensitive to variable network and server response times. Avoid waiting for network idle as a universal rule: analytics, streaming requests, and long polling can keep a page active even after the target content is ready.

Handle missing elements safely

$eval() throws when the selector has no match. If absence is valid for your workflow, query first and return a deliberate fallback:

const element = await page.$('main');
const mainHtml = element
  ? await element.evaluate(node => node.outerHTML)
  : null;

if (mainHtml === null) {
  console.error('The page has no main element');
}

For several matches, $$eval() returns an empty array when there are no matches, which is convenient when zero results is acceptable. Use a selector that is stable for the site and page state you expect; a selector for a transient loading shell can succeed while the real content is still absent.

DOM HTML versus the original HTTP response

Puppeteer’s page.content() describes the current document. Page scripts may add, remove, or change elements, so the result can differ from the original response source. The same distinction applies to page.evaluate(), which executes in the browser page context. If you need the original network response body exactly as received, inspect and save the navigation response separately; DOM serialization should not be treated as byte-for-byte source recovery.

Practical options and edge cases

  • Navigation readiness: choose a waitUntil state such as domcontentloaded or load based on what the page needs. These indicate navigation milestones, not proof that all application data is rendered.
  • Frames: page.content() reads the page’s main document. Content inside an iframe belongs to that frame’s document; access the frame and evaluate or select within it separately.
  • Shadow DOM: ordinary document selectors do not cross shadow-root boundaries. Query the relevant shadow root in page context when you need its contents.
  • Large documents: returning the full HTML transfers it from the browser process to Node.js and can consume substantial memory. Extract only the needed elements or fields when possible.
  • Forms and live properties: serialized markup may not represent every live property exactly as the browser currently holds it. Read values such as an input’s value directly in page context if that current value is what you need.
  • Authentication and access: pages requiring a session may need navigation and login steps before extraction. Handle credentials and saved HTML as sensitive data.

cURL and Python alternatives

These examples fetch a URL over HTTP without running a browser. They return the server response body, which may not include content inserted later by client-side JavaScript. Use Puppeteer when the rendered DOM is the target.

cURL

curl -L --fail --show-error 'https://example.com' -o response.html

Python

import requests

response = requests.get('https://example.com', timeout=30)
response.raise_for_status()
with open('response.html', 'w', encoding=response.encoding or 'utf-8') as output:
    output.write(response.text)

Use these for a simple HTTP response inspection, not as a substitute for JavaScript rendering. Sites can serve different markup based on cookies, headers, or authentication.

Or skip the browser setup

If your goal is a visual capture rather than HTML markup, ScreenshotNeo returns a screenshot or PDF through one GET request. Its API documentation lists the available capture options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server lets AI agents use screenshot, page information, and PDF capture tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for free.

Common errors and fixes

Symptom Likely cause Fix
$eval() reports no element The selector is wrong, the content is not rendered yet, or it is inside a frame or shadow root. Confirm the selector, wait for the target state, and query the correct frame or shadow root.
HTML is missing content seen in the browser Extraction ran before client rendering finished, or content is in an iframe. Wait for a meaningful selector or condition and inspect the relevant frame.
Navigation times out The site is slow, blocked, or has requests that do not settle. Use a suitable navigation milestone, set an intentional timeout, and wait separately for the target content.
Output differs from View Source You captured the current DOM, which scripts may have changed. Decide whether you need rendered DOM or the original response body and capture the correct representation.
Node process remains open The browser was not closed after an error. Put browser.close() in a finally block.
Extraction is slow or memory-heavy The page is large, images and scripts are still loading, or too much markup is returned. Wait only for the needed state and extract a smaller subtree or structured values.

Performance, reliability, and cost

Launching a browser has more CPU, memory, and startup cost than a direct HTTP request. Reuse a browser process for multiple pages in a controlled worker, close each page when finished, and limit concurrency according to available memory. A fresh browser context per job helps isolate cookies and storage, while a reused context is faster but can carry state between captures.

Set explicit navigation and selector timeouts, close browser resources in cleanup paths, and record the URL and failure stage when extraction fails. Retrying can help with transient network errors, but use a bounded retry count and avoid retrying deterministic selector errors without changing the readiness condition. Puppeteer itself does not charge per capture; the compute and hosting costs are yours. If you only need an image or PDF and do not want to operate a browser, ScreenshotNeo has a free tier and published paid tiers described above.

FAQ

Does page.content() include the DOCTYPE?

Yes. Puppeteer documents the full HTML contents, including the DOCTYPE.

Can I get only the HTML inside the body?

Yes: use await page.evaluate(() => document.body.innerHTML).

Does Puppeteer return the same source as View Source?

Not necessarily. Puppeteer reads the current browser document, which client-side scripts may have changed.

Can Puppeteer extract HTML from a page that requires JavaScript?

Yes. Navigate with Puppeteer, wait for the content you need to appear, then read the current DOM.

References