ScreenshotNeo

BlogHow-to

How to Get HTML from a NodeList with Puppeteer

Use Puppeteer’s $$eval to turn every matched element into an outerHTML string, with complete patterns, edge cases, and fixes.

By the ScreenshotNeo team1 October 20265 min read

Use page.$$eval() and map each matched element to its outerHTML:

const htmlByElement = await page.$$eval('.item', elements =>
  elements.map(element => element.outerHTML)
);

console.log(htmlByElement);

$$eval finds every element matching the selector, runs the callback in the page, and serializes its return value back to Node.js. Because outerHTML includes the selected element itself, the result is an array of complete HTML strings. See the Puppeteer $$eval API reference.

1. Complete runnable example

Install Puppeteer, open a page, wait for the elements you need, and extract the strings:

npm install puppeteer
const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch({headless: true});
  try {
    const page = await browser.newPage();
    await page.goto('https://example.com', {waitUntil: 'domcontentloaded'});

    const htmlByElement = await page.$$eval('p', elements =>
      elements.map(element => element.outerHTML)
    );

    for (const [index, html] of htmlByElement.entries()) {
      console.log(`Element ${index}:\n${html}`);
    }
  } finally {
    await browser.close();
  }
})();

Replace p with your selector. The callback must return serializable data; an array of strings is safe.

2. NodeList versus Puppeteer selectors

Inside browser code, a NodeList from document.querySelectorAll() is converted before mapping:

const html = await page.evaluate(() => {
  const nodeList = document.querySelectorAll('.item');
  return Array.from(nodeList, node => node.outerHTML);
});

For a selector supplied by Node.js, page.$$eval(selector, callback) already provides the matching array, so a second querySelectorAll() is normally unnecessary.

3. Choose the HTML scope you need

Need API Result
HTML for every match page.$$eval() Map all matches to outerHTML.
Element handles page.$$() Array of ElementHandle objects; returns [] when none match. See the API reference.
HTML for first match page.$eval() One value; throws when no match. See the API reference.
Element contents only innerHTML Use page.$eval(selector, element => element.innerHTML).
Whole document page.content() Full markup including the DOCTYPE. See the API reference.
const first = await page.$eval('.item', element => element.outerHTML);
const contents = await page.$eval('.item', element => element.innerHTML);
const documentHtml = await page.content();

4. Useful extraction patterns

Keep an index

const items = await page.$$eval('.item', elements =>
  elements.map((element, index) => ({index, html: element.outerHTML}))
);

Return attributes and HTML

const records = await page.$$eval('[data-id]', elements =>
  elements.map(element => ({
    id: element.getAttribute('data-id'),
    html: element.outerHTML
  }))
);

Normalize whitespace

const compact = await page.$$eval('.item', elements =>
  elements.map(element => element.outerHTML.replace(/\s+/g, ' ').trim())
);

Normalize only when whitespace is not meaningful. The result represents the live DOM after scripts have run, not necessarily the original response body.

5. Timing, dynamic pages, and frames

Navigation completion does not guarantee that client-rendered elements exist. Wait for the state you need:

await page.goto('https://example.com/products', {waitUntil: 'networkidle2'});
await page.waitForSelector('.item');
const html = await page.$$eval('.item', nodes => nodes.map(node => node.outerHTML));

For optional content, an empty result is safe:

const html = await page.$$eval('.item', nodes => nodes.map(n => n.outerHTML));

Elements inside an iframe belong to that frame’s document:

const frame = page.frames().find(f => f.url().includes('/embedded/'));
if (!frame) throw new Error('Embedded frame was not found');
await frame.waitForSelector('.item');
const html = await frame.$$eval('.item', nodes => nodes.map(n => n.outerHTML));

Normal selectors do not cross a shadow root:

const html = await page.$eval('my-widget', host =>
  Array.from(host.shadowRoot?.querySelectorAll('.item') ?? [], node => node.outerHTML)
);

6. Page-context rules

Puppeteer serializes the callback and executes it in the browser page context. Node.js variables, imports, and helper functions are not automatically available. Keep logic inside the callback or pass serializable arguments:

const suffix = '<!-- extracted -->';
const html = await page.$$eval('.item', (nodes, suffix) =>
  nodes.map(node => node.outerHTML + suffix), suffix
);

Return strings or plain objects. Do not return element handles when your goal is HTML.

7. Edge cases

  • No matches: $$eval returns []. $eval throws, so use it only when a match is required.
  • Detached nodes: a framework can replace nodes while extraction runs. Wait for a stable state and perform one short evaluation.
  • Large results: thousands of full subtrees increase memory and serialization time. Extract required fields, process batches, or write incrementally.
  • Unsafe HTML: treat returned strings as untrusted. Sanitize before inserting them into another document.
  • SVG and XML: serialization follows browser namespace rules and may differ from source bytes.
  • Closed shadow roots: page JavaScript cannot query a closed shadow root through shadowRoot.

8. Troubleshooting

Symptom Cause Fix
page.$$eval is not a function Wrong object or incompatible Puppeteer setup. Confirm the variable is a Puppeteer Page and check the installed version’s API reference.
Result is [] Selector is wrong, content is not rendered yet, or the element is in another frame or shadow root. Check the selector, wait for it, and evaluate in the correct context.
$eval cannot find an element No match existed at evaluation time. Use waitForSelector for required content or switch to $$eval for optional content.
Client-rendered HTML is missing Extraction ran before rendering finished. Wait for a meaningful selector or application state.
Node helper is undefined Callback runs in the page context. Move logic into the callback or pass arguments.
Browser hangs Navigation, scripts, or selector wait never completes. Set explicit timeouts, log the URL and selector, and close the browser in finally.

9. Performance, reliability, and cost

  • Performance: one $$eval serializes one array and avoids creating a handle for each element. Return the smallest useful value.
  • Reliability: wait on a meaningful DOM condition, isolate iframe work, and close the browser even after failures.
  • Concurrency: reuse a browser process for many pages, create separate pages for independent tasks, and cap concurrency for available CPU and memory.
  • Cost: a self-managed Puppeteer browser requires compute, memory, bandwidth, and maintenance. An API can remove that setup when you need rendered images or PDFs.

10. Or skip the browser setup

If you need a rendered image or PDF rather than HTML strings, ScreenshotNeo provides a single screenshot request. See the ScreenshotNeo API docs.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie and consent banners, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing result. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

11. FAQ

Does outerHTML include the element itself?

Yes. Use innerHTML for only its children.

What does $$eval return when nothing matches?

An empty array.

Can I get the original server response HTML?

No. outerHTML and page.content() serialize the current DOM.

When should I use page.$$?

Use it when later Puppeteer operations require element handles; use $$eval for serializable results.