How to Get HTML from a NodeList with Puppeteer
Use Puppeteer’s $$eval to turn every matched element into an outerHTML string, with complete patterns, edge cases, and fixes.
Use page.$$eval() and map each matched element to its outerHTML:
const htmlByElement = await page.$$eval('.item', elements =>
elements.map(element => element.outerHTML)
);
console.log(htmlByElement);
$$eval finds every element matching the selector, runs the callback in the page, and serializes its return value back to Node.js. Because outerHTML includes the selected element itself, the result is an array of complete HTML strings. See the Puppeteer $$eval API reference.
1. Complete runnable example
Install Puppeteer, open a page, wait for the elements you need, and extract the strings:
npm install puppeteer
const puppeteer = require('puppeteer');
(async () => {
const browser = await puppeteer.launch({headless: true});
try {
const page = await browser.newPage();
await page.goto('https://example.com', {waitUntil: 'domcontentloaded'});
const htmlByElement = await page.$$eval('p', elements =>
elements.map(element => element.outerHTML)
);
for (const [index, html] of htmlByElement.entries()) {
console.log(`Element ${index}:\n${html}`);
}
} finally {
await browser.close();
}
})();
Replace p with your selector. The callback must return serializable data; an array of strings is safe.
2. NodeList versus Puppeteer selectors
Inside browser code, a NodeList from document.querySelectorAll() is converted before mapping:
const html = await page.evaluate(() => {
const nodeList = document.querySelectorAll('.item');
return Array.from(nodeList, node => node.outerHTML);
});
For a selector supplied by Node.js, page.$$eval(selector, callback) already provides the matching array, so a second querySelectorAll() is normally unnecessary.
3. Choose the HTML scope you need
| Need | API | Result |
|---|---|---|
| HTML for every match | page.$$eval() |
Map all matches to outerHTML. |
| Element handles | page.$$() |
Array of ElementHandle objects; returns [] when none match. See the API reference. |
| HTML for first match | page.$eval() |
One value; throws when no match. See the API reference. |
| Element contents only | innerHTML |
Use page.$eval(selector, element => element.innerHTML). |
| Whole document | page.content() |
Full markup including the DOCTYPE. See the API reference. |
const first = await page.$eval('.item', element => element.outerHTML);
const contents = await page.$eval('.item', element => element.innerHTML);
const documentHtml = await page.content();
4. Useful extraction patterns
Keep an index
const items = await page.$$eval('.item', elements =>
elements.map((element, index) => ({index, html: element.outerHTML}))
);
Return attributes and HTML
const records = await page.$$eval('[data-id]', elements =>
elements.map(element => ({
id: element.getAttribute('data-id'),
html: element.outerHTML
}))
);
Normalize whitespace
const compact = await page.$$eval('.item', elements =>
elements.map(element => element.outerHTML.replace(/\s+/g, ' ').trim())
);
Normalize only when whitespace is not meaningful. The result represents the live DOM after scripts have run, not necessarily the original response body.
5. Timing, dynamic pages, and frames
Navigation completion does not guarantee that client-rendered elements exist. Wait for the state you need:
await page.goto('https://example.com/products', {waitUntil: 'networkidle2'});
await page.waitForSelector('.item');
const html = await page.$$eval('.item', nodes => nodes.map(node => node.outerHTML));
For optional content, an empty result is safe:
const html = await page.$$eval('.item', nodes => nodes.map(n => n.outerHTML));
Elements inside an iframe belong to that frame’s document:
const frame = page.frames().find(f => f.url().includes('/embedded/'));
if (!frame) throw new Error('Embedded frame was not found');
await frame.waitForSelector('.item');
const html = await frame.$$eval('.item', nodes => nodes.map(n => n.outerHTML));
Normal selectors do not cross a shadow root:
const html = await page.$eval('my-widget', host =>
Array.from(host.shadowRoot?.querySelectorAll('.item') ?? [], node => node.outerHTML)
);
6. Page-context rules
Puppeteer serializes the callback and executes it in the browser page context. Node.js variables, imports, and helper functions are not automatically available. Keep logic inside the callback or pass serializable arguments:
const suffix = '<!-- extracted -->';
const html = await page.$$eval('.item', (nodes, suffix) =>
nodes.map(node => node.outerHTML + suffix), suffix
);
Return strings or plain objects. Do not return element handles when your goal is HTML.
7. Edge cases
- No matches:
$$evalreturns[].$evalthrows, so use it only when a match is required. - Detached nodes: a framework can replace nodes while extraction runs. Wait for a stable state and perform one short evaluation.
- Large results: thousands of full subtrees increase memory and serialization time. Extract required fields, process batches, or write incrementally.
- Unsafe HTML: treat returned strings as untrusted. Sanitize before inserting them into another document.
- SVG and XML: serialization follows browser namespace rules and may differ from source bytes.
- Closed shadow roots: page JavaScript cannot query a closed shadow root through
shadowRoot.
8. Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
page.$$eval is not a function |
Wrong object or incompatible Puppeteer setup. | Confirm the variable is a Puppeteer Page and check the installed version’s API reference. |
Result is [] |
Selector is wrong, content is not rendered yet, or the element is in another frame or shadow root. | Check the selector, wait for it, and evaluate in the correct context. |
$eval cannot find an element |
No match existed at evaluation time. | Use waitForSelector for required content or switch to $$eval for optional content. |
| Client-rendered HTML is missing | Extraction ran before rendering finished. | Wait for a meaningful selector or application state. |
| Node helper is undefined | Callback runs in the page context. | Move logic into the callback or pass arguments. |
| Browser hangs | Navigation, scripts, or selector wait never completes. | Set explicit timeouts, log the URL and selector, and close the browser in finally. |
9. Performance, reliability, and cost
- Performance: one
$$evalserializes one array and avoids creating a handle for each element. Return the smallest useful value. - Reliability: wait on a meaningful DOM condition, isolate iframe work, and close the browser even after failures.
- Concurrency: reuse a browser process for many pages, create separate pages for independent tasks, and cap concurrency for available CPU and memory.
- Cost: a self-managed Puppeteer browser requires compute, memory, bandwidth, and maintenance. An API can remove that setup when you need rendered images or PDFs.
10. Or skip the browser setup
If you need a rendered image or PDF rather than HTML strings, ScreenshotNeo provides a single screenshot request. See the ScreenshotNeo API docs.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie and consent banners, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing result. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
11. FAQ
Does outerHTML include the element itself?
Yes. Use innerHTML for only its children.
What does $$eval return when nothing matches?
An empty array.
Can I get the original server response HTML?
No. outerHTML and page.content() serialize the current DOM.
When should I use page.$$?
Use it when later Puppeteer operations require element handles; use $$eval for serializable results.


