Screenshot a Webpage with Node.js Puppeteer and Block Ads and Trackers
Capture a webpage with Puppeteer, save a screenshot, and block selected ad or tracker requests with explicit, maintainable rules.
Use Puppeteer’s request interception to apply explicit rules to outgoing requests, then capture the page with page.screenshot(). Every intercepted request must be continued, aborted, or answered; a request left unresolved can stall page loading. Puppeteer gives you request URLs and resource types, but it does not provide a built-in, authoritative list of ads and trackers. The result depends on the rules you choose and the site you capture.
1. Install Puppeteer and save a screenshot
This runnable ES module example blocks requests whose host matches an example list and captures the page as a full-page PNG. Replace the sample domains with rules you maintain for your use case. They are examples of matching mechanics, not a verified tracker list.
npm init -y
npm install puppeteer
Save this as screenshot.mjs:
import puppeteer from 'puppeteer';
import { URL } from 'node:url';
const targetUrl = process.argv[2] ?? 'https://example.com';
const blockedHosts = new Set([
'ads.example.net',
'tracker.example.org',
]);
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage({
viewport: { width: 1440, height: 1000 },
deviceScaleFactor: 1,
});
await page.setRequestInterception(true);
page.on('request', request => {
// A single handler owns the decision for each request.
if (request.isInterceptResolutionHandled()) return;
let shouldBlock = false;
try {
const host = new URL(request.url()).hostname;
shouldBlock = blockedHosts.has(host) ||
[...blockedHosts].some(domain => host.endsWith(`.${domain}`));
} catch {
// If a request URL cannot be parsed, allow it through.
}
if (shouldBlock) {
void request.abort().catch(() => {});
} else {
void request.continue().catch(() => {});
}
});
await page.goto(targetUrl, {
waitUntil: 'networkidle2',
timeout: 45_000,
});
// Optional: wait for a meaningful page-specific element instead of
// assuming network quiet means the visual content is ready.
// await page.waitForSelector('main article', { timeout: 10_000 });
await page.screenshot({ path: 'page.png', fullPage: true });
console.log('Saved page.png');
} finally {
await browser.close();
}
Run it with:
node screenshot.mjs https://example.com
The domain check includes subdomains, so a rule for example.net also matches ads.example.net. Keep rules narrow: a domain can host both useful content and tracking scripts. If you prefer a path or URL pattern, make the condition explicit and review it against requests from the pages you capture.
2. Choose what to block
There are three common rule styles. Puppeteer exposes the request information and resolution methods; deciding whether a request is an ad or tracker is your policy.
| Rule style | Example signal | Trade-off |
|---|---|---|
| Exact host or domain | new URL(request.url()).hostname |
Readable and relatively narrow, but shared domains may serve needed content too. |
| URL path or substring | request.url().includes('/analytics/') |
Useful for a known endpoint; broad substrings can match unrelated URLs. |
| Resource type | request.resourceType() |
Easy to apply, but types such as script, image, or stylesheet also contain legitimate page resources. |
For example, a resource-type rule can be added inside the handler:
const type = request.resourceType();
const blockAllImages = false;
if (blockAllImages && type === 'image') {
void request.abort().catch(() => {});
} else {
void request.continue().catch(() => {});
}
This blocks every image, not just ad images. It can remove product photos, charts, logos, or other content and change the screenshot. Avoid blanket rules for scripts, images, or stylesheets unless that change is intentional.
Keep request handling reliable
- Enable interception before navigation and before calling
abort(),continue(), orrespond(). - Resolve every intercepted request. A pass-through request still needs
continue(). - If other code registers interception handlers, check
request.isInterceptResolutionHandled()before resolving the request. - Puppeteer’s cooperative interception mode uses numeric priorities when every handler supplies one; the highest priority wins. Without a priority, a handler uses legacy immediate resolution. Avoid mixing handlers casually.
page.authenticate()enables interception internally. Account for that if authentication is added elsewhere in your code.
3. Pick screenshot scope and format
page.screenshot() captures the current viewport by default and produces PNG output. Use fullPage to capture the full document, or clip to capture a defined rectangle. The screenshot guide also documents element screenshots through ElementHandle.screenshot(); it scrolls a hidden element into view by default.
// Viewport PNG (default scope and format)
await page.screenshot({ path: 'viewport.png' });
// Entire document
await page.screenshot({ path: 'full-page.png', fullPage: true });
// A rectangle in page coordinates
await page.screenshot({
path: 'region.png',
clip: { x: 100, y: 120, width: 900, height: 600 },
});
// A specific element
const card = await page.waitForSelector('.product-card');
if (!card) throw new Error('Product card was not found');
await card.screenshot({ path: 'product-card.png' });
Other relevant screenshot options include type (png, jpeg, or webp where supported by the installed Puppeteer/browser), quality for formats that support it, and omitBackground for a transparent background where supported. Quality does not apply to PNG. Check the API reference for the Puppeteer version installed in your project before relying on version-specific options.
4. Wait for the page you mean to capture
waitUntil: 'networkidle2' is a useful navigation condition, not proof that a page is visually complete. Applications can render after navigation, lazy-load content on scroll, or keep connections active. If a screenshot misses content, wait for a selector or application-specific ready condition before capture.
await page.goto(targetUrl, { waitUntil: 'domcontentloaded', timeout: 45_000 });
await page.waitForSelector('main article', { timeout: 15_000 });
await page.screenshot({ path: 'article.png', fullPage: true });
For a full-page capture of content loaded on scroll, scroll in steps before the screenshot so the page can trigger lazy loading. The exact strategy depends on the site; avoid assuming that one short fixed delay works for every page.
5. cURL, Python, and Node.js options
Puppeteer is a Node.js library, so the interception implementation belongs in Node.js. cURL and Python can call a screenshot service, but neither runs Puppeteer’s browser interception handler by itself. These examples use ScreenshotNeo’s API for a one-request screenshot; see the ScreenshotNeo API documentation for supported parameters.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com \
-o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
f.write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));
6. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server from ScreenshotNeo. One GET request returns an image or PDF, and its API supports blocking requests and resource types, along with full-page capture, selectors, device presets, custom CSS and JavaScript, waits, caching, and other capture controls. The examples below save a screenshot; use the docs for the relevant request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, with page verdict and billing information in response headers. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Read the API docs and sign up for 1,000 free screenshots a month, with no card.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Navigation hangs or times out after interception is enabled | A request was neither continued, aborted, nor answered. | Ensure every handler path resolves the request, including the allow path and exceptions. |
| Useful page content is missing | A broad domain, path, or resource-type rule blocked a required request. | Log request URLs and types, narrow the rule, and compare a capture with blocking disabled. |
| Handler reports a request was already handled | Another handler or code path resolved it first. | Check isInterceptResolutionHandled() before acting and coordinate priorities if using cooperative interception. |
| Screenshot is blank or incomplete | The page may not have rendered its main content before capture, or a needed script/style request was blocked. | Wait for a meaningful selector or app-ready signal; temporarily allow suspected blocked resources. |
| Full-page shot omits lazy content | Content may only load after it enters the viewport. | Scroll through the page before capturing and wait for the relevant content to appear. |
| Authentication changes request behavior or slows capture | page.authenticate() turns interception on internally. |
Review all interception handlers and keep their decisions narrow and resolved. |
| Output type or quality is ignored | The chosen option may not apply to that format or installed version; quality does not apply to PNG. | Use a supported format, set quality only for supported formats, and consult the API reference for the installed version. |
8. Performance, reliability, and cost
- Performance: interception adds a decision for each request. Keep handlers synchronous and lightweight, avoid expensive regular expressions or remote lookups per request, and use narrow rules.
- Reliability: use
try/finallyto close the browser, set navigation and selector timeouts, and capture only after a page-specific readiness condition when timing matters. Treat network-idle as a signal rather than a universal guarantee. - Rule maintenance: review blocked and allowed request logs as target sites change. A domain can change purpose, and resource types do not reveal whether a request is advertising-related.
- Cost: local Puppeteer has no per-screenshot API charge, but browser compute, memory, storage, and maintenance have costs in your environment. ScreenshotNeo offers 1,000 shots monthly free, then paid plans from $5 for 3,000; yearly billing gives two months free. Every feature is on every plan.
9. FAQ
Does Puppeteer automatically know which requests are ads?
No. It provides request details and controls. Your rules determine what is blocked.
Can I block trackers without changing the page?
Sometimes, but not universally. A request classified as a script or image may also provide visible page content. Test rules against each target site.
Should I use full-page capture for every screenshot?
No. Use viewport capture for the visible screen, full-page for the document, and a clip or element screenshot for a specific region.
Is network idle always the best wait condition?
No. It can be a useful navigation condition, but a selector or application-specific ready signal is often a better indication that the content you need is present.


