ScreenshotNeo

BlogHow-to

How to Block Ads and Analytics Requests in Puppeteer Screenshots

Use Puppeteer request interception to block known ad and analytics endpoints while allowing page resources through, then capture when the content you need is ready.

By the ScreenshotNeo team4 October 20269 min read

To block ads and analytics requests in a Puppeteer screenshot, enable request interception, match known ad or analytics URLs in a request handler, abort matches, and continue every other unresolved request. Then wait for the page content your screenshot needs and call page.screenshot().

Keep the filter specific. A resource type such as script or image describes what the browser is loading; it does not tell you whether it is an ad. Blocking all scripts or images can remove legitimate page content or stop the site from rendering. Puppeteer also warns that each intercepted request stalls until it is continued, answered, aborted, or served from cache. See the official request interception guide and Page.setRequestInterception reference.

1. Install Puppeteer and choose the endpoints to block

Use a current Puppeteer installation in a Node.js project. The puppeteer package downloads a compatible browser during installation; if your environment manages its own Chrome, use the appropriate Puppeteer setup for that environment.

npm install puppeteer

There is no universal ad or analytics URL list in Puppeteer. Create rules for endpoints you have identified on the sites you capture. Match hostnames rather than loose substrings where possible, so an unrelated URL containing a familiar word is less likely to be blocked.

const blockedHosts = new Set([
  'www.googletagmanager.com',
  'www.google-analytics.com',
  'doubleclick.net',
]);

function shouldBlockRequest(rawUrl) {
  let host;
  try {
    host = new URL(rawUrl).hostname.toLowerCase();
  } catch {
    return false;
  }

  return [...blockedHosts].some(
    blockedHost => host === blockedHost || host.endsWith(`.${blockedHost}`),
  );
}

The hostnames above illustrate the shape of a project-specific filter; they are not a complete tracker list or a guarantee that every request to those hosts is safe to block. Inspect the target site’s requests and verify that blocking a chosen endpoint does not remove content or behavior required in the capture.

2. Capture a page with request interception

Save this as screenshot.mjs. It reads the target URL from the command line, installs the interception handler before navigation, waits for a page-specific selector, and saves a PNG. Replace the example host rules and selector with the ones appropriate for your target.

import puppeteer from 'puppeteer';

const targetUrl = process.argv[2] ?? 'https://example.com';
const outputPath = process.argv[3] ?? 'page.png';

const blockedHosts = new Set([
  'www.googletagmanager.com',
  'www.google-analytics.com',
  'doubleclick.net',
]);

function shouldBlockRequest(rawUrl) {
  let host;
  try {
    host = new URL(rawUrl).hostname.toLowerCase();
  } catch {
    return false;
  }

  return [...blockedHosts].some(
    blockedHost => host === blockedHost || host.endsWith(`.${blockedHost}`),
  );
}

const browser = await puppeteer.launch({ headless: true });

try {
  const page = await browser.newPage();
  await page.setViewport({ width: 1440, height: 1000, deviceScaleFactor: 1 });

  await page.setRequestInterception(true);
  page.on('request', request => {
    // Another listener or package may already have resolved this request.
    if (request.isInterceptResolutionHandled()) return;

    if (shouldBlockRequest(request.url())) {
      void request.abort().catch(error => {
        console.error('Could not abort request:', request.url(), error.message);
      });
    } else {
      void request.continue().catch(error => {
        console.error('Could not continue request:', request.url(), error.message);
      });
    }
  });

  const response = await page.goto(targetUrl, {
    waitUntil: 'domcontentloaded',
    timeout: 45000,
  });

  if (!response) {
    throw new Error('Navigation did not produce a main-resource response.');
  }
  if (!response.ok()) {
    throw new Error(`Navigation returned HTTP ${response.status()}.`);
  }

  // Replace this selector with content that must be present in your screenshot.
  await page.waitForSelector('body', { timeout: 15000 });

  await page.screenshot({ path: outputPath, fullPage: true });
  console.log(`Saved ${outputPath}`);
} finally {
  await browser.close();
}

Run it with a URL and optional output path:

node screenshot.mjs https://example.com example.png

The event handler deliberately makes its decision and resolves each request synchronously. Puppeteer documents that with multiple interception handlers, another handler may resolve a request first; check isInterceptResolutionHandled() immediately before resolving it. If you add asynchronous work to a handler, check again after each await and before calling abort(), continue(), or respond(). Read the details in the multiple-handler section of Puppeteer’s guide.

3. Tune matching without breaking the page

Approach When it fits Trade-off
Match known hostnames You know which third-party endpoints serve advertising or analytics on the target sites. More selective, but rules need maintenance as sites change providers or endpoints.
Match URL paths or query patterns A host serves both required content and known tracking endpoints, and the paths are distinguishable. More precise when the pattern is stable; brittle if the provider changes its URLs.
Block a resource type You intentionally want to omit that whole class of requests, such as images in a text-only capture. Easy to express, but can block legitimate scripts, stylesheets, fonts, or images. Resource type is not an ad classifier.

Puppeteer exposes both request.url() and request.resourceType() on the request. Use the URL and host to identify purpose; use the resource type only when you actually want to filter by browser-perceived category. See the HTTPRequest API.

For path-based matching, extend shouldBlockRequest carefully. For example, check the parsed hostname and pathname together rather than using a broad substring over the entire URL. Keep a record of which sites a rule is intended for, and inspect failures when a site changes.

Avoid installing duplicate request handlers that all resolve requests independently. If other libraries register interception handlers, follow Puppeteer’s resolution guidance. Cooperative interception can use numeric priorities, but it only behaves cooperatively when all handlers provide priorities; a legacy handler can resolve immediately. For a single-purpose script, one small handler with the handled-state guard is easier to reason about.

4. Wait for the right screenshot moment

waitUntil: 'domcontentloaded' means the initial HTML has been parsed; it does not mean a client-rendered application or its images are ready. Pick a readiness signal that corresponds to what the screenshot must show.

  • Known content: wait for a selector that marks the content as rendered, such as a report heading or product grid.
  • Known state: wait for a site-specific condition with page.waitForFunction(), such as a loading indicator disappearing.
  • Fixed delay: use page.waitForTimeout() only when the page offers no better signal; a delay can be too short on a slow run and waste time on a fast one.
  • Network idle: navigation with waitUntil: 'networkidle2' can be useful, but persistent connections or background activity can prevent it from matching visual readiness.

Puppeteer’s screenshots guide shows page.screenshot() and an example using networkidle2. Treat that as an option, not a universal guarantee. The most reliable choice is usually a page-specific selector or state.

// Wait for a meaningful page-specific element before capture.
await page.waitForSelector('[data-testid="report-ready"]', { timeout: 15000 });
await page.screenshot({ path: 'report.png', fullPage: true });

For a single element, wait for it and call its screenshot method:

const card = await page.waitForSelector('.summary-card', { timeout: 15000 });
if (!card) throw new Error('Summary card was not found.');
await card.screenshot({ path: 'summary-card.png' });

5. Screenshot and interception options

Need Puppeteer option or method Notes
Capture the full document page.screenshot({ fullPage: true }) Captures beyond the current viewport.
Capture a region page.screenshot({ clip: { x, y, width, height } }) Coordinates are page screenshot coordinates; ensure the clip is within the rendered content.
Capture an element elementHandle.screenshot({ path }) Puppeteer’s screenshots guide documents this method and says it tries to scroll a hidden element into view.
Choose image format type: 'png', 'jpeg', or 'webp' where supported by the installed version Quality applies to lossy formats, not PNG. The screenshot type can also be inferred from the file extension.
Transparent background omitBackground: true Useful for transparent output where the page background permits it.
Retina-like rendering page.setViewport({ width, height, deviceScaleFactor }) Higher device scale factors increase pixel dimensions and memory use.
Disable interception page.setRequestInterception(false) Turn it off when filtering is no longer needed; do not leave requests unresolved during teardown.

See Puppeteer’s ScreenshotOptions reference for the options available in the installed version. If your capture service needs to block whole categories such as images or fonts, verify the resulting layout: removing those resources can change page geometry or cause scripts to fail.

6. Common errors and fixes

Symptom Likely cause Fix
Navigation hangs after interception is enabled A request handler returned without resolving an unhandled request. Ensure every unresolved request reaches continue(), abort(), or respond(). Keep the handler small and log its decision while diagnosing.
Request is already handled! Another listener or package resolved the same request first. Check request.isInterceptResolutionHandled() immediately before resolving. If the handler awaits anything, check again afterward.
The screenshot is missing page content A rule blocked an endpoint the page needs, or the screenshot ran before client rendering completed. Temporarily log matched URLs, remove the suspect rule, and wait for a page-specific selector or state before capturing.
Ads still appear The site uses a different host, a first-party endpoint, or content already present in the document. Inspect requests from that page, add a narrow rule for the confirmed endpoint, and determine whether the ad is network-loaded or already rendered.
waitForSelector times out The selector is wrong, content is gated, the page failed to render, or a blocked request supplies the content. Confirm the selector in the rendered DOM, inspect navigation and console errors, and review the blocked-request log.
Navigation timeout The page did not reach the chosen navigation condition before the timeout, often because it stays active. Use a less strict navigation condition such as domcontentloaded, then wait for the specific content required. Set a suitable timeout for your workload.
Correct page, wrong dimensions or clipped content The viewport or screenshot options do not match the intended capture. Set the viewport before navigation when responsive layout matters, and check fullPage, clip, and device scale.

7. Performance, reliability, and cost considerations

  • Interception overhead: every request passes through your handler and remains paused until resolved. Keep matching synchronous and inexpensive; avoid network lookups or slow asynchronous rule checks inside the handler.
  • Filter maintenance: endpoint rules can become stale as a site changes vendors or moves requests to first-party hosts. Review blocked URLs when a visual regression appears.
  • Capture reliability: a successful navigation response does not prove the screenshot content is ready. Check the response status, wait for the content required, and use explicit timeouts so failures surface clearly.
  • Memory and output size: full-page captures and high device scale factors produce larger images and require more browser memory. Use an element or clip capture when the full page is unnecessary.
  • Cost: self-hosted Puppeteer has no per-screenshot API price in this method, but your browser runtime, compute, storage, and maintenance have costs. Choose the workflow based on the volume and operational work you can support; there is no universal speed or savings figure.

8. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. Its request options include blocking ads, trackers, requests, or resource types, alongside capture readiness controls. See the ScreenshotNeo API documentation for the available parameters. For a simple capture, send one GET request:

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

In Node.js environments without Bun, save the response body with Node’s filesystem API:

import { writeFile } from 'node:fs/promises';

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
  • Cookie and consent banners are accepted like a visitor and more than 60 known consent platforms, newsletter popups, and chat widgets are removed before the shot; each step can be turned off.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Response headers say which page verdict applied and whether the capture was billed.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
  • The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

9. FAQ

Can Puppeteer block ads without an extension?

Yes. Request interception lets your script abort matching requests without installing a browser extension. The rules must identify requests that the target page actually makes.

Can I block analytics but keep the site functional?

Often, if the analytics endpoints are separate from the resources the page needs. Confirm endpoint purpose and inspect the resulting page; hostname alone does not prove a request is nonessential.

Does blocking analytics remove tracking already in the HTML?

Request interception controls network requests. It does not erase markup or scripts already delivered in the main document, and it does not retroactively undo requests completed before interception was enabled.

Should I block every script to remove ads?

No. Sites commonly rely on JavaScript to render navigation, content, and layout. Block identified endpoints and verify the capture rather than disabling an entire resource category by default.

Sources