ScreenshotNeo

BlogHow-to

How to Capture Ad Screenshots and Automate Tear Sheets

Capture defensible ad proof, preserve metadata, and automate client-ready tear sheets with browser automation, APIs, validation, and PDF workflows.

By the ScreenshotNeo team29 September 20269 min read

How to Capture Ad Screenshots and Automate Tear Sheets

A digital tear sheet is a dated proof-of-placement record. It shows what ad creative rendered on a publisher page at a specific time, market, device, and URL. The reliable workflow is: collect the ad and page identifiers, capture the page or placement in a controlled browser session, save the original image, record UTC metadata, validate the evidence, and assemble a dated PDF.

For recurring work, schedule that workflow and keep a manifest beside every image. A screenshot proves what rendered in the captured session; it does not prove spend, reach, conversions, or total campaign delivery. Geography, login state, personalization, auction timing, viewport, and device can all change what appears.

What to capture for a defensible tear sheet

Before writing code, define the evidence record. For every ad, preserve:

The evidence pipeline connects ad discovery, controlled capture, metadata, and report assembly.
The evidence pipeline connects ad discovery, controlled capture, metadata, and report assembly.
  • Advertiser, campaign, platform, publisher, and page URL.
  • Capture timestamp in UTC, market or country, timezone, viewport, device preset, and browser version.
  • Campaign ID, ad ID, creative ID, or the source query and API response ID.
  • Placement dimensions, format, visible headline or offer, and any notes about position.
  • The original PNG, JPEG, WebP, or PDF without recompression.

Microsoft’s Ad Library API supports advertiser lookup and ad queries by search text, date range, country codes, advertiser ID, and pagination. Request single-ad details when you need the complete record, and record the market because its documented library covers ads served in the European Economic Area. Use the library response as your discovery source, then capture the publisher page as evidence.

Choose full-page, placement, or both

Capture Use it when Evidence it provides
Full page The client needs context or the ad can move between modules. Publisher identity, surrounding content, page URL, and placement context.
Element crop The report needs a compact creative proof image. The ad itself at a consistent size.
Both You need auditability and a presentation-ready card. Full context plus a readable placement crop.

Use a consistent viewport and device for a campaign. A mobile ad captured at a desktop width is a different observation and should be labeled as such. If a placement is below the fold, load lazy content before taking the screenshot and wait for the ad selector rather than relying on a fixed delay.

DIY capture with Playwright

A self-managed browser job gives you control over sessions, headers, cookies, retries, and storage. The following Python example captures a full page and an optional CSS-selected ad element, writes a JSON manifest, and uses a fixed UTC timestamp.

from datetime import datetime, timezone
import json
from pathlib import Path
from playwright.sync_api import sync_playwright

PAGE_URL = "https://publisher.example/article"
AD_SELECTOR = "[data-ad-slot='leaderboard']"
OUT = Path("tear_sheet")
OUT.mkdir(exist_ok=True)

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page(viewport={"width": 1440, "height": 1000}, device_scale_factor=1)
    page.goto(PAGE_URL, wait_until="networkidle", timeout=90000)
    page.wait_for_selector(AD_SELECTOR, timeout=30000)
    page.locator(AD_SELECTOR).scroll_into_view_if_needed()
    page.wait_for_timeout(1500)  # allow lazy creative and animation to settle

    captured = datetime.now(timezone.utc).isoformat()
    page.screenshot(path=str(OUT / "page.png"), full_page=True)
    page.locator(AD_SELECTOR).screenshot(path=str(OUT / "placement.png"))

    manifest = {
        "url": PAGE_URL,
        "captured_at_utc": captured,
        "viewport": {"width": 1440, "height": 1000},
        "device_scale_factor": 1,
        "placement_selector": AD_SELECTOR,
        "files": ["page.png", "placement.png"]
    }
    (OUT / "manifest.json").write_text(json.dumps(manifest, indent=2))
    browser.close()

Install the dependency with pip install playwright and playwright install chromium. For a Node.js job, the equivalent is:

import { chromium } from "playwright";

const browser = await chromium.launch();
const page = await browser.newPage({ viewport: { width: 1440, height: 1000 }, deviceScaleFactor: 1 });
await page.goto("https://publisher.example/article", { waitUntil: "networkidle", timeout: 90000 });
await page.locator("[data-ad-slot='leaderboard']").waitFor({ state: "visible", timeout: 30000 });
await page.locator("[data-ad-slot='leaderboard']").scrollIntoViewIfNeeded();
await page.waitForTimeout(1500);
await page.screenshot({ path: "page.png", fullPage: true });
await page.locator("[data-ad-slot='leaderboard']").screenshot({ path: "placement.png" });
await browser.close();

Make the browser session reproducible

  1. Pin the viewport, device scale factor, timezone, locale, and user agent.
  2. Use a named browser context per market or login state. Never mix cookies between advertisers.
  3. Wait for the ad selector, then wait for the creative frame or image to become visible.
  4. Save response status, final URL, redirect chain, and a hash of each original file.
  5. Retry transient navigation failures with exponential backoff, but keep the first failure in the run log.

For pages that require authentication, load a short-lived storage state from a secret store. Do not put credentials in the manifest. If an ad is rendered inside an iframe, locate the frame and screenshot the iframe element or its bounding box. Cross-origin restrictions can prevent DOM inspection even though the browser can render the frame; in that case, retain a full-page capture and document the limitation.

Automate discovery and scheduling

Ad-library results are inputs, not proof. A daily job can fetch records, deduplicate by ad ID and URL, then enqueue captures only when the creative or publisher page has changed. Preserve the original library query and response identifier with the capture.

Use a scheduler such as GitHub Actions, a container cron job, or your existing workflow runner. The self-managed shot-scraper project documents configuration-driven screenshots that can run in GitHub Actions and write results to a repository. Hosted screenshot services are useful when you do not want to maintain Chromium, fonts, browser patches, or worker scaling.

Idempotency and retries

  • Use a key such as platform:ad_id:publisher_url:capture_date:market:viewport.
  • Do not overwrite an existing original. Store a new version and mark the prior result.
  • Retry navigation timeouts and HTTP 5xx responses; do not blindly retry a bot challenge.
  • Set a maximum run duration and send failed items to a review queue.
  • Keep the page HTML or response metadata only when policy allows; the screenshot and manifest are the minimum audit set.

Validate each screenshot before it enters a report

Google Ad Manager’s screenshot-inspection guidance requires a PNG, one creative, no zoom-out, and no compression for its inspection workflow. Apply the same discipline to tear sheets:

  • Confirm the image opens and its dimensions match the manifest.
  • Check that exactly one intended creative is visible in the placement crop.
  • Confirm the crop is readable and not accidentally zoomed out.
  • Look for blank frames, consent dialogs, bot checks, broken images, and loading spinners.
  • Compare the visible headline or offer with the ad-library record.
  • Record whether the page redirected, required login, or showed a market-specific variant.

Automated checks can reject files with very small dimensions, an all-white pixel histogram, or a missing selector. A human review is still needed for creative identity and misleading overlays.

Build a client-ready PDF tear sheet

Put a title page or header on the report with client, campaign, platform, market, capture window, and source. Give every ad a consistent card containing the placement image, full-page context when needed, URL, UTC timestamp, ad or campaign ID, dimensions, and notes. Export a dated PDF and retain the source images plus the metadata manifest.

PDF generation can compile multiple screenshots into a report. Keep the source images at their original quality and generate a separate presentation PDF so a resized page never replaces the evidence file. If you use a hosted service such as Urlbox, compare its screenshot and PDF options with your requirements for scheduling, storage, and retention.

Suggested manifest

{
  "client": "Example Client",
  "campaign": "Spring Launch",
  "platform": "Publisher site",
  "market": "DE",
  "captured_at_utc": "2025-03-21T09:42:11Z",
  "source": {"ad_library_query": "...", "response_id": "..."},
  "items": [{
    "ad_id": "...",
    "publisher_url": "https://publisher.example/article",
    "placement": "leaderboard",
    "viewport": "1440x1000",
    "files": ["page.png", "placement.png"]
  }]
}

Or skip the browser setup

ScreenshotNeo is a hosted website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. It can capture a full page with lazy images loaded, select one element by CSS selector, set a viewport or device preset, use retina scale, apply custom CSS or JavaScript, click an element, wait for a selector, delay, or network idle, and set headers, cookies, user agent, Authorization, timezone, and geolocation. You can also block ads, trackers, requests, or resource types; resize images; choose a cache TTL; create signed links; run async jobs with signed webhooks; capture up to 100 URLs per bulk call; and use the usage API or OpenAPI specification. See the ScreenshotNeo documentation for parameter details.

Consent and overlay handling determines whether the captured creative is actually usable.
Consent and overlay handling determines whether the captured creative is actually usable.

For a simple evidence capture:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie and consent banners are accepted and more than 60 known consent platforms, newsletter popups, and chat widgets are removed before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

Free usage includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is on every plan. Create a free ScreenshotNeo account.

Troubleshooting common failures

Symptom Likely cause Fix
Blank or white image Navigation completed before content rendered, blocked resources, or a bot check. Wait for a selector or network idle, inspect the verdict, and capture the page state for review.
Ad missing Wrong market, viewport, login state, auction timing, or lazy loading. Record geography and device, scroll the slot into view, and retry at the campaign’s intended market.
Cookie dialog covers the creative Consent state was not established. Persist an approved consent state or use a capture service that handles consent before capture.
Element selector times out The selector changed, is inside an iframe, or the ad was not served. Inspect the DOM, locate the frame, allow a documented fallback, and mark the item as not observed.
Screenshot is unreadable Viewport too wide, device scale too low, or an accidental zoom-out. Use the placement crop, increase retina scale, and validate dimensions before PDF assembly.
Repeated 403 or CAPTCHA Publisher bot mitigation or rate limiting. Reduce concurrency, use an allowed authenticated session, or record the bot verdict instead of retrying indefinitely.
PDF pages differ in size Mixed image dimensions or inconsistent page templates. Normalize the report canvas while retaining original files separately.

Performance, reliability, and cost

Full-page captures cost more time and storage than placement crops. Capture both only when the report needs context. Reuse a browser process for a batch, limit concurrency to what the publisher tolerates, and cache stable pages with a clearly documented TTL. For hosted APIs, use async jobs and webhooks for large queues, and bulk capture when you have many URLs.

Cost accounting should include browser workers, storage, PDF rendering, retries, and review time. A cheap screenshot that cannot be audited is expensive to replace. Track success, failed-load, bot-check, and cache outcomes separately. With ScreenshotNeo, only clean shots are billed and response headers expose the billing and page verdict, which makes reconciliation easier.

FAQ

Does a tear sheet prove an ad was served?

It proves that the creative rendered in one captured session. It does not prove impressions, spend, reach, or conversions.

Should I capture the whole page?

Use a full page for context and an element crop for a compact proof image. Many audits retain both.

Why must timestamps be in UTC?

UTC removes daylight-saving and cross-market ambiguity. Store the local market separately when it matters.

Can I automate 100 URLs at once?

Yes, with a queue or a bulk capture API. Keep per-site concurrency controlled and preserve one manifest item per URL.

What if the ad is personalized?

Record the market, device, login state, and capture conditions. Treat the screenshot as evidence of that session’s rendering.

When should I use a PDF?

Use a PDF for client delivery and archival review, while retaining original images and metadata for auditability.