How to Screenshot an Indian News Website with Playwright When Ads Shift the Layout
Capture a news article after its content is ready, handle shifting ad slots, and choose between a full-page image and a stable visual test.
To screenshot a news article whose layout shifts as ads load, wait for the article content that matters, then capture the page in the intended visual state. Do not treat load, a fixed sleep, or network quiet as proof that every ad has settled. Preserve ads for an archival image of the page as seen; mask or hide them only when the purpose is a stable comparison of editorial layout.
This guide uses a generic article URL and selector because no publisher or page was specified. Replace them with a page you are allowed to access and selectors that match that page. A publisher’s automation rules and the site’s current terms were not checked for this guide; review them before repeated capture.
1. Install Playwright and choose a capture goal
The example uses Playwright for Node.js and Chromium. It waits for an article heading, captures a full-page PNG, and closes the browser even if navigation or capture fails.
mkdir news-capture
cd news-capture
npm init -y
npm install playwright
npx playwright install chromium
Save the following as capture.mjs. Set ARTICLE_URL to the article URL. If the page uses a different heading structure, change ARTICLE_HEADING to a selector that identifies meaningful article content.
import { chromium } from 'playwright';
const url = process.env.ARTICLE_URL;
const articleHeading = process.env.ARTICLE_HEADING ?? 'article h1';
if (!url) {
throw new Error('Set ARTICLE_URL to the article URL.');
}
const browser = await chromium.launch({ headless: true });
try {
const context = await browser.newContext({
viewport: { width: 1365, height: 900 },
deviceScaleFactor: 1,
locale: 'en-IN',
timezoneId: 'Asia/Kolkata',
});
const page = await context.newPage();
page.setDefaultTimeout(15_000);
const response = await page.goto(url, {
waitUntil: 'domcontentloaded',
timeout: 45_000,
});
if (!response || !response.ok()) {
throw new Error(`Navigation failed: ${response?.status() ?? 'no response'} ${url}`);
}
// This checks for article content, rather than assuming navigation means ready.
await page.locator(articleHeading).first().waitFor({ state: 'visible' });
// For an article archive image, retain the ad layout and capture the full page.
await page.screenshot({
path: 'article.png',
fullPage: true,
animations: 'disabled',
});
} finally {
await browser.close();
}
Run it with an environment variable. The URL is quoted so shell characters in query strings are not interpreted:
ARTICLE_URL='https://example.com/news/article' node capture.mjs
domcontentloaded is a navigation milestone, not a guarantee that the page is visually complete. Modern pages can fetch data and update their interface later. Playwright recommends relying on assertions about the web state that matters; its navigation guide explains why “loaded” depends on the page and framework. See Playwright navigation guidance and the Page API.
2. Decide whether ads belong in the image
First decide what the image is meant to represent:
- Archive or documentation: retain ads and capture after the relevant ad slots have reached the state you want to preserve. A specific ad may be different on every run.
- Editorial layout check: compare the article and surrounding layout while treating volatile ad content as a masked region or a documented visual style.
- Article-only reference: screenshot the article locator instead of the full page, if excluding surrounding page content fits the purpose.
Ads, cookie notices, images, fonts, and widgets can arrive or resize after initial rendering. Inserting an element above the article can move everything below it; resizing an ad slot can shift nearby content. A fixed delay may happen to catch one run, but it is not a readiness condition. If you own the page, reserving space for known ad dimensions can reduce shifts. If you do not own it, wait for the meaningful page state and make your capture goal explicit.
When you know the ad-slot selector and need to preserve its rendered state, wait for the slot to exist and become visible. This confirms visibility, not that the ad will never change again:
const adSlot = page.locator('[data-ad-slot="top"]');
await adSlot.waitFor({ state: 'visible', timeout: 10_000 });
await page.screenshot({ path: 'article-with-ad.png', fullPage: true });
Do not copy that selector unless it exists on your target page. Inspect the page and choose a selector appropriate to that publisher. A third-party slot can remain empty, be blocked, or refresh after becoming visible.
3. Choose full-page, article-only, or clipped capture
Use fullPage: true when the target is the full scrollable page. Use a locator screenshot for a single article region; it captures the element’s bounds and avoids an extremely tall image of the whole site.
// Full scrollable page
await page.screenshot({ path: 'full-page.png', fullPage: true });
// Just the article region (adjust selector for the target site)
await page.locator('article').screenshot({ path: 'article-region.png' });
Full-page screenshots can be very tall, consume more memory, and expose more content than needed. If your goal is the story rather than the surrounding page, a locator screenshot is often easier to inspect. If you need a specific viewport region, use clip:
await page.screenshot({
path: 'top-of-article.png',
clip: { x: 0, y: 0, width: 1200, height: 1600 },
});
Clip coordinates are page coordinates in CSS pixels. Ensure the clip fits the rendered page and viewport assumptions you set. Screenshot options and locator capture behavior are documented in the Page API and Locator API.
4. Stabilize visual regression captures without hiding the problem
For repeatable visual checks, use Playwright Test’s toHaveScreenshot assertion. It waits until two consecutive screenshots match before comparing with the baseline. That helps control short-lived rendering differences, but it cannot make third-party ads serve identical content across separate sessions. Masks and screenshot styles are useful only when those dynamic areas are intentionally outside the comparison.
Install the test runner if needed:
npm install --save-dev @playwright/test
npx playwright install chromium
Example test in tests/article.spec.mjs:
import { test, expect } from '@playwright/test';
test('article layout matches its baseline', async ({ page }) => {
const response = await page.goto(process.env.ARTICLE_URL, {
waitUntil: 'domcontentloaded',
});
expect(response?.ok()).toBeTruthy();
await expect(page.locator('article h1').first()).toBeVisible();
await expect(page).toHaveScreenshot('article-layout.png', {
fullPage: true,
animations: 'disabled',
mask: [page.locator('[data-ad-slot]')],
maskColor: '#888888',
});
});
Run with an article URL:
ARTICLE_URL='https://example.com/news/article' npx playwright test
The mask selector is only an example. If it matches nothing or matches the article itself, adjust it for the page. A mask makes a region visually uniform in the assertion; it does not remove the underlying element from the page or prove that the page is stable. Playwright documents screenshot assertion options including animations, fullPage, masks, and screenshot styles in its visual comparisons guide.
Use a screenshot style when you need a repeatable CSS treatment for dynamic regions; document the style as part of the test so reviewers know which page regions are excluded from visual comparison. Do not hide ads in an archival capture if their presence or layout is the thing you need to preserve.
5. Avoid using network idle as the finish line
Playwright defines networkidle around a 500 ms period without network connections and discourages using it as a testing readiness condition. Ad auctions, analytics, polling, and long-lived connections can keep traffic active; conversely, a quiet network does not prove that the content you need is visible or that a later timer will not move the layout. Prefer a semantic assertion such as the article heading or a known slot becoming visible. See the Page API load-state documentation.
If a page has a known, finite transition with no useful DOM condition, a short delay can be used as a last-mile settling window after the meaningful content appears. It is a heuristic, not a guarantee:
await page.locator('article h1').waitFor({ state: 'visible' });
await page.waitForTimeout(800); // heuristic only; tune for your page and capture purpose
await page.screenshot({ path: 'article.png', fullPage: true });
For stronger observation, record the relevant element’s bounding box at intervals and proceed only after it remains unchanged for a chosen window. This still cannot prove a future third-party update will not occur:
async function waitForStableBox(locator, { samples = 4, intervalMs = 300 } = {}) {
let previous = null;
let stable = 0;
while (stable < samples) {
const box = await locator.boundingBox();
if (!box) throw new Error('Target element is not visible or has no box');
const current = [box.x, box.y, box.width, box.height].map(v => Math.round(v));
if (JSON.stringify(current) === JSON.stringify(previous)) stable += 1;
else stable = 0;
previous = current;
await locator.page().waitForTimeout(intervalMs);
}
}
const heading = page.locator('article h1').first();
await heading.waitFor({ state: 'visible' });
await waitForStableBox(heading);
await page.screenshot({ path: 'article.png', fullPage: true });
This checks the heading’s box, not every ad slot or all pixels on the page. For a page where ad movement is the problem, observe the specific slot or article container whose movement matters.
6. Keep rendering settings consistent
For visual comparisons, use the same browser engine and version, operating system, viewport, device scale factor, locale, timezone, fonts, and headless setting for baseline and later runs. Playwright notes that screenshots can differ with the host OS, browser version, settings, hardware, power source, and headless mode. Keep capture configuration under version control alongside the baseline.
- Viewport: set a fixed width and height; responsive breakpoints can change both article and ad layout.
- Device scale: the default device scale can produce more image pixels on high-DPI contexts. Use CSS-pixel scale for one output pixel per CSS pixel where supported by your installed version.
- Locale and timezone: set these if date formatting or location-sensitive rendering matters.
- Fonts and assets: ensure the capture environment can load the same fonts and resources; font substitution changes wrapping and page height.
- Browser release: pin and update Playwright deliberately, then review visual baseline changes after an upgrade.
Example explicit context configuration:
const context = await browser.newContext({
viewport: { width: 1365, height: 900 },
deviceScaleFactor: 1,
locale: 'en-IN',
timezoneId: 'Asia/Kolkata',
});
const page = await context.newPage();
await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.locator('article h1').waitFor({ state: 'visible' });
await page.screenshot({ path: 'css-pixel-layout.png', fullPage: true, scale: 'css' });
Check the API for your installed Playwright version before using version-specific options such as screenshot scale or screenshot style. The browser and capture conditions are described in the Page API.
7. cURL, Python, and Node.js alternatives
Playwright’s browser automation API is JavaScript/TypeScript, Java, Python, and .NET. The title’s method uses Node.js above; here is a minimal Python Playwright equivalent for readers who prefer Python.
from playwright.sync_api import sync_playwright
import os
url = os.environ["ARTICLE_URL"]
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
try:
page = browser.new_page(
viewport={"width": 1365, "height": 900},
device_scale_factor=1,
locale="en-IN",
timezone_id="Asia/Kolkata",
)
response = page.goto(url, wait_until="domcontentloaded", timeout=45_000)
if response is None or not response.ok:
raise RuntimeError(f"Navigation failed: {response.status if response else 'no response'}")
page.locator("article h1").first.wait_for(state="visible", timeout=15_000)
page.screenshot(path="article.png", full_page=True, animations="disabled")
finally:
browser.close()
Install with python -m pip install playwright and python -m playwright install chromium. Run with ARTICLE_URL='https://example.com/news/article' python capture.py after saving the code as capture.py.
cURL alone cannot render a JavaScript-driven page or capture a browser screenshot. It can fetch the HTML response for diagnosis, but it is not an equivalent screenshot method:
curl -L --fail --show-error 'https://example.com/news/article' -o article.html
8. Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Screenshot has a blank article area | The capture ran after navigation but before client-rendered content appeared, or the selector is wrong. | Wait for the page’s actual article heading or content locator; inspect the DOM and replace the generic selector. |
| Article is cut off or image is unexpectedly tall | The capture mode does not match the goal; full-page mode includes the whole scrollable document. | Use a locator screenshot for the article region or a deliberate clip. Keep full-page capture for a full-page record. |
| Layout still shifts after the wait | A later ad, image, font, or widget update changed the layout after the observed condition. | Wait for the relevant slot or observe the target’s geometry. A delay is only a heuristic. For a visual test, mask the volatile area only if it is outside the test’s purpose. |
TimeoutError waiting for the heading |
Wrong selector, slow or failed navigation, consent overlay, or article not available in this browser context. | Check the response status and URL, inspect the rendered DOM, and use a locator that exists on this page. Handle consent only in a way consistent with the capture purpose and site rules. |
net::ERR_* or no response |
Network, DNS, TLS, proxy, or publisher-side access issue. | Retry with bounded retries for transient failures, inspect browser logs and response status, and verify the URL is reachable from the capture environment. |
| Visual assertion fails on every run | Third-party ad creative, changing timestamps, rotating content, font differences, or unstable viewport. | Make environment settings consistent. Mask or style only known volatile regions, or test a more focused article locator. Review baseline changes instead of blindly updating snapshots. |
| Chromium executable missing | The Playwright package is installed but its browser binaries are not. | Run npx playwright install chromium (or the matching Python install command) in the runtime environment. |
9. Performance, reliability, and cost
Browser startup, navigation, fonts, images, third-party ads, and full-page rasterization all contribute to runtime. Reuse a browser process for batches of captures while creating a fresh context per independent session; close pages and contexts when finished. Limit concurrency to what the machine can render without memory pressure. A full-page screenshot of a long article uses more memory and storage than an article-only image.
Reliability improves when you check navigation status, assert the content that matters, give navigation and locator waits explicit timeouts, and save diagnostic information when a capture fails. Avoid retrying permanent failures indefinitely. If you retry transient errors, cap attempts and use a delay between attempts. Third-party ads cannot be made deterministic by Playwright; define whether their content is part of the expected result.
Self-hosted Playwright has no per-shot API price in this example, but it uses your compute, storage, maintenance time, and outbound network. Account for browser updates and the operational cost of handling failed or slow pages. A hosted capture API trades browser setup and maintenance for a service charge; compare the exact workflow and billing rules before moving production captures.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF. The following request uses the documented API shape; replace YOUR_API_KEY with your key. See the ScreenshotNeo API documentation for options and parameter details.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com/news/article \
-o article.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={
"access_key": "YOUR_API_KEY",
"url": "https://example.com/news/article",
},
timeout=90,
)
r.raise_for_status()
open("article.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com/news/article',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('article.webp', new Uint8Array(await res.arrayBuffer()));
ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. It bills only clean shots: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with response headers identifying the page verdict and billing state. Its MCP server provides screenshot tools for AI agents. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. This is useful when you want a hosted browser workflow without managing browser installation, although the DIY Playwright path gives you direct control over selectors and capture behavior. Create a free ScreenshotNeo account for 1,000 screenshots a month, with no card.
FAQ
Does fullPage: true wait for ads to finish loading?
No. It controls capture extent. Wait for the content or ad state you care about before calling the screenshot method.
Can I get the same ad creative in every run?
Not reliably from a public third-party page. Ads may rotate or vary by session and location. For a repeatable editorial test, exclude the ad region from comparison if that matches the test’s purpose.
Should I always hide ads to prevent layout shifts?
No. Hiding ads changes the rendered page and may change the layout being recorded. Keep them for an as-seen archive; hide or mask only for a clearly defined comparison that excludes them.
Why can two screenshots differ when the page content is unchanged?
Browser, operating system, fonts, viewport, device scale, rendering mode, and dynamic page elements can all affect pixels. Keep the environment consistent and isolate regions that are intentionally dynamic.


