ScreenshotNeo

BlogHow-to

How to Capture Screenshots of Indian News Article URLs in Bulk

Use Playwright to capture Indian news URLs in bulk, save consistent screenshots, and track failures. Includes runnable Python code and a hosted API option.

By the ScreenshotNeo team4 October 202611 min read

To capture screenshots of Indian news article URLs in bulk, put one absolute URL per line in a UTF-8 text file, then use a Playwright script to open each page, save a screenshot, and record the result in a manifest. Start with sequential captures and a consistent viewport; use full-page mode when you need the entire scrollable article. Inspect the output because news pages can show consent screens, login walls, popups, sticky elements, or incomplete dynamic content.

This guide uses Python and Playwright for the do-it-yourself workflow. It also covers viewport and full-page choices, output naming, failures, batch reliability, India-specific rendering, and a hosted API option.

1. Choose the screenshot you need

Capture mode Use it for Trade-off
Viewport A consistent record of what appeared in a particular browser window. Content outside the viewport is not included.
Full page The full scrollable article in one image. Very long pages produce tall, potentially large files; fixed and sticky elements may appear unexpectedly.

Playwright defines a full-page screenshot as a capture of the full scrollable page, as if the page fit on a very tall screen (Playwright screenshots documentation). For comparisons or records, use the same viewport, scale, image format, and wait strategy for each URL, and record those settings with the batch.

Use viewport captures when you want a comparable above-the-fold view. Choose full-page captures when the article content below the fold matters. Check representative results before processing the whole list: a navigation event does not guarantee that every publisher has finished rendering its article.

2. Install Playwright and prepare the URL list

Install Python 3, then create a project environment and install Playwright and its Chromium browser:

python -m venv .venv
# macOS or Linux:
source .venv/bin/activate
# Windows PowerShell:
# .venv\Scripts\Activate.ps1
python -m pip install playwright
python -m playwright install chromium

Save one absolute https:// article URL per line in urls.txt. Blank lines and lines beginning with # are ignored. Keep this file as the input record; the script writes a separate CSV manifest so each result can be matched back to its original row.

https://example.com/news/article-one
https://example.in/story/article-two
# Add one URL per line

Only include pages you can access through ordinary, authorized browsing. Do not try to bypass a paywall, login, CAPTCHA, or other access control. If the page is inaccessible, record the failure and seek an authorized route.

3. Run a sequential batch capture

Save this as capture_news.py beside urls.txt. It launches Chromium once and creates a fresh page for every URL. Each input row gets a stable, zero-padded filename. The CSV records the original row, requested and final URLs, capture time, result, output path, and any error. A failure for one URL does not stop later rows.

import csv
import re
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlparse

from playwright.sync_api import TimeoutError as PlaywrightTimeoutError
from playwright.sync_api import sync_playwright

INPUT = Path("urls.txt")
OUTPUT_DIR = Path("screenshots")
MANIFEST = Path("manifest.csv")
FULL_PAGE = True
VIEWPORT = {"width": 1365, "height": 900}
NAVIGATION_TIMEOUT_MS = 45_000
WAIT_UNTIL = "domcontentloaded"  # alternatives: "load", "networkidle", "commit"


def read_urls(path):
    rows = []
    with path.open("r", encoding="utf-8-sig") as source:
        for row_number, raw in enumerate(source, start=1):
            url = raw.strip()
            if not url or url.startswith("#"):
                continue
            rows.append((row_number, url))
    return rows


def safe_filename_part(value):
    value = re.sub(r"[^A-Za-z0-9._-]+", "-", value).strip("-._")
    return (value or "page")[:80]


def output_name(row_number, url):
    parsed = urlparse(url)
    domain = safe_filename_part(parsed.netloc.lower())
    slug = safe_filename_part(parsed.path.rstrip("/").split("/")[-1])
    return f"{row_number:04d}-{domain}-{slug}.png"


def main():
    OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
    rows = read_urls(INPUT)
    fields = [
        "input_row", "requested_url", "final_url", "captured_at_utc",
        "status", "output_path", "error", "full_page", "viewport_width",
        "viewport_height", "wait_until",
    ]

    with MANIFEST.open("w", newline="", encoding="utf-8") as manifest_file:
        writer = csv.DictWriter(manifest_file, fieldnames=fields)
        writer.writeheader()

        with sync_playwright() as playwright:
            browser = playwright.chromium.launch(headless=True)
            try:
                for row_number, url in rows:
                    page = None
                    path = OUTPUT_DIR / output_name(row_number, url)
                    record = {
                        "input_row": row_number,
                        "requested_url": url,
                        "final_url": "",
                        "captured_at_utc": datetime.now(timezone.utc).isoformat(),
                        "status": "failed",
                        "output_path": str(path),
                        "error": "",
                        "full_page": FULL_PAGE,
                        "viewport_width": VIEWPORT["width"],
                        "viewport_height": VIEWPORT["height"],
                        "wait_until": WAIT_UNTIL,
                    }
                    try:
                        parsed = urlparse(url)
                        if parsed.scheme not in ("http", "https") or not parsed.netloc:
                            raise ValueError("Expected an absolute http:// or https:// URL")

                        page = browser.new_page(viewport=VIEWPORT)
                        page.goto(
                            url,
                            wait_until=WAIT_UNTIL,
                            timeout=NAVIGATION_TIMEOUT_MS,
                        )
                        record["final_url"] = page.url
                        page.screenshot(
                            path=str(path),
                            full_page=FULL_PAGE,
                            type="png",
                            animations="disabled",
                        )
                        record["status"] = "success"
                    except (PlaywrightTimeoutError, Exception) as exc:
                        record["error"] = f"{type(exc).__name__}: {exc}"
                        if page is not None:
                            record["final_url"] = page.url
                    finally:
                        if page is not None:
                            page.close()
                        writer.writerow(record)
                        manifest_file.flush()
                        print(f"{record['status']}: row {row_number} -> {path}")
            finally:
                browser.close()


if __name__ == "__main__":
    main()

Run it with:

python capture_news.py

Images go in screenshots/, and manifest.csv records each attempt. The script uses PNG to preserve readable text without a quality setting. Playwright also supports JPEG and WebP, with a quality option for those formats; change the extension and screenshot type together if you choose another format.

The exception clause catches timeouts and other per-URL errors so the batch continues. If you want more specific handling in your own version, catch and report navigation, validation, and screenshot errors separately. Avoid putting page titles or other untrusted page content into filenames.

4. Tune navigation waits and screenshot options

Playwright’s page.goto supports several navigation milestones. The best choice depends on each publisher’s page behavior:

Wait value What it indicates When to consider it
commit The response has been received and document loading has begun. When you want the earliest navigation milestone and can handle page readiness separately.
domcontentloaded The initial HTML document has been parsed. A practical starting point for a mixed batch; client-side content may still be loading.
load The page load event has fired. When the page’s load event is a useful readiness signal.
networkidle Network activity has been idle for a period. Only when suitable for the site; analytics or long-lived requests can make it slow or unreliable.

After navigation, you can wait for a site-specific article selector or a short delay if representative pages show that the content appears later. A selector can be absent on other publishers, so handle its timeout rather than treating it as proof that the URL failed. Do not assume one wait condition works for every news site.

Playwright’s screenshot options include output path, image format, JPEG/WebP quality, full-page capture, scale, clipping, and masking. The script uses PNG, full-page capture, and disabled animations. Set full_page=False for the current viewport. Use a clip when only a known region matters, and a mask when a dynamic region should be obscured in a screenshot; masking is not a substitute for permission to capture or share page content.

Scale controls whether the output follows CSS pixels or device pixels. Device-scale output can be substantially larger on high-DPI pages. Keep scale consistent across a batch and choose it based on text legibility and storage. See the Playwright Page screenshot API for the current option details.

5. Capture pages as they appear from India

A local Playwright browser renders from the machine where the script runs. If the page varies its content by location, running the browser in India may help capture the version served to that location. This does not guarantee a particular result: publishers can vary behavior for many reasons, and ordinary access restrictions still apply.

For a hosted service, check whether it supports the location you need, how it routes traffic, and what latency or usage limits apply. ScreenshotOne documents an India location option based on proxies and says proxy routing is slower; that is a vendor-documented capability, not a compatibility test of any particular Indian news site. See its location and proxy documentation. Confirm current options and terms before relying on them.

6. Review results and recover from common problems

A successful screenshot call means an image was saved; it does not prove the image contains the expected article. Review a sample from each publisher and inspect every failed capture. Check for a consent screen, popup, login wall, paywall, empty content area, sticky overlay, clipped article, or page that had not finished rendering.

Symptom Likely cause What to do
Navigation timeout The publisher is slow, the network is unavailable, or the chosen wait condition never arrives. Check the URL manually, try a longer timeout or an earlier navigation milestone, then add a site-appropriate readiness wait. Keep the timeout recorded.
Screenshot is blank or incomplete The page is still rendering, content is loaded by script, or the site returned an interstitial. Inspect the final URL and page image. Wait for a relevant selector or a measured delay, then recapture. Record interstitials as such; do not bypass access controls.
Consent banner or popup covers the story The site presents an overlay during ordinary browsing. For local automation, handle consent only through an authorized, user-like flow and follow the publisher’s terms. Otherwise retain the result and mark it for review.
Article text is missing below the fold The capture used viewport mode, or lazy-loaded content was not present before capture. Use full-page mode and inspect the resulting image. If needed, scroll through the page before capture and verify the content appears.
Output files overwrite each other Names were based only on a slug, or different URLs share the same last path segment. Include the original row number, as this script does, and preserve the requested URL in the manifest.
Browser fails to launch Playwright’s Chromium browser was not installed for the active environment. Run python -m playwright install chromium in the same environment used to run the script.
Batch stops after a bad URL An exception escaped the per-URL capture block or the manifest could not be written. Keep errors isolated per row, check that the output directory is writable, and ensure the manifest destination is writable.
Very large images or slow processing Full-page and device-scale captures create large outputs; pages may also be unusually long. Use viewport mode or CSS scale if that meets the need, and consider JPEG/WebP with an appropriate quality setting. Keep the settings consistent when comparing results.

7. Run the batch reliably and control its cost

  • Start sequentially. One page at a time makes failures easier to diagnose and avoids adding load to publishers before you understand their response.
  • Use a fresh page per URL. This limits accidental state carryover. If a workflow intentionally uses cookies or authentication, handle that state explicitly and securely.
  • Preserve a manifest. Keep requested URL, final URL, timestamp, status, error, output path, viewport, full-page setting, and wait strategy. Flush records as the batch proceeds so an interrupted run still leaves useful progress.
  • Retry selectively. Retry transient network or timeout failures with a small bounded retry count and a delay. Do not repeatedly hammer a site or retry access denials as if they were transient errors.
  • Add concurrency only after measuring. Parallel pages can increase memory use and load on both your machine and target sites. Increase gradually and respect publisher rules.
  • Plan storage. PNG and full-page/device-scale captures can consume substantial disk space. Choose format and scale for the task, and define a retention period for images and manifests.
  • Budget for local resources. Playwright itself does not charge per screenshot, but browser runtime, bandwidth, machine capacity, and storage have costs. Managed APIs may have plan limits and fees; compare current terms before choosing.
  • Protect credentials and data. Do not commit API keys, authenticated cookies, or sensitive screenshots to a public repository. Restrict access to manifests if they contain private URLs or workflow details.

For evidence or later comparison, record the viewport dimensions and whether full-page mode was used. A screenshot is a record of one capture at one time and configuration; it does not establish that the page is unchanged or that the image may be republished.

8. Hosted options for bulk capture

A managed service can remove local browser installation and provide an API or no-code workflow. It also adds a third-party processing dependency. Compare batch limits, concurrency, location support, dynamic-page options, output retention, key security, retries, error reporting, privacy terms, and current pricing before moving a URL list to a provider.

  1. ScreenshotNeo — clean screenshots with cookie and consent banners, newsletter popups, and chat widgets removed before capture; only clean shots are billed, and its lowest paid plan is $5 for 3,000 shots.
  2. ScreenshotOne documents a /bulk endpoint with shared defaults and per-request overrides. Its documentation says bulk requests use the same one-minute request bucket as regular requests and advises checking concurrency before draining a large queue.
  3. Urlbox documents API and no-code batch workflows, including URL lists from CSV files, Google Sheets, or Airtable, and lists news-site archiving among its use cases.

These are provider-documented capabilities, not independent compatibility tests or guarantees. The local Playwright workflow remains useful when you need direct browser control; a hosted service may suit a managed or no-code process. Verify current limits and pricing with each provider.

Or skip the browser setup

ScreenshotNeo documentation · API base: https://api.screenshotneo.com/v1/shot. One GET request returns an image; this example saves a WebP response:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Change the example target URL to an article URL. For bulk capture, call the endpoint once per URL and record each response alongside your input row. ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000.

Sign up for 1,000 free screenshots a month, with no card required.

9. Permissions and responsible use

Capturing a page you can access does not by itself establish permission to redistribute or publish its screenshot. Check the publisher’s terms and obtain permission where required for your intended use. This guide does not resolve India-specific copyright or publisher-term questions. Keep private documentation separate from public reuse, and do not circumvent login, paywall, CAPTCHA, or other access controls.

FAQ

Yes. The script reads one URL per line and writes a screenshot plus a manifest row for each non-empty, non-comment line.

Should I use full-page mode for every article?

No. Use it when the full scrollable article matters. Use a viewport capture for a controlled view of what was above the fold.

Does running the browser in India guarantee the same page every time?

No. Location can affect content, but each publisher’s behavior and the capture conditions can vary. Record where and how the browser ran if location matters to your workflow.

Can I publish the screenshots I collect?

That depends on the publisher’s terms and the intended use. Check the applicable terms and obtain permission where needed before redistribution.

Sources