ScreenshotNeo

BlogHow-to

How to Bulk Screenshot Indian Job Listing Pages from a URL List

Use Playwright to save one screenshot per job listing URL, with stable filenames, failure logs, and careful checks of each site's automation rules.

By the ScreenshotNeo team4 October 202611 min read

Direct answer: Read the listing URLs from a text or CSV file, validate and deduplicate them, then use a browser automation script to visit each permitted page and save a screenshot. Playwright provides the page navigation and screenshot operations; batching is a loop in your script, not a built-in bulk operation. Choose a viewport screenshot for the visible part of a listing or a full-page screenshot when below-the-fold details matter. Playwright Page API.

Before automating, check the rules for every site on your list. Do not bypass logins, CAPTCHAs, paywalls, or other access controls. For example, LinkedIn says third-party software and extensions that scrape or automate activity are prohibited, and WorkIndia’s terms restrict scraping and automated extraction. These are specific platform policies, not a rule that applies identically to every job site. See LinkedIn’s policy and WorkIndia’s terms. For a site whose terms prohibit automation or are unclear, use an authorized export or API, or get permission.

Choose what each screenshot should preserve

Mode What it captures Use it when
Viewport The currently visible browser area The top of the listing is the record you need, or you want compact, consistent images.
Full page The full scrollable page Responsibilities, qualifications, or other details below the fold matter. Playwright documents full-page capture directly.
Element A specific page element You have inspected the page and know a stable selector for the listing content. Selector-based capture can fail if a site’s markup changes.

For a useful record, keep the source URL and capture time with each image. A screenshot shows what was rendered at capture time; by itself, it does not establish that a listing is still open or current. Limit capture and storage to what you need, especially where pages expose personal or account-specific information.

Prepare the URL list

Use one URL per line in a UTF-8 text file called urls.txt:

https://example.com/jobs/listing-one
https://example.org/careers/listing-two

Replace these examples with URLs you are authorized to capture. The script below accepts only HTTP and HTTPS URLs, removes duplicates, assigns stable numbered filenames, records timestamps and errors in a CSV manifest, and retries failed navigations a limited number of times. It does not attempt to defeat access controls.

DIY: capture a URL list with Python and Playwright

1. Install Playwright

Install the Python package and its Chromium browser:

python -m pip install playwright
python -m playwright install chromium

2. Save the script

Save the following as bulk_screenshot.py. It reads either a plain text file or a CSV with a url column. Captures run sequentially to keep load modest and make failures easier to diagnose.

import argparse
import csv
import re
import time
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlparse

from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError


def read_urls(path: Path) -> list[str]:
    if path.suffix.lower() == ".csv":
        with path.open(newline="", encoding="utf-8-sig") as f:
            rows = csv.DictReader(f)
            if not rows.fieldnames or "url" not in rows.fieldnames:
                raise ValueError("CSV must have a column named 'url'")
            candidates = [row.get("url", "").strip() for row in rows]
    else:
        candidates = [line.strip() for line in path.read_text(encoding="utf-8").splitlines()]

    unique = []
    seen = set()
    for url in candidates:
        if not url or url.startswith("#"):
            continue
        parsed = urlparse(url)
        if parsed.scheme not in ("http", "https") or not parsed.netloc:
            print(f"Skipping invalid URL: {url}")
            continue
        if url not in seen:
            seen.add(url)
            unique.append(url)
    return unique


def safe_name(url: str, index: int) -> str:
    parsed = urlparse(url)
    # The index prevents collisions when two URLs have the same path.
    tail = re.sub(r"[^A-Za-z0-9_-]+", "-", parsed.path.strip("/") or parsed.netloc)
    return f"{index:04d}-{tail[:70] or 'listing'}.png"


def main() -> None:
    parser = argparse.ArgumentParser()
    parser.add_argument("input", type=Path, help="UTF-8 .txt or .csv containing URLs")
    parser.add_argument("--out", type=Path, default=Path("screenshots"))
    parser.add_argument("--full-page", action="store_true", help="capture the full scrollable page")
    parser.add_argument("--wait-ms", type=int, default=1000, help="extra render wait after navigation")
    parser.add_argument("--timeout-ms", type=int, default=30000)
    parser.add_argument("--retries", type=int, default=1, help="additional navigation attempts")
    args = parser.parse_args()

    if args.wait_ms < 0 or args.timeout_ms < 1 or args.retries < 0:
        parser.error("wait and retry values must be non-negative; timeout must be positive")

    urls = read_urls(args.input)
    if not urls:
        raise SystemExit("No valid URLs found")

    args.out.mkdir(parents=True, exist_ok=True)
    manifest_path = args.out / "manifest.csv"
    with manifest_path.open("w", newline="", encoding="utf-8") as manifest_file:
        fields = ["index", "url", "final_url", "captured_at_utc", "file", "status", "error"]
        writer = csv.DictWriter(manifest_file, fieldnames=fields)
        writer.writeheader()

        with sync_playwright() as p:
            browser = p.chromium.launch(headless=True)
            page = browser.new_page(viewport={"width": 1365, "height": 900}, device_scale_factor=1)
            page.set_default_navigation_timeout(args.timeout_ms)

            for index, url in enumerate(urls, start=1):
                filename = safe_name(url, index)
                row = {
                    "index": index, "url": url, "final_url": "", "captured_at_utc": "",
                    "file": filename, "status": "failed", "error": "",
                }
                for attempt in range(args.retries + 1):
                    try:
                        response = page.goto(url, wait_until="domcontentloaded")
                        # This is a bounded, configurable render pause, not a guarantee
                        # that every site's client-side content has finished loading.
                        if args.wait_ms:
                            page.wait_for_timeout(args.wait_ms)
                        row["final_url"] = page.url
                        if response is not None and response.status >= 400:
                            raise RuntimeError(f"HTTP status {response.status}")
                        page.screenshot(path=str(args.out / filename), full_page=args.full_page)
                        row["captured_at_utc"] = datetime.now(timezone.utc).isoformat()
                        row["status"] = "captured"
                        break
                    except (PlaywrightTimeoutError, Exception) as exc:
                        row["error"] = f"{type(exc).__name__}: {exc}"
                        if attempt < args.retries:
                            time.sleep(min(2 ** attempt, 8))
                writer.writerow(row)
                manifest_file.flush()
                print(f"{row['status']}: {url}" + (f" ({row['error']})" if row["error"] else ""))

            browser.close()


if __name__ == "__main__":
    main()

The exception clause can be simplified to except Exception as exc; it is shown with the Playwright timeout type named explicitly for readability. If you prefer to avoid a redundant exception tuple, replace that clause with except Exception as exc:.

3. Run it

python bulk_screenshot.py urls.txt
python bulk_screenshot.py urls.txt --full-page
python bulk_screenshot.py listings.csv --out captures --wait-ms 1800 --timeout-ms 45000 --retries 2

For CSV input, include a header and a url column:

url,notes
https://example.com/jobs/listing-one,operations role
https://example.org/careers/listing-two,engineering role

Output consists of numbered PNG files and manifest.csv. The manifest records the original URL, final URL after redirects, UTC capture time, status, and error. Review failed rows rather than treating them as successful screenshots.

Wait for the content you need

domcontentloaded waits for the initial document parse, not every client-rendered component, image, or embedded resource. A short fixed wait is simple but may be wasteful or insufficient. If you know a stable selector for the listing, wait for it before capture:

page.goto(url, wait_until="domcontentloaded")
page.locator("article.job-listing").wait_for(state="visible", timeout=15000)
page.screenshot(path="listing.png", full_page=True)

Replace article.job-listing with a selector confirmed for that site. Sites often use different markup, so a selector that works on one portal may not work on another. For a page that loads content only as you scroll, inspect whether scrolling or a site-provided export is necessary; do not assume a full-page screenshot triggers every lazy-load behavior.

Configuration and edge cases

  • Viewport: The example uses 1365 × 900 CSS pixels and scale factor 1. Change the viewport if you need a particular desktop layout. Responsive pages may show different content at different sizes.
  • Full page: Enable --full-page to include below-the-fold content. Very long pages can produce large images or fail due to memory limits; use viewport captures or an authorized targeted export when appropriate.
  • Wait behavior: Increase --wait-ms when ordinary rendering needs more time. Prefer a relevant selector wait when available. A fixed pause cannot guarantee content completeness.
  • Timeout and retries: Set a larger timeout for slow pages and a small retry count for transient failures. Retries do not fix access restrictions or persistently broken pages.
  • Redirects: The manifest keeps the final URL. Check it when a listing redirects to a sign-in page, a generic landing page, or an unavailable notice.
  • Duplicate and malformed URLs: Exact duplicate strings are removed; invalid schemes and missing hosts are skipped. URLs with different query strings remain distinct, even if the site renders the same listing.
  • Filename collisions: The sequence number makes names unique within one run. A later run using the same output directory can overwrite files with the same names. Use a new directory for each batch if you need to retain prior runs.
  • Personalized pages: The script starts a fresh browser context without your logged-in session. Do not add credentials or automate access unless the site permits it and you are authorized.
  • Failure classification: An HTTP error status is logged as a failure. A portal may instead return a normal status with a bot check or generic page; inspect the image and final URL because a successful navigation does not prove the intended listing loaded.

cURL, Python requests, and Node.js options

These examples call ScreenshotNeo’s hosted screenshot API for one URL at a time. They are useful when you want to avoid managing a local browser. For a local bulk workflow, loop over your authorized URL list and make one request per URL, recording the response and outcome for every item. The API accepts common screenshot parameters; see the ScreenshotNeo API documentation for available options and exact parameter names.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com/jobs/listing-one \
  -o listing.webp

Python requests

import requests

url = "https://example.com/jobs/listing-one"
r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": url},
    timeout=90,
)
r.raise_for_status()
with open("listing.webp", "wb") as f:
    f.write(r.content)

Node.js

const target = 'https://example.com/jobs/listing-one';
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: target });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('listing.webp', bytes));

Keep API keys out of source control and logs. For a batch, use a conservative number of simultaneous requests and write each response to a unique filename. Check the response headers and API documentation to distinguish clean captures from non-captures or cache hits.

Or skip the browser setup

ScreenshotNeo takes a screenshot with one API call. Its capture process accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses include X-Page-Verdict and X-Billed headers. It also has an MCP server with take_screenshot, get_page_info, and capture_pdf tools for AI agents.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/jobs/listing-one -o listing.webp

See the ScreenshotNeo docs for request options and bulk capture. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Performance, reliability, and cost

  • Local browser workload: A browser process uses memory and CPU, and full-page images use more resources than viewport images. Start sequentially; only increase concurrency after checking the target site’s rules and monitoring resource use.
  • Batch reliability: Keep per-URL results in a manifest, flush progress as the script runs, and retry only a limited number of transient failures. Resume by processing only failed rows or writing each run to its own directory.
  • Freshness: A capture is a point-in-time visual record. Save its timestamp and source URL, and recapture when you need a later state.
  • Storage: PNG files can be large, especially for long pages. Choose image format and capture scope based on the evidence you need, and define a retention period for images and URL lists.
  • Hosted API pricing: ScreenshotNeo has a free tier of 1,000 shots per month with no card. Paid plans are Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing gives two months free. Every feature is available on every plan. Check current plan details on the product site before choosing.

Troubleshooting

Symptom Likely cause What to do
Navigation timeout The page is slow, never completes navigation, or holds open network activity. Try a longer timeout and a less strict navigation wait such as domcontentloaded. Wait for the content selector you need instead of expecting all network activity to stop.
Screenshot shows a blank or partial listing Client rendering or lazy content was not ready when capture ran. Wait for a visible listing selector, increase the bounded render pause, or inspect the page manually. A full-page capture alone may not load lazy content.
Screenshot shows a bot check, login, or generic page The site redirected or restricted the request. Check the manifest’s final URL and image. Follow the site’s terms and use an authorized method; do not try to bypass the control.
HTTP status error The server returned an error response, or the listing was removed. Check the URL in an ordinary browser, confirm the listing still exists, then retry only if the error appears transient.
Selector wait fails The selector is wrong, changed, or differs by portal. Inspect the page DOM and use a site-specific selector, or use a bounded delay and review the resulting capture.
Output image overwritten A later run reused numbered filenames in the same folder. Use a separate output directory for each batch or add a run identifier to filenames.
API request returns an error The key, URL encoding, request, or account configuration may be incorrect. Check the key and encoded target URL, inspect the HTTP status and response headers, and consult the API docs.

FAQ

Can I screenshot every Indian job portal with the same script?

The browser steps can be reused, but each portal has its own terms, access behavior, page layout, and rendering needs. Review the rules and verify captures site by site.

Does a successful screenshot mean the job is still open?

No. It records the rendered page at capture time. Confirm listing status with the source site if that matters.

Should I use full-page capture for every listing?

Only when below-the-fold details are needed. Viewport captures are smaller and can be easier to compare; full-page captures preserve more page context.

Can I use these screenshots as a permanent archive?

That depends on the site’s terms and your authorization, as well as your own retention and data-handling needs. Check the applicable rules before storing or redistributing captures.