ScreenshotNeo

BlogHow-to

How to Bulk Screenshot Websites with Python on Ubuntu in India

Capture a list of websites from Ubuntu with Python and Playwright. Learn setup, viewport and full-page captures, failure handling, and when to use a screenshot API.

By the ScreenshotNeo team4 October 20267 min read

To bulk screenshot websites with Python on Ubuntu, use Playwright: install its Python package and Chromium, put one URL per line in a text file, then loop over the URLs and save each page as an image. The same local workflow applies in India; the research sources do not identify an India-specific setting. Network access and the websites you visit can still affect individual captures.

1. Install Python and Playwright on Ubuntu

Use a virtual environment so the project’s Python dependencies stay separate from system packages. These commands assume Python 3 and its virtual environment package are available in your Ubuntu release:

sudo apt update
sudo apt install python3 python3-venv
mkdir website-shots
cd website-shots
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install playwright
python -m playwright install chromium

Playwright’s Python guide documents installing the package and browser binaries. Chromium is launched headlessly by default, which suits terminal and server runs. If installation or launch reports missing system libraries, consult the current Playwright installation guidance for your Ubuntu release rather than applying a guessed package list. If your network requires a proxy to download browsers, Playwright documents browser download proxy configuration.

2. Make a URL list

Create urls.txt in the project directory, with one complete URL per line. Blank lines and comment lines beginning with # are ignored by the script below.

https://example.com
https://www.python.org
# Add more pages here

3. Capture every URL with Python

Save this as capture.py. It creates a screenshots directory, names files by sequence number and hostname, applies a navigation timeout, reports the HTTP status when available, and continues when one URL fails.

from pathlib import Path
from urllib.parse import urlparse
import re

from playwright.sync_api import sync_playwright

INPUT = Path("urls.txt")
OUTPUT = Path("screenshots")
OUTPUT.mkdir(exist_ok=True)


def safe_name(url: str, index: int) -> str:
    host = urlparse(url).netloc or f"page-{index}"
    host = re.sub(r"[^A-Za-z0-9.-]+", "_", host)
    return f"{index:03d}_{host}.png"


urls = [
    line.strip()
    for line in INPUT.read_text(encoding="utf-8").splitlines()
    if line.strip() and not line.lstrip().startswith("#")
]

with sync_playwright() as p:
    browser = p.chromium.launch()
    for index, url in enumerate(urls, start=1):
        page = browser.new_page(viewport={"width": 1440, "height": 1000})
        try:
            response = page.goto(url, wait_until="load", timeout=45000)
            page.screenshot(
                path=str(OUTPUT / safe_name(url, index)),
                full_page=True,
            )
            status = response.status if response else "no HTTP response"
            print(f"OK {status} {url}")
        except Exception as exc:
            print(f"ERROR {url}: {exc}")
        finally:
            page.close()
    browser.close()

Run it from the project directory while the virtual environment is active:

python capture.py

The browser launch, navigation, and screenshot calls use Playwright’s documented Python API. The filename scheme, 45-second timeout, and per-URL error handling are choices in this example; adjust them to your workload. For a quick viewport image instead of the complete scrollable page, change full_page=True to full_page=False or remove the argument.

4. Choose the capture and wait behavior

Viewport or full page

  • Viewport: captures what is visible within the configured 1440 × 1000 viewport. Use this for above-the-fold review or consistent viewport comparisons.
  • Full page: full_page=True captures the full scrollable page as if it fit on a very tall screen. Long pages produce larger images and may take longer to capture.

Playwright also supports capturing an element and returning screenshot bytes instead of writing directly to a path. Use those modes when you need a specific component or want to send image data elsewhere in your program.

Wait for content that loads after navigation

wait_until="load" is a reasonable general starting point, but it does not guarantee that every client-rendered widget, image, or late API response has settled. For dynamic pages, wait for a meaningful page-specific selector before calling screenshot(), for example:

page.goto(url, wait_until="domcontentloaded", timeout=45000)
page.locator("main article").wait_for(state="visible", timeout=15000)
page.screenshot(path="page.png", full_page=True)

Replace main article with a selector that exists on the target site. A fixed sleep may help diagnose a timing issue, but it is not a universal readiness condition and makes every capture wait the same amount.

Validate input and avoid filename collisions

Use full URLs with a scheme such as https://. The sequence prefix in the example keeps two URLs on the same host from overwriting each other. For repeated runs, either clear or version the output directory, or include a run identifier in filenames. If URLs come from outside your own trusted list, validate allowed schemes and hosts before navigating; do not treat arbitrary input as safe to browse.

5. Handle failures and troubleshoot

Symptom Likely cause What to do
Chromium fails to launch with a missing library error System libraries required by the browser are unavailable. Use the Playwright installation guidance for the Ubuntu release and environment in use; install the dependencies it specifies.
Browser download fails Network restrictions, proxy requirements, or an interrupted download. Retry the browser installation when the connection is available. If your network uses a proxy, follow Playwright’s browser download proxy instructions.
Navigation times out The site is slow, unreachable from this machine, or does not finish loading within the chosen timeout. Check the URL and connectivity, then set an appropriate timeout. If the page’s main content appears before all resources finish, consider a different navigation condition plus a selector wait.
Screenshot is blank or missing late content The page may need client-side rendering, a consent action, authentication, or a site-specific readiness condition. Inspect the page state and wait for a meaningful selector. Handle authentication only where you are authorized. A browser automation script does not guarantee that every site can be captured.
Some URLs fail but later ones still run The example catches errors separately for each page. Read the logged URL and exception, fix or skip that input, and rerun. This isolation is intentional so one unavailable site does not stop the batch.
Output files overwrite or have awkward names Custom naming may omit a unique component, or a URL may have no hostname. Keep the sequence prefix or add a run identifier. The helper sanitizes hostname characters and falls back to a numbered page name.

For debugging, keep the failed URL and exception in a log, and open a representative output image to check whether it contains the expected content. Redirects, consent prompts, sign-in pages, and bot challenges can change what the browser sees. Do not assume a successful navigation status means the desired visual content was captured.

6. Make batch runs more reliable and manageable

  • Close each page: the example closes a page after every attempt and closes the browser after the loop.
  • Keep concurrency modest: a sequential loop is simple to debug and avoids issuing a burst of navigations. If you later add parallel workers, account for memory use, target-site load, and the policies that apply to those sites.
  • Choose timeouts deliberately: a larger timeout accommodates slow responses but also delays moving past an unavailable URL. Record failures so a long batch remains auditable.
  • Make reruns safe: choose whether to overwrite, skip existing files, or write each run into a new directory. The example overwrites a matching path.
  • Review the target’s behavior: automated requests can encounter access controls or bot checks. Respect site permissions and applicable organizational rules.

No benchmark in the cited documentation establishes a universally fastest browser tool or safe request rate. Actual batch duration depends on the URLs, network, page content, and capture mode. Full-page shots can require more work than viewport shots because more page area is involved.

7. Playwright or Selenium?

For a new script, Playwright is a direct choice because its Python API includes navigation and screenshot calls and its tooling installs browser builds. Selenium is sensible if your project already uses it or needs one of its supported browsers. Selenium’s documentation describes support for Chrome, Edge, Firefox, Safari, and others, and Selenium Manager handles driver setup for most supported platform and browser combinations. Choose based on the browser and existing project needs; the sources do not provide a speed comparison.

Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. One GET request can return an image or PDF; the API uses the parameter names other screenshot APIs use, which can make a switch easier. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners are accepted and removed along with 60+ known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. An MCP server lets Claude, Cursor, and other MCP clients use screenshot tools. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000. Learn about ScreenshotNeo or sign up for 1,000 free screenshots a month with no card.

Frequently asked questions

Does running the script from India require a different Python setting?

The cited sources document no India-specific setting for this local workflow. Use the same Python and Playwright steps; network and site behavior depend on your environment and the destination.

Can I capture a site that requires sign-in?

Only if you are authorized to access it. The example does not implement login or session persistence; those need site-appropriate handling and secure treatment of credentials.

Can I save JPEG instead of PNG?

Playwright’s screenshot API supports image options such as type and quality. Choose a supported format and quality for your use case, and use the matching file extension.

Will every URL produce a usable screenshot?

No. Network errors, timeouts, access controls, bot checks, or unexpected page states can prevent a useful capture. Log and inspect exceptions and output images rather than assuming every navigation yielded the intended page.