ScreenshotNeo

BlogHow-to

How to screenshot a batch of URLs and upload them to Google Drive

Capture a list of URLs with Playwright, save uniquely named screenshots, and upload them to Google Drive with resumable recovery for larger files.

By the ScreenshotNeo team4 October 202610 min read

To screenshot a batch of URLs and put the images in Google Drive, use a browser automation script to visit each URL and save a uniquely named screenshot, then upload each image with the Google Drive API. The example below uses Playwright for capture and the Drive API’s resumable upload protocol for transfer. Playwright supplies the screenshot operations; the URL loop, file naming, waits, retries, and Drive upload are workflow code you provide.

This separates the job into two stages: capture one image per URL, then transfer those files. Drive API batch requests do not combine media uploads into one operation.

1. Choose the capture and upload approach

Decision Choose this when Trade-off
Viewport screenshot You want the visible browser viewport. Content below the fold is not included.
Full-page screenshot You want the full scrollable page in one image. Images can be very tall, and page layouts may render differently during a full-page capture.
Drive web interface You want a manual upload after capture. It is less suitable for an unattended repeatable workflow.
Drive API You want uploads handled by the script. You need Google authentication and a Drive API upload flow.

Google documents simple upload for small files up to 5 MB without metadata, multipart upload for small files up to 5 MB with metadata in one request, and resumable upload for larger files or connections that may be interrupted. Resumable uploads also work for small files, but add a request to start each upload. The runnable example uses resumable uploads so it has a defined recovery path.

2. Set up the project and credentials

Install Playwright and Google’s API client for Python:

python -m venv .venv
source .venv/bin/activate
python -m pip install playwright google-api-python-client google-auth google-auth-oauthlib
python -m playwright install chromium

Create OAuth credentials for a desktop application in Google Cloud, enable the Google Drive API for the project, and save the downloaded credentials as credentials.json beside the script. On the first run, the script opens a browser for authorization and saves a refreshable token in token.json. Keep both files private. This example requests the Drive file scope and creates screenshots in the authorizing user’s Drive.

Save the following as capture_and_upload.py. Replace the sample URLs with the pages to capture. Set FULL_PAGE to True for full-page images.

3. Capture and upload every URL

import re
import time
from pathlib import Path
from urllib.parse import urlparse

from playwright.sync_api import sync_playwright
from google.auth.transport.requests import Request
from google.oauth2.credentials import Credentials
from google_auth_oauthlib.flow import InstalledAppFlow
from googleapiclient.discovery import build
from googleapiclient.http import MediaFileUpload

URLS = [
    "https://example.com/",
    "https://www.python.org/",
]
OUTPUT_DIR = Path("screenshots")
DRIVE_FOLDER_ID = None  # Optional: set to an existing Drive folder ID.
FULL_PAGE = False
NAVIGATION_TIMEOUT_MS = 45_000
SCOPES = ["https://www.googleapis.com/auth/drive.file"]


def safe_name(value: str) -> str:
    value = re.sub(r"[^A-Za-z0-9._-]+", "-", value).strip("-._")
    return value[:80] or "page"


def screenshot_name(index: int, url: str) -> str:
    parsed = urlparse(url)
    host = safe_name(parsed.netloc or "page")
    path = safe_name(parsed.path.strip("/").replace("/", "-") or "home")
    return f"{index:03d}-{host}-{path}.png"


def drive_service():
    creds = None
    token_path = Path("token.json")
    if token_path.exists():
        creds = Credentials.from_authorized_user_file(str(token_path), SCOPES)
    if not creds or not creds.valid:
        if creds and creds.expired and creds.refresh_token:
            creds.refresh(Request())
        else:
            creds = InstalledAppFlow.from_client_secrets_file(
                "credentials.json", SCOPES
            ).run_local_server(port=0)
        token_path.write_text(creds.to_json(), encoding="utf-8")
    return build("drive", "v3", credentials=creds)


def upload_file(service, path: Path):
    metadata = {"name": path.name}
    if DRIVE_FOLDER_ID:
        metadata["parents"] = [DRIVE_FOLDER_ID]
    media = MediaFileUpload(
        str(path), mimetype="image/png", resumable=True, chunksize=1024 * 1024
    )
    request = service.files().create(
        body=metadata, media_body=media, fields="id,name,webViewLink"
    )
    response = None
    while response is None:
        _status, response = request.next_chunk(num_retries=3)
    return response


def main():
    OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
    service = drive_service()
    failures = []

    with sync_playwright() as playwright:
        browser = playwright.chromium.launch(headless=True)
        page = browser.new_page(viewport={"width": 1440, "height": 1000})
        page.set_default_navigation_timeout(NAVIGATION_TIMEOUT_MS)

        for index, url in enumerate(URLS, start=1):
            path = OUTPUT_DIR / screenshot_name(index, url)
            try:
                response = page.goto(url, wait_until="domcontentloaded")
                # Let the page settle briefly after initial HTML arrives.
                page.wait_for_timeout(1000)
                page.screenshot(path=str(path), full_page=FULL_PAGE)
                result = upload_file(service, path)
                print(f"Uploaded {path.name}: {result.get('id')}")
            except Exception as exc:
                failures.append((url, str(exc)))
                print(f"Failed {url}: {exc}")
        browser.close()

    if failures:
        print("\nFailures:")
        for url, message in failures:
            print(f"- {url}: {message}")
        raise SystemExit(1)


if __name__ == "__main__":
    main()

Run it with:

python capture_and_upload.py

The script writes a local PNG before attempting the Drive upload. If an upload fails, that file remains available to retry without revisiting the page. Each run uses an index and sanitized host/path in the filename to avoid collisions within that URL list. Query parameters are intentionally omitted from filenames, so if two entries have the same host and path but different query values, the index keeps their names distinct.

Adjust page readiness

domcontentloaded waits for the initial document to be parsed; it does not guarantee that every image, font, animation, or client-rendered component is ready. Choose a readiness condition that matches each target site:

  • Wait for a specific element with page.locator("main").wait_for() when that element signals usable content.
  • Use page.goto(url, wait_until="load") when the page’s load event is a better fit.
  • Use page.goto(url, wait_until="networkidle") only where the page becomes idle; analytics, polling, or streaming requests can prevent it from settling.
  • For lazy-loaded content, scroll through the page before taking a full-page screenshot if the site loads images only when they approach the viewport.

There is no universal wait that makes every site visually stable. Validate the capture on the pages you care about, and use a selector or site-specific delay when necessary. The example’s one-second pause is a tunable starting point, not a guarantee.

4. Configure Google Drive uploads

Upload to a specific folder

Set DRIVE_FOLDER_ID to the ID of a folder the authorized account can access. The folder ID is the identifier in that folder’s Drive URL. The example adds it as the file’s parent. If the value is unset, files are created in the user’s Drive root.

Choose an upload method

Method Use it for Drive behavior
Simple Small media, up to 5 MB, with no file metadata to send. Uploads the content only.
Multipart Small media, up to 5 MB, with metadata such as the filename in one request. Sends metadata and media together.
Resumable Larger media or a connection that may fail during transfer. Starts a session, then transfers content in chunks that can be checked and resumed.

In the Python client example, resumable=True selects resumable transfer and chunksize controls the client chunk size. The example uses 1 MiB chunks. Choose a value supported by the client and appropriate for your environment; larger chunks mean fewer transfer requests but more data per request to retry after an interruption.

Recover a resumable upload

The Google client library retries some requests, but a production worker that must survive process restarts should persist enough state to recover each transfer, including the upload session URI. Google’s protocol uses PUT requests after session initiation. To check progress after an interruption, query the session and use the server’s reported Range to determine which bytes arrived. A 308 Resume Incomplete response means the upload is not finished. Do not assume the last chunk you sent was accepted. A resumable session URI expires after one week, so expired sessions must be started again.

For a small script, retaining the local screenshot and rerunning the upload may be sufficient. If you need guaranteed recovery after a machine restart, add a durable job record for each file and implement Google’s session status and resume steps against the current Drive upload documentation.

5. Upload through the Drive interface instead

If you prefer to capture locally and upload interactively, open Drive and select the screenshot files. For a browser-driven upload, Playwright can assign multiple local paths to an HTML file input in one call:

await page.locator('input[type="file"]').set_input_files([
    "screenshots/001-example-com-home.png",
    "screenshots/002-python-org-home.png",
])

This works only if the interface exposes a compatible file input and accepts multiple files. The exact Drive interface flow can change, so inspect the current page and verify the upload completes. For repeatable unattended uploads, the Drive API avoids relying on UI selectors.

6. Handle failures and make the batch reliable

  • Keep per-URL results: record the URL, local filename, capture status, upload status, and error. Continue after one URL fails, as the example does.
  • Retry selectively: retry transient navigation or network failures with a small bounded retry count and backoff. Do not retry invalid URLs or persistent access-denied pages indefinitely.
  • Make reruns safe: keep deterministic filenames or store completed Drive file IDs in a job manifest. Otherwise a rerun can create duplicate Drive files.
  • Keep the local copy: remove it only after Drive confirms a successful upload and your retention policy allows deletion.
  • Limit concurrency: start sequentially, as in the example. If you add parallel workers, tune them against target-site behavior, browser memory, network capacity, and Drive quota responses. The reviewed sources do not establish a universal safe concurrency level.
  • Protect credentials: do not commit OAuth credentials or tokens to source control. Restrict access to the machine and files that hold them.

7. Troubleshooting

Symptom Likely cause Fix
Playwright says the browser executable is missing. The installed Playwright package has no browser binary installed. Run python -m playwright install chromium in the same environment.
Navigation times out. The site is slow, never reaches the chosen wait state, or keeps network activity open. Increase the timeout for that site, use a more suitable wait condition, or wait for a meaningful selector instead of requiring network idle.
The image is blank or incomplete. The page may render after the chosen readiness event, require authentication, or load content lazily. Wait for the page’s content selector, scroll to trigger lazy loading, and inspect the page state before capture.
Several files have unexpected names or overwrite. Names were derived from non-unique page data or an old naming rule. Include a stable index or unique record ID and sanitize characters that are unsuitable for filenames.
Google authorization fails. Drive API is not enabled, credentials are for the wrong OAuth client type, or the token has stale scopes. Enable the API, use desktop-app OAuth credentials, remove the old token after changing scopes, then authorize again.
Drive returns a permission error. The authorizing account cannot create files in the selected folder, or the folder ID is wrong. Confirm the account and folder access, then check the folder ID.
An upload is interrupted or returns 308. A resumable transfer is incomplete; 308 is a continuation status. Query the session status and resume from the server-reported byte range. Restart if the session URI has expired.
Rerunning the script creates duplicates. Drive file creation does not automatically replace a previous file with the same name. Track file IDs in a manifest and update existing files, or define a deliberate duplicate policy.

8. Performance and cost considerations

Browser launch and page navigation usually dominate the work in a small sequential job; upload time depends on image size and connection speed. Reusing one browser and page, as the example does, avoids launching a browser for every URL. Full-page images can be substantially larger than viewport images, so select capture scope based on what the record needs. The research sources provide no universal timing or concurrency benchmark.

Drive API batch requests can combine eligible API calls to reduce connection overhead, with a documented maximum of 100 calls per batch. They do not support batching media uploads or downloads. Each screenshot’s media therefore remains its own upload operation. Account for browser compute, network transfer, and Drive API quota in your environment; exact costs depend on your infrastructure and account.

Or skip the browser setup

ScreenshotNeo can return a screenshot from one GET request; see the API documentation. Capture each URL with your API key, then upload the returned bytes or saved file to Drive using the API workflow above. Repeat the call for each URL in your list.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie banners, popups, and chat widgets before capture. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan.

Sign up free for 1,000 screenshots a month, no card required.

FAQ

Can I upload all screenshots in one Google Drive API batch request?

No. Drive API batching can combine eligible API calls, but media uploads and downloads are not supported in batch requests. Upload each image as its own media operation.

Can one failed URL stop the rest of the batch?

It does not have to. Catch errors per URL, save the failure in a report, and continue with the remaining entries. The example follows that pattern and exits with an error status after the batch if any item failed.

Can I capture screenshots directly into memory?

Yes. Playwright can return screenshot bytes in a buffer. Saving to a file is convenient for resumable Drive uploads and recovery; a buffer is useful when passing content to another in-memory step.

How long can I wait before resuming a Drive upload?

A resumable session URI expires after one week. Check upload progress with the server before resending content, and initiate a new session if the old one expired.

Sources