How to Capture Bulk Website Screenshots and Save Them to Google Drive in India
Capture a few pages manually or automate a URL list with Playwright and the Google Drive API. Includes runnable Python code, upload setup, retries, and India-specific scope.
For a few pages, use Google’s Save to Google Drive Chrome extension and save each page from the browser. For a large or recurring URL list, use Playwright to capture each page and the Google Drive API to upload each image file. These are separate steps: Playwright handles browser rendering and capture; Drive API authorization, file naming, uploads, and retries need to be implemented. Google’s Drive API does not support media uploads in batch requests.
The documented browser, Playwright, and Drive API behavior below is general, not specific to India. The sources reviewed do not establish India-specific pricing, quotas, availability changes, legal conditions, or typical network performance.
1. Choose a workflow
| Approach | Best for | What to expect |
|---|---|---|
| Save to Google Drive Chrome extension | A few pages you can open and save yourself | Browser action or context-menu saving; not a documented multi-URL queue |
| Playwright plus Drive API | Repeated captures or a long URL list | A script can loop through URLs, capture files, then upload each file separately |
| Screenshot API | Recurring captures when you want to avoid maintaining browser infrastructure | Captures can be requested over an API; uploading the resulting files to Drive is still a separate integration step |
When the extension is enough
Google’s Save to Google Drive Chrome Web Store listing describes saving web content or browser screenshots through browser actions and context menus. It offers choices such as the entire image, visible image, raw HTML, MHTML, or a Google Doc. It is useful when the pages are already open and a person can save them individually. The listing says Chrome must be signed in to the Google account used for saving; to save to another Drive account, change Chrome profiles. It also says chrome:// pages and Chrome Web Store pages cannot be captured. These are listing descriptions, not independently tested results.
Third-party extensions may advertise full-page capture or Drive folders, but check the listing, requested permissions, privacy policy, and current behavior. Do not assume an extension that automates a particular business workflow is a general-purpose URL crawler.
When to automate
Use a script when you have a reviewed list of URLs, need repeatable filenames, or expect to run the task again. The script below demonstrates the capture stage and writes images locally. Drive authorization and upload are separate; the official Drive API supports media uploads, but its media operations cannot be put into API batch requests.
2. Prepare a URL list and output folder
Put one fully qualified URL per line in urls.txt. Keep the source list under your control: capturing a URL may access public pages, authenticated pages, or pages that trigger site-side protections. Remove duplicates if you do not want multiple copies, and decide how to handle pages that require login before starting.
https://example.com/
https://www.example.org/
https://www.example.net/
Use descriptive image names with an extension. Google’s Drive file creation reference recommends an extension in the file’s name field. The Python example below derives a readable name from the URL and adds a short hash so different URLs with the same path are less likely to collide.
3. Capture the pages with Playwright in Python
Install Playwright and its Chromium browser, then save this script as capture.py. It captures full-page PNGs, limits concurrency, records failures, and continues after an individual page fails. The script is a practical synthesis of Playwright’s documented screenshot building block; it is not a turnkey Google Drive upload workflow.
python -m pip install playwright
python -m playwright install chromium
import asyncio
import hashlib
import re
from pathlib import Path
from urllib.parse import urlparse
from playwright.async_api import async_playwright
URLS_FILE = Path("urls.txt")
OUTPUT_DIR = Path("screenshots")
ERROR_LOG = Path("capture-errors.txt")
CONCURRENCY = 3
NAVIGATION_TIMEOUT_MS = 45_000
def read_urls():
urls = [line.strip() for line in URLS_FILE.read_text(encoding="utf-8").splitlines()]
return list(dict.fromkeys(url for url in urls if url and not url.startswith("#")))
def filename_for(url):
parsed = urlparse(url)
host = parsed.netloc.lower().replace(":", "_") or "page"
path = parsed.path.strip("/") or "home"
readable = re.sub(r"[^a-zA-Z0-9._-]+", "-", path).strip("-._")[:70] or "page"
digest = hashlib.sha256(url.encode("utf-8")).hexdigest()[:10]
return f"{host}-{readable}-{digest}.png"
async def main():
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
semaphore = asyncio.Semaphore(CONCURRENCY)
failures = []
async with async_playwright() as p:
browser = await p.chromium.launch()
async def capture(url):
async with semaphore:
page = await browser.new_page(viewport={"width": 1440, "height": 1000}, device_scale_factor=1)
try:
response = await page.goto(url, wait_until="domcontentloaded", timeout=NAVIGATION_TIMEOUT_MS)
# Some sites continue rendering after DOM content is loaded. A short bounded wait
# lets common client-side rendering settle without waiting indefinitely for all traffic.
await page.wait_for_timeout(1200)
filename = OUTPUT_DIR / filename_for(url)
await page.screenshot(path=str(filename), full_page=True, type="png")
status = response.status if response else "no HTTP response"
print(f"saved {filename} (status: {status})")
except Exception as exc:
failures.append(f"{url}\t{type(exc).__name__}: {exc}")
print(f"failed {url}: {exc}")
finally:
await page.close()
try:
await asyncio.gather(*(capture(url) for url in read_urls()))
finally:
await browser.close()
ERROR_LOG.write_text("\n".join(failures), encoding="utf-8")
print(f"Finished. Failures: {len(failures)}. Log: {ERROR_LOG}")
if __name__ == "__main__":
asyncio.run(main())
python capture.py
Adjust capture behavior
- Viewport: change
widthandheightto control the visible layout. For a mobile view, use a narrower viewport. Device emulation settings can also be used when the target requires a particular device profile. - Full page:
full_page=Truecaptures the scrollable page. For a viewport-only image, omit it or set it toFalse. - Wait strategy:
domcontentloadedis a bounded starting point. Use a selector wait when a known element indicates the page is ready. Network-idle waits can hang on pages with long polling or analytics, so use a timeout and page-specific logic. - Element capture: locate the target with a locator and call its screenshot method when only one component is needed. Check that the element exists and is visible first.
- Format: Playwright screenshots support image output such as PNG and JPEG; choose the format and extension together. PNG is lossless and can be larger; JPEG is generally smaller for photographic content.
- Authentication: for pages behind a login, use an authorized browser context or saved authentication state. Protect any cookies or state files as credentials.
- Dynamic content: wait for a page-specific selector or a known UI state. A fixed delay is simple but cannot guarantee that every site has finished rendering.
4. Upload the image files to Google Drive
Create a Google Cloud project, enable the Google Drive API, configure OAuth for the application type you are building, and authorize the account that owns the destination folder. Store client secrets and refresh tokens securely; do not commit them with the script. Google’s API supports simple media upload, multipart upload, and resumable upload. The documented simple and multipart approaches are for files at or below 5 MB; Google recommends resumable uploads for larger files or unreliable connections, and resumable upload also works for small files.
For a small demonstration, Google’s Python client can create one Drive file per image. Install the client libraries and set up OAuth credentials according to Google’s current Python quickstart for Drive before running this upload stage. This code assumes an authorized service Drive API client is available and a destination folder ID has been configured; it does not embed credentials or pretend that OAuth setup is one-size-fits-all.
python -m pip install google-api-python-client google-auth-oauthlib google-auth-httplib2
from pathlib import Path
from googleapiclient.http import MediaFileUpload
FOLDER_ID = "YOUR_DRIVE_FOLDER_ID"
SCREENSHOT_DIR = Path("screenshots")
def upload_screenshots(service):
for image_path in sorted(SCREENSHOT_DIR.glob("*.png")):
metadata = {
"name": image_path.name,
"mimeType": "image/png",
"parents": [FOLDER_ID],
}
media = MediaFileUpload(str(image_path), mimetype="image/png", resumable=True)
request = service.files().create(
body=metadata,
media_body=media,
fields="id,name,mimeType,size,webViewLink",
)
response = None
while response is None:
_status, response = request.next_chunk()
print(f"uploaded {response['name']} ({response['id']})")
# After completing Google's OAuth flow and constructing the Drive API client:
# upload_screenshots(service)
The example requests resumable uploads so it can handle interruptions more gracefully. For an unattended production job, add structured logs, bounded retries with backoff for transient failures, and a record of each source URL, local file, Drive file ID, and final status. Avoid retrying permanent authorization or invalid-request errors without fixing their cause.
Why Drive API batching does not upload the screenshots
Drive API batch requests can group supported API calls, up to 100 calls in a batch. Google explicitly excludes media uploads and downloads from batch operations. Each screenshot’s media must therefore be uploaded as its own upload operation. You may batch eligible metadata operations, but that does not combine image transfers into one batch upload.
5. Verify the result and rerun safely
- Count the unique input URLs and compare them with the successful local captures and upload records.
- Open a sample of the images and inspect page content, viewport, and full-page height.
- Confirm files are in the expected Drive folder and have the intended names and formats.
- Keep the error log and retry only failed URLs after addressing the cause.
- For repeat runs, decide whether to overwrite, create a dated folder, or detect existing files. Drive file names need not be unique, so filename stability alone does not prevent duplicates.
6. Use cURL, Python, or Node.js for a screenshot API
If you want screenshot generation without installing and maintaining a browser, ScreenshotNeo provides a website screenshot API and MCP server. Its GET endpoint returns an image or PDF. The following one-call example captures a page; the response is an image file that you can then upload to Drive using the API upload approach above. See the ScreenshotNeo API documentation for request options and formats.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
To process a list, loop over the URLs, save each successful response under a stable filename, and then upload those files to Drive. Keep API credentials out of source control, limit concurrency to a level your plan and workflow can handle, and record response status and file mapping. ScreenshotNeo supports bulk capture of up to 100 URLs per call, async jobs with signed webhooks, and a usage API; consult its docs for the relevant request shapes. Its response headers identify the page verdict and whether a capture was billed.
7. Or skip the browser setup
ScreenshotNeo can capture the page through one API call, then your workflow can upload the returned image to Google Drive. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots, and 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See ScreenshotNeo for the service and the API docs for options. Sign up for 1,000 free screenshots a month with no card.
8. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Navigation timeout | The site is slow, blocks automation, or keeps connections open | Use a bounded timeout, wait for a specific ready selector, and retry selectively. Do not wait indefinitely for network idle. |
| Screenshot is blank or incomplete | Client rendering had not finished, a consent overlay obscured content, or content loads on scroll | Wait for the relevant selector or state; scroll or use a page-specific loading strategy where appropriate; inspect the image before uploading. |
| Login page captured instead | The browser context is unauthenticated or the session expired | Authenticate through an approved flow and verify the destination page before capture. Treat saved session state as a secret. |
| 403, CAPTCHA, or bot check | The site restricted automated access | Do not assume the capture can bypass the restriction. Check permission and site policy; use an authorized access method or omit the URL. |
| Google API 401 or 403 | OAuth token is missing, expired, lacks required scope, or the API is not enabled for the project | Complete the Drive OAuth setup, enable the API, and authorize the destination account with appropriate access. |
| File is in the wrong folder | The folder ID is wrong or the authorized account cannot access it | Verify the folder ID and account access; inspect the returned file metadata. |
| Upload interrupted | Connection dropped during transfer | Use resumable uploads, preserve per-file status, and retry transient failures with bounded backoff. |
| Duplicate images after rerun | Drive permits files with identical names | Record Drive file IDs or use a deterministic overwrite/deduplication policy before creating another file. |
| Extension cannot capture a page | It may be a restricted Chrome page or a Chrome Web Store page | Use a normal web page or an automated capture workflow; the cited Google extension listing says these restricted page types cannot be captured. |
9. Performance, reliability, and cost
- Concurrency: more parallel browser pages may shorten a run, but consume more memory and can increase load on destination sites. Start conservatively and tune for your machine and permissions; no universal benchmark is established here.
- Image size: full-page PNGs can be large, especially for long pages. JPEG may reduce size for photographic pages, while PNG preserves sharp text and edges. Larger files take longer to upload.
- Reliability: keep capture and upload status separately. A successful screenshot does not prove the Drive upload succeeded. Retry transient navigation or upload failures selectively and retain enough metadata to resume.
- Drive upload mode: Google describes simple and multipart upload for files up to 5 MB and recommends resumable uploads for larger files or unreliable connections. Resumable mode also supports small files.
- Batching: API batching has a 100-call limit for supported operations, but does not support media uploads or downloads. It is not a shortcut for transferring a directory of screenshot images.
- Costs and regional scope: the reviewed sources do not establish India-specific Google Drive pricing, quotas, or network performance. Check current account and API terms for your use. ScreenshotNeo offers 1,000 shots per month free without a card; paid plans are Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan.
10. Frequently asked questions
Can I save a whole URL list with Google’s Chrome extension?
The cited listing describes browser-action and context-menu saving, not a multi-URL batch queue. For a URL list, automate capture and upload or use a service that supports bulk capture.
Can Drive API batch requests upload all screenshot images?
No. Google documents that media uploads and downloads are not supported in batch requests; upload each image as its own media operation.
Does the workflow change for someone in India?
The sources reviewed document general Google and Playwright behavior and do not establish an India-specific variant. Confirm current account, API, and service terms for your setup.
Can I capture pages that require a login?
Yes, if you are authorized and provide the browser or service with an appropriate authenticated session. Keep cookies and tokens private, and account for session expiry.


