How to Capture Bulk Screenshots of Websites as JPEG Files
Capture a list of websites as JPEGs with Playwright, shot-scraper, or R webshot. Set quality, full-page behavior, filenames, and retries for reliable batches.
To capture bulk website screenshots as JPEG files, prepare a list of URLs, open each page in a browser automation tool, wait for it to render, and save a JPEG with a predictable filename. Playwright is a flexible choice when you need code-level control over JPEG quality, viewport size, and full-page capture. For a repeatable configuration-file workflow, use shot-scraper; if your project uses R, the webshot package accepts URL vectors and JPEG output paths.
This guide uses Playwright with Node.js for a complete batch workflow. It also covers Playwright in Python, cURL examples for a hosted API, and the documented shot-scraper and R options. Do not assume a basic navigation-and-screenshot loop will handle every site’s lazy loading, login, consent dialog, or bot defense. Start with a small sample and check the resulting files.
1. Choose a batch workflow
Batching and file format are separate decisions: a tool can accept many URLs while each individual capture still needs to be configured as JPEG.
| Route | Choose it when | Batch and JPEG approach |
|---|---|---|
| Playwright API or CLI | You want browser automation, custom waits, or error reporting in code. | Loop over URLs with the API and set type: 'jpeg'; the CLI supports JPEG filenames and full-page capture. |
| shot-scraper | You want URL jobs declared in a YAML file and run as a batch. | Define one URL/output entry per job and use shot-scraper multi. Its documentation examples name PNG, so check your installed version’s format behavior before relying on a JPEG option or extension. |
| R webshot | Your existing workflow is in R. | Pass URL vectors and corresponding output paths ending in .jpeg. Check the installed package version before using commands verbatim. |
Compare tools by whether you prefer code or configuration, how much control you need over page loading and interaction, whether the whole page or just the viewport is needed, and how you want filenames and failures recorded. The cited documentation does not establish a universal performance winner.
2. Prepare URLs and output names
- Put one fully qualified URL per input record. Include the scheme, such as
https://, and validate that each URL is non-empty. - Create a destination directory and plan stable filenames from a site or page identifier. Avoid relying only on list order: a reordered input should not silently change which page a file represents.
- Sanitize names for your operating system and ensure they are unique. For repeated domains, add a short page identifier or sequence number.
- Keep a mapping from each source URL to its output path and result status. This makes failed URLs and retries traceable.
A viewport screenshot captures the visible browser area. A full-page screenshot captures the page’s scrollable content, but very long pages can produce large images and may need extra time or special handling for lazy-loaded content.
3. Capture a batch with Playwright in Node.js
Install Playwright and its browser in a Node.js project using the official installation instructions. The example below reads a JSON array of objects with name and url fields, takes one browser page through the list, saves JPEGs, and writes a CSV-style failure log. It makes three attempts per URL and records the final error rather than aborting the whole batch.
import { chromium } from 'playwright';
import { mkdir, readFile, writeFile } from 'node:fs/promises';
import path from 'node:path';
const jobs = JSON.parse(await readFile('urls.json', 'utf8'));
const outputDir = 'screenshots';
const fullPage = false; // Set true to capture the scrollable page.
const quality = 80; // Playwright JPEG quality range: 0–100.
const viewport = { width: 1440, height: 900 };
const failures = [];
function safeName(value) {
const cleaned = String(value).toLowerCase()
.replace(/[^a-z0-9]+/g, '-').replace(/^-|-$/g, '');
if (!cleaned) throw new Error('Each job needs a usable name');
return cleaned;
}
await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch();
try {
const page = await browser.newPage({ viewport, deviceScaleFactor: 1 });
for (const job of jobs) {
const filename = path.join(outputDir, `${safeName(job.name)}.jpg`);
let lastError;
for (let attempt = 1; attempt <= 3; attempt++) {
try {
const response = await page.goto(job.url, {
waitUntil: 'load', timeout: 45_000
});
if (response && response.status() >= 400) {
throw new Error(`HTTP ${response.status()}`);
}
// Add a page-specific wait here if important content renders later.
await page.screenshot({
path: filename, type: 'jpeg', quality, fullPage
});
lastError = undefined;
break;
} catch (error) {
lastError = error;
if (attempt < 3) await new Promise(r => setTimeout(r, 1000 * attempt));
}
}
if (lastError) failures.push({ url: job.url, error: String(lastError) });
}
await writeFile('failures.json', JSON.stringify(failures, null, 2));
} finally {
await browser.close();
}
Create urls.json like this:
[
{ "name": "example-home", "url": "https://example.com/" },
{ "name": "example-about", "url": "https://example.com/about" }
]
The script reuses one page sequentially, which keeps concurrency and memory use low. Use separate pages or browser contexts for parallel work only after deciding how you will limit concurrency and isolate cookies or other session state. The retries are for transient failures; repeated access denial or bot checks should be investigated rather than retried indefinitely.
Wait for the content that matters
waitUntil: 'load' waits for the page load event, but it does not guarantee that client-rendered content, web fonts, animations, or below-the-fold lazy images are ready. If the page has a stable target element, wait for it before taking the screenshot:
await page.goto(job.url, { waitUntil: 'domcontentloaded', timeout: 45_000 });
await page.locator('main').waitFor({ state: 'visible', timeout: 15_000 });
await page.screenshot({ path: filename, type: 'jpeg', quality: 80 });
For a deliberate fixed pause, use await page.waitForTimeout(1500), but prefer a meaningful selector when possible. Network-idle waits can hang or be inappropriate on pages with persistent connections. A full-page option does not necessarily trigger every site’s lazy-loaded content; scrolling or page-specific interaction may be needed.
4. JPEG quality, dimensions, and full-page settings
- Format: Set
type: 'jpeg'explicitly in the API. The filename extension alone can also communicate output format, but explicit configuration makes intent clear. - Quality: Playwright documents a JPEG quality range of 0–100 and a default of 80. Lower quality usually reduces file size while increasing compression artifacts; choose based on the text, fine lines, and image detail your use case must preserve.
- Viewport: Set a consistent width and height when comparing pages. Responsive layouts change with viewport dimensions.
- Scale: Playwright’s
cssscale produces one image pixel per CSS pixel.devicescale can produce more pixels on high-DPI screens and larger output files. Use a consistent scale across a batch. - Full page: Set
fullPage: trueto capture the full scrollable page. Leave it false for the current viewport. Extremely long pages can consume more memory and yield tall images that are awkward to inspect. - Transparency: JPEG has no alpha channel. If transparent output is required, select a format that supports it rather than JPEG.
For example, these are the key Playwright settings:
await page.screenshot({
path: 'screenshots/example.jpg',
type: 'jpeg',
quality: 80,
fullPage: true,
scale: 'css'
});
See the official Playwright Page API and Playwright screenshot CLI documentation for supported screenshot and CLI options.
5. Other do-it-yourself routes
Playwright in Python
Install the Playwright Python package and its browser following the official Python setup guide. This runnable asynchronous example reads the same JSON format, writes JPEGs, and reports failures:
import asyncio
import json
import re
from pathlib import Path
from playwright.async_api import async_playwright
async def main():
jobs = json.loads(Path('urls.json').read_text())
out = Path('screenshots')
out.mkdir(parents=True, exist_ok=True)
failures = []
async with async_playwright() as p:
browser = await p.chromium.launch()
try:
page = await browser.new_page(viewport={"width": 1440, "height": 900})
for job in jobs:
name = re.sub(r'[^a-z0-9]+', '-', job['name'].lower()).strip('-')
if not name:
failures.append({"url": job['url'], "error": "invalid name"})
continue
try:
response = await page.goto(job['url'], wait_until='load', timeout=45000)
if response and response.status >= 400:
raise RuntimeError(f"HTTP {response.status}")
await page.screenshot(path=str(out / f'{name}.jpg'), type='jpeg', quality=80, full_page=False)
except Exception as exc:
failures.append({"url": job['url'], "error": str(exc)})
Path('failures.json').write_text(json.dumps(failures, indent=2))
finally:
await browser.close()
asyncio.run(main())
shot-scraper configuration
shot-scraper documents a YAML jobs file and a multi command for multiple URL/output entries. Its documentation examples identify PNG; verify your installed version’s accepted image formats and options before treating this as a JPEG recipe.
- url: https://example.com/
output: screenshots/example-home.png
- url: https://example.com/about
output: screenshots/example-about.png
Run the jobs with:
shot-scraper multi jobs.yml
See the shot-scraper documentation for multi-job configuration, selecting jobs by output, selectors, and scaling.
R webshot
The webshot package documentation describes URL vectors, JPEG filenames, viewport dimensions, selectors, and a configurable delay. This example follows that interface; check your installed package version and setup requirements before running it.
library(webshot)
urls <- c('https://example.com/', 'https://example.org/')
files <- c('screenshots/example.jpg', 'screenshots/example-org.jpg')
dir.create('screenshots', showWarnings = FALSE)
webshot(urls, file = files, vwidth = 1440, vheight = 900, delay = 2)
The delay gives pages additional time to render assets; choose it based on the pages and inspect the output. Refer to the webshot package documentation.
6. cURL and hosted screenshot capture
A local Playwright batch gives you direct browser control. A hosted API can be simpler when you do not want to install and maintain browser binaries on the machine running captures. For a single URL, ScreenshotNeo returns a screenshot from one GET request. The API accepts a URL and supports image output options; see the ScreenshotNeo API documentation for current parameters and JPEG configuration.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
The supplied example saves a WebP response. For JPEG, request the JPEG output option documented by ScreenshotNeo and use a .jpg output filename. Repeat the request for each URL in your input list, using safe filenames and recording the HTTP result for each request. Do not put a real API key in a shared script or commit it to source control.
Python request
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
For a batch, loop over your URL/name records, check each response, and write successful output under its mapped filename. The example uses WebP as supplied; request JPEG in the API parameters documented by ScreenshotNeo docs when JPEG is required.
Node.js request
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: HTTP ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
For repeated captures, create one request per input record and retain a URL-to-file mapping. The specific API examples above use WebP to match the supplied call; consult ScreenshotNeo’s docs for the JPEG format parameter and all available request options.
7. Make a large batch reliable
- Begin with a sample: Capture a few representative pages and inspect dimensions, text legibility, missing images, compression artifacts, overlays, and filenames.
- Limit concurrency: Sequential capture is easier on memory and makes failures easier to isolate. If you add parallel pages, cap the number in flight and account for browser memory and the target sites’ load.
- Retry selectively: Retry timeouts and transient network failures with a small limit and backoff. Do not repeatedly retry authentication failures, access restrictions, or bot defenses.
- Keep a manifest: Record input URL, output path, timestamp, status, and error. Write to a temporary filename and rename only after a successful capture so a failed attempt does not leave a partial file that looks complete.
- Control browser state: Reusing a page can carry cookies and local storage between sites. Create a fresh browser context per job where isolation matters, or intentionally reuse an authenticated context where permitted.
- Respect site access: Authentication, bot defenses, and site terms are specific to each site. Confirm you are allowed to capture the pages and use credentials only through a secure workflow.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Output is PNG or another format | The format was inferred from an extension or the chosen tool’s defaults. | Set Playwright type: 'jpeg' explicitly and inspect the actual file format. For shot-scraper, verify installed-version support before using JPEG. |
| Screenshot is cut off | The capture used viewport mode. | Enable fullPage in the API or the CLI’s full-page option. |
| Images or page sections are missing | Content had not loaded, is lazy-loaded below the fold, or depends on interaction. | Wait for a stable selector, add a justified delay, scroll or interact as the page requires, and capture a sample to confirm. |
| Page is blank or incomplete | Navigation failed, content is client-rendered, or the site returned an access challenge. | Check navigation status and logs, wait for the relevant content, and handle access restrictions through an authorized route. |
| Repeated timeout | The site is slow, never reaches the chosen load condition, or has persistent network activity. | Set a suitable timeout and wait for a specific selector or a less strict navigation event. Retry only transient cases. |
| JPEG looks soft or blocky | Quality is too low, or the screenshot has too few pixels for the intended display size. | Raise JPEG quality, review viewport dimensions, and consider device scale while accounting for the larger output. |
| Files overwrite one another | Names collide after sanitization or outputs use a shared name. | Check uniqueness before capture and include stable page identifiers in filenames. |
| Batch stops on one bad URL | An exception is not caught per job. | Catch errors inside the loop, append the URL and error to a manifest, and continue with later jobs. |
9. Performance, reliability, and cost
Capture time depends on browser startup, page rendering, network conditions, chosen wait behavior, and whether you capture a full page. The research sources provide no benchmark that predicts throughput across sites. A persistent browser with sequential page reuse avoids relaunching for each URL, while bounded concurrency can increase throughput at the cost of more memory and less isolation. Measure a representative sample from your own URL set before estimating completion time.
For local capture, account for the machine running the browser, browser installation, storage for output files, and the time spent handling failures. Smaller viewport captures, CSS scale, and lower JPEG quality can reduce output size; higher device scale and full-page images can increase it. The right setting depends on how much detail you need to retain.
Hosted capture shifts browser operations to an API and may suit managed workloads. ScreenshotNeo’s stated pricing is 1,000 shots per month free with no card; Starter is $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with page verdict and billed status in response headers. See ScreenshotNeo and its docs for API details and current output parameters.
Or skip the browser setup
Use ScreenshotNeo when you would rather send a request than install and maintain a browser. The service removes cookie banners, newsletter popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
This supplied one-call example saves WebP. To make a JPEG batch, consult the ScreenshotNeo docs for the JPEG output parameter, then repeat the request for each URL and save under a unique .jpg filename. Sign up for 1,000 free screenshots a month with no card.
10. FAQ
Can I use JPEG for every website screenshot?
JPEG is suitable when a compressed photographic image is acceptable. It does not preserve transparency and can show artifacts around text or sharp edges; choose quality and format based on the output’s purpose.
Should I capture the viewport or the full page?
Use viewport capture for a consistent above-the-fold view. Use full-page capture when below-the-fold content matters, and account for lazy loading and very tall output.
Is there a best JPEG quality value?
There is no universal value. Playwright’s default is 80; inspect representative pages at the intended display size and adjust for legibility and file size.
Can a batch guarantee identical results each time?
No. Websites can change between captures, and dynamic content, timing, personalization, and access controls can affect the result. Keep settings and timestamps in the capture manifest.


