How to Build a Website Screenshot Downloader With JavaScript
Build a JavaScript screenshot downloader with Playwright: capture a viewport, full page, or element, save the image, and plan for safe server use.
A JavaScript website screenshot downloader opens a URL in a real browser, waits for the page to render, captures the viewport, full page, or a selected element, then saves the image or returns its bytes. Playwright is a practical choice for a new implementation; Puppeteer offers a similar workflow. Install both the JavaScript package and its browser binary.
1. Create a local Playwright downloader
The example below accepts a URL and optional output path, validates that the input is an HTTP or HTTPS URL, captures a full-page PNG, and closes the browser even if navigation or capture fails. It uses domcontentloaded as a starting wait condition; change it when the target page needs more time to render.
import { chromium } from 'playwright';
import { pathToFileURL } from 'node:url';
function parseArgs(args) {
const [rawUrl, output = 'screenshot.png'] = args;
if (!rawUrl) throw new Error('Usage: node screenshot.js <url> [output.png]');
const url = new URL(rawUrl);
if (!['http:', 'https:'].includes(url.protocol)) {
throw new Error('Only http: and https: URLs are supported');
}
return { url: url.href, output };
}
export async function downloadScreenshot(rawUrl, output = 'screenshot.png') {
const { url } = parseArgs([rawUrl, output]);
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage({
viewport: { width: 1280, height: 800 },
deviceScaleFactor: 1
});
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 });
await page.screenshot({ path: output, type: 'png', fullPage: true });
return output;
} finally {
await browser.close();
}
}
if (process.argv[1] && import.meta.url === pathToFileURL(process.argv[1]).href) {
try {
const { url, output } = parseArgs(process.argv.slice(2));
console.log(`Saved ${await downloadScreenshot(url, output)}`);
} catch (error) {
console.error(error.message);
process.exitCode = 1;
}
}
Set up a project and install the package and Chromium browser:
mkdir screenshot-downloader
cd screenshot-downloader
npm init -y
npm pkg set type=module
npm install playwright
npx playwright install chromium
Save the code as screenshot.js, then run:
node screenshot.js https://example.com ./example.png
Playwright’s browser binaries and system dependencies are separate from the package installation. Consult the Playwright library installation guide for environment-specific setup and the ScreenshotNeo API documentation for the hosted alternative below.
2. Choose what to capture and how to save it
Viewport or full page
The default screenshot captures the visible viewport. For a single tall image of the scrollable document, set fullPage: true. Very long pages can create large images, use more memory, and take longer to encode. For pages with lazy-loaded content, scrolling may be needed to trigger loading before capture; a full-page option alone does not guarantee every site loads every image.
// Visible viewport
await page.screenshot({ path: 'viewport.png', type: 'png' });
// Entire scrollable page
await page.screenshot({ path: 'full-page.png', type: 'png', fullPage: true });
Capture one element
Use a locator when the downloader should save a component, card, chart, or other specific region. Make sure the selector matches exactly what you intend; a missing or ambiguous match can fail or capture an unintended element.
const card = page.locator('.product-card').first();
await card.waitFor({ state: 'visible', timeout: 10_000 });
await card.screenshot({ path: 'card.png', type: 'png' });
Format, quality, and scale
PNG preserves sharp edges and is lossless. JPEG and WebP can reduce file size with lossy compression. Quality applies to lossy formats, not PNG. Check the current Playwright screenshot documentation for supported options in your installed version.
// JPEG with lossy quality setting
await page.screenshot({ path: 'page.jpg', type: 'jpeg', quality: 82, fullPage: true });
// A higher device scale factor produces more pixels per CSS pixel
const retinaPage = await browser.newPage({
viewport: { width: 1280, height: 800 },
deviceScaleFactor: 2
});
A device scale factor of 2 roughly doubles each dimension and quadruples the pixel count relative to scale 1, increasing image size and memory use. For reproducible output, explicitly set viewport, scale, color scheme, locale, and other rendering inputs that matter to your use case.
Wait for the right page state
domcontentloaded means the initial document has been parsed; it does not mean every image, font, or client-rendered widget is ready. load waits for load-event resources. networkidle can suit some pages, but analytics, streaming connections, and long-lived requests may prevent it from completing. Prefer waiting for a known page-specific signal when available.
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 });
await page.locator('main article').waitFor({ state: 'visible', timeout: 10_000 });
await page.screenshot({ path: 'article.png', fullPage: true });
For a simple fixed delay, use await page.waitForTimeout(1500) sparingly: it may waste time on fast pages and still be too short on slow ones. Page screenshots can also be returned as bytes instead of written to disk:
const imageBytes = await page.screenshot({ type: 'png', fullPage: true });
// For example, return imageBytes from an HTTP handler with Content-Type: image/png
3. Offer a downloader over HTTP
A local command-line script only visits URLs you choose. An HTTP service that accepts arbitrary destinations adds a serious security boundary: the browser can follow redirects and make subrequests from your server’s network. URL parsing alone does not prevent server-side request forgery (SSRF).
This small Express example shows request validation, timeout handling, and returning PNG bytes. It is a teaching baseline, not a complete public-service security policy. Before exposing such an endpoint, add destination and redirect controls, network egress restrictions, concurrency and rate limits, request and response size limits, authentication, and isolated browser processes or containers. Block loopback, private, link-local, and cloud metadata destinations at the network boundary, accounting for DNS resolution and changes across redirects. Restrict outbound network access so browser subrequests cannot reach protected infrastructure.
import express from 'express';
import { chromium } from 'playwright';
const app = express();
app.use(express.json({ limit: '2kb' }));
app.post('/screenshot', async (req, res) => {
let target;
try {
target = new URL(req.body?.url);
if (!['http:', 'https:'].includes(target.protocol)) throw new Error();
} catch {
return res.status(400).json({ error: 'Provide a valid http or https URL' });
}
let browser;
try {
browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1280, height: 800 } });
await page.goto(target.href, { waitUntil: 'domcontentloaded', timeout: 20_000 });
const png = await page.screenshot({ type: 'png', fullPage: false });
res.type('png').send(png);
} catch (error) {
res.status(502).json({ error: 'Could not capture the requested page' });
} finally {
await browser?.close();
}
});
app.listen(3000, () => console.log('Listening on port 3000'));
Install the HTTP dependency with npm install express, save this as server.js, and run node server.js. Submit a request with:
curl -X POST http://localhost:3000/screenshot \
-H 'Content-Type: application/json' \
-d '{"url":"https://example.com"}' \
-o page.png
Do not treat this sample’s scheme check as sufficient protection for a public service. A malicious destination can redirect, resolve to a protected address, or load additional URLs through scripts and embedded resources. Browser isolation helps, but the network policy must also constrain where the process can connect.
4. Use the downloader from cURL, Python, or Node.js
These clients call the sample HTTP endpoint above. They are ways to use the JavaScript service; the browser rendering still happens in the Node.js process.
cURL
curl -X POST http://localhost:3000/screenshot \
-H 'Content-Type: application/json' \
-d '{"url":"https://example.com"}' \
--fail-with-body \
-o page.png
Python
import requests
response = requests.post(
'http://localhost:3000/screenshot',
json={'url': 'https://example.com'},
timeout=45,
)
response.raise_for_status()
with open('page.png', 'wb') as image_file:
image_file.write(response.content)
Node.js
const response = await fetch('http://localhost:3000/screenshot', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ url: 'https://example.com' }),
signal: AbortSignal.timeout(45_000),
});
if (!response.ok) throw new Error(`Screenshot request failed: ${response.status}`);
const bytes = Buffer.from(await response.arrayBuffer());
await import('node:fs/promises').then(({ writeFile }) => writeFile('page.png', bytes));
5. Install and operate it reliably
- Match package and browser versions. Install browser binaries through the Playwright CLI for the version in your project. A package/browser mismatch can prevent launch.
- Use bounded waits. Set navigation and selector timeouts; decide what your caller should receive when a site never becomes ready.
- Limit work. Cap concurrency, full-page dimensions, capture duration, and input size. Each active browser page consumes memory and CPU.
- Clean up. Close pages and browsers in
finallyblocks. For a long-lived service, consider a managed browser lifecycle and worker limits instead of launching unbounded browser processes per request. - Keep untrusted pages isolated. The Playwright Docker guidance describes its image as intended for testing and development and not recommended for visiting untrusted websites. It advises aligning the image’s Playwright version with the project, using
--init, and using--ipc=hostfor Chromium to reduce memory-related crashes. For crawling untrusted sites, it recommends a separate user with a seccomp profile. Treat this as starting guidance, not a complete production threat model. - Expect variable costs. Self-hosting uses compute, memory, bandwidth, storage, and operational time. Larger full-page images and concurrent captures require more resources. The cited documentation provides no universal benchmark, so measure your own pages and deployment before setting service limits or pricing.
6. Common problems and fixes
| Symptom | Likely cause | What to do |
|---|---|---|
| Browser executable missing | The Playwright package was installed, but its browser binary was not. | Run npx playwright install chromium; install required OS dependencies for the environment using the official installation guide. |
| Browser fails to launch in a container | Missing system dependencies, incompatible package and image versions, or container process/memory configuration. | Align Playwright versions, install browser dependencies, and review the official Docker guidance for --init and Chromium shared memory. |
| Navigation times out | The page is slow, never settles, or the selected wait condition is too strict. | Use a suitable wait condition such as domcontentloaded, then wait for the specific content needed. Keep a finite timeout. |
| Screenshot is blank or missing content | Client rendering, fonts, lazy images, or a consent/interstitial screen may not be ready or may block the page. | Wait for a visible content selector, use a page-specific readiness signal, and inspect the actual page state before capture. |
| Target element is not found | The selector is wrong, the element appears later, or it is inside a frame or shadow tree. | Check the selector in the target page, wait for visibility, and use the appropriate frame or locator strategy. |
| Image is too large or capture runs out of memory | Full-page capture, high device scale, or many concurrent pages create large bitmaps. | Capture the viewport or a specific element, lower scale, cap page dimensions, and reduce concurrency. |
| Works locally but not on a server | The deployment lacks browser libraries, fonts, permissions, writable output storage, or enough memory. | Build from a compatible browser environment, verify dependencies and filesystem permissions, and test the same target deployment configuration. |
| Unexpected access to internal services | A public endpoint lets the browser reach user-supplied hosts, redirects, or subresources. | Restrict network egress and destinations outside the browser; validate resolved addresses throughout redirects and enforce isolation. |
7. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF, so you do not need to install or operate the browser for this capture.
cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the API documentation for the request options and response details. ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; those steps can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server lets AI agents use screenshot, page-info, and PDF-capture tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up free for 1,000 screenshots a month, with no card required.
8. Frequently asked questions
Can a JavaScript downloader capture a page that requires login?
It can when the browser session has the required authentication, but handling credentials and cookies safely is your responsibility. Do not accept arbitrary credential values in a public endpoint without a clear security design.
Does a screenshot preserve text as selectable text?
No. PNG, JPEG, and WebP screenshots are raster images. If you need a document with selectable text, choose a PDF workflow or extract page content separately.
Should I choose Playwright or Puppeteer?
Both document page and element screenshots. Use the library that best fits your existing tooling and browser requirements; the research sources establish no universal performance winner.
Can I use this for screenshots of arbitrary public URLs?
A local script can visit chosen URLs, but a publicly callable service needs destination restrictions, network isolation, timeouts, resource limits, and abuse controls before it should accept arbitrary input.


