How to Bulk Screenshot URLs from a Database Query with Puppeteer
Query URL rows, capture each page with Puppeteer, and save reliable, uniquely named screenshots with bounded concurrency and per-row results.
To bulk screenshot URLs stored in a database, query each row with a stable unique ID, validate its URL, then use one Puppeteer browser with a small pool of pages to navigate and save one screenshot per row. Give each output a filename derived from the row ID, close every page in a finally block, and record failures per row so one bad URL does not hide the rest of the batch.
Puppeteer’s documented capture primitive is Page.screenshot(): navigate a page, then save the screenshot. Its default is a viewport screenshot, so set fullPage explicitly when you need the whole page. Puppeteer screenshot guide · ScreenshotOptions reference.
1. Install Puppeteer and choose a browser
This example uses SQLite and the better-sqlite3 package to make the database query concrete. Adapt the query and database driver to your application; Puppeteer does not prescribe a database, queue, or storage system.
npm init -y
npm install puppeteer better-sqlite3
The puppeteer package downloads a compatible Chrome during installation. Use puppeteer-core instead when your environment supplies or manages the browser; it does not download Chrome and requires explicit browser management. If installation scripts were blocked and launch reports that Chrome is missing, follow Puppeteer’s installation guide and install the browser with npx puppeteer browsers install.
2. Create a URL table
The row ID is used for output naming and result tracking. Avoid using the URL itself as a path: URLs can contain slashes, query strings, unsafe characters, or duplicate values.
CREATE TABLE pages (
id INTEGER PRIMARY KEY,
url TEXT NOT NULL,
enabled INTEGER NOT NULL DEFAULT 1
);
INSERT INTO pages (url) VALUES
('https://example.com/'),
('https://www.iana.org/domains/reserved');
Save that SQL as schema.sql, or create the same table through your usual migration system. For real jobs, query only the intended rows and include whatever stable ID or key your application uses.
3. Save this batch script
Create capture.mjs next to a database file named pages.db. The script creates the example table if needed, selects enabled rows, validates HTTP(S) URLs, captures with two workers by default, retries transient navigation or screenshot errors once, and writes a JSON-lines result file. Adjust the query, output directory, navigation readiness, timeout, and worker count for your workload.
import fs from 'node:fs/promises';
import path from 'node:path';
import Database from 'better-sqlite3';
import puppeteer from 'puppeteer';
const DB_FILE = './pages.db';
const OUTPUT_DIR = './screenshots';
const RESULTS_FILE = './screenshot-results.jsonl';
const CONCURRENCY = 2;
const MAX_ATTEMPTS = 2;
const NAVIGATION_TIMEOUT_MS = 45_000;
const FULL_PAGE = true;
function validateHttpUrl(value) {
let parsed;
try {
parsed = new URL(value);
} catch {
throw new Error('Invalid URL');
}
if (parsed.protocol !== 'http:' && parsed.protocol !== 'https:') {
throw new Error(`Unsupported URL scheme: ${parsed.protocol}`);
}
return parsed.href;
}
function safeFileKey(id) {
// Keep output names deterministic and prevent path traversal.
const key = String(id);
if (!/^[a-zA-Z0-9_-]{1,100}$/.test(key)) {
throw new Error('Row ID must contain only letters, digits, underscore, or hyphen');
}
return key;
}
const db = new Database(DB_FILE);
db.exec(`
CREATE TABLE IF NOT EXISTS pages (
id INTEGER PRIMARY KEY,
url TEXT NOT NULL,
enabled INTEGER NOT NULL DEFAULT 1
)
`);
// Replace this query with your application's filtered or queued work set.
const rows = db.prepare(
'SELECT id, url FROM pages WHERE enabled = 1 ORDER BY id'
).all();
db.close();
await fs.mkdir(OUTPUT_DIR, { recursive: true });
// A new run starts a fresh result log. Use appendFile instead if runs share a log.
await fs.writeFile(RESULTS_FILE, '');
const browser = await puppeteer.launch({ headless: true });
const results = new Array(rows.length);
let nextIndex = 0;
async function captureRow(row, index) {
const startedAt = new Date().toISOString();
let page;
let outputPath;
try {
const key = safeFileKey(row.id);
const url = validateHttpUrl(row.url);
outputPath = path.resolve(OUTPUT_DIR, `${key}.png`);
page = await browser.newPage();
page.setDefaultNavigationTimeout(NAVIGATION_TIMEOUT_MS);
let lastError;
for (let attempt = 1; attempt <= MAX_ATTEMPTS; attempt += 1) {
try {
// domcontentloaded avoids waiting indefinitely for analytics or long polling.
const response = await page.goto(url, { waitUntil: 'domcontentloaded' });
if (response && response.status() >= 400) {
throw new Error(`HTTP ${response.status()} from target`);
}
await page.screenshot({ path: outputPath, type: 'png', fullPage: FULL_PAGE });
return {
id: row.id, url, status: 'success', outputPath,
httpStatus: response?.status() ?? null,
attempts: attempt, startedAt, finishedAt: new Date().toISOString()
};
} catch (error) {
lastError = error;
// A navigation timeout can leave the page in a transitional state.
if (attempt < MAX_ATTEMPTS) {
try { await page.goto('about:blank', { timeout: 5_000 }); } catch {}
}
}
}
throw lastError;
} catch (error) {
// Remove a partial file from a failed attempt, if one was created.
if (outputPath) await fs.rm(outputPath, { force: true }).catch(() => {});
return {
id: row.id, url: row.url, status: 'failed',
error: error instanceof Error ? error.message : String(error),
startedAt, finishedAt: new Date().toISOString()
};
} finally {
if (page && !page.isClosed()) await page.close().catch(() => {});
}
}
async function worker() {
while (true) {
const index = nextIndex++;
if (index >= rows.length) return;
results[index] = await captureRow(rows[index], index);
}
}
try {
const workerCount = Math.min(Math.max(1, CONCURRENCY), rows.length || 1);
await Promise.all(Array.from({ length: workerCount }, () => worker()));
} finally {
await browser.close();
}
for (const result of results) {
await fs.appendFile(RESULTS_FILE, `${JSON.stringify(result)}\n`);
}
const succeeded = results.filter(result => result.status === 'success').length;
const failed = results.length - succeeded;
console.log(`Finished ${results.length} rows: ${succeeded} succeeded, ${failed} failed.`);
console.log(`Images: ${path.resolve(OUTPUT_DIR)}; results: ${path.resolve(RESULTS_FILE)}`);
Run it with node capture.mjs. Each successful row produces screenshots/<id>.png. Failures remain in screenshot-results.jsonl with their row ID, URL, and error message.
4. Tune navigation and screenshot behavior
| Choice | Use it when | Tradeoff |
|---|---|---|
waitUntil: 'domcontentloaded' |
The document markup is enough to capture, or the page has persistent network activity. | Some images or client-rendered content may not have finished. |
waitUntil: 'load' |
You need the browser load event and the site completes it reliably. | Third-party assets can delay the event. |
waitUntil: 'networkidle0' or 'networkidle2' |
The page settles its network requests and the result needs that settled state. | Analytics, polling, and streaming can prevent or delay network idle. |
fullPage: true |
You need a capture of the full document. | Very long pages may use substantial memory and take longer to capture. |
fullPage: false |
You want the viewport only. This is Puppeteer’s documented default. | Content below the fold is omitted. |
clip: {x, y, width, height} |
You need one rectangular region. | The clip must fit the page dimensions and is not a CSS selector. |
type: 'jpeg' or 'webp' with quality |
Smaller lossy output is acceptable. | Quality applies to supported lossy formats, not PNG. |
omitBackground: true |
You need transparency where the page background allows it. | Page elements may still paint their own backgrounds. |
Options such as path, type, quality, clip, captureBeyondViewport, fullPage, and omitBackground are documented in Puppeteer’s ScreenshotOptions reference. If a single element is required, locate it and call its screenshot method; Puppeteer’s screenshot guide includes an element capture example.
For lazy-loaded images, a full-page screenshot alone may not cause every site to load its lazy content. A project-specific approach is to scroll through the page in increments and wait briefly for content to appear, then capture. Sites differ: scrolling may trigger more content indefinitely, and fixed or sticky elements can repeat in full-page output. Define the required visual state before adding scrolling logic.
5. Database, naming, and batch reliability
- Use a stable key. Database IDs prevent duplicate URLs from overwriting each other. If IDs are not path-safe, encode them or map them to a generated UUID; never concatenate untrusted URL text into a filesystem path.
- Validate before navigation. Restrict accepted schemes to HTTP and HTTPS, and apply your product’s host allowlist or network policy. URL parsing does not itself prevent access to internal services or otherwise make a target safe.
- Bound concurrency. A browser can have multiple pages, but no universal worker count is safe for every machine or target. Start low, observe memory, CPU, timeouts, and target behavior, then adjust.
- Keep browser lifetime bounded. Reuse one browser for a batch, create a page per active capture, close pages in
finally, and close the browser even when the batch throws. - Persist outcomes. Keep per-row success/failure records. For large or restartable jobs, persist status in the database or queue rather than keeping all results in memory.
- Retry selectively. Retrying transient network failures can help; repeated retries of invalid URLs, access-denied pages, or deterministic HTTP errors waste time. Use a bounded attempt count and record attempts.
- Make reruns safe. Stable paths intentionally replace the previous screenshot for that row. If you need history, include a job ID or timestamp in the path and clean up old artifacts under your retention policy.
- Separate database and browser pressure. Read the work set in a short query, then release the connection before opening browser pages. For a very large table, fetch in pages or claim jobs atomically rather than loading every URL into memory.
6. Performance, reliability, and cost
Browser startup has a fixed cost, so launching one browser per URL is usually unnecessary. A reused browser with bounded active pages avoids that repeated startup while limiting resource use. Increasing concurrency can improve throughput when the machine and target sites have spare capacity, but it also increases memory, CPU, network traffic, and the risk of throttling. Measure the actual workload; this guide does not assume a throughput figure.
Timeouts should reflect the pages and service-level needs. A timeout is not proof the URL is permanently bad. Keep the failed row and error, and decide whether to retry later. For dependable large batches, use a queue, checkpoint each result, cap retry attempts, and make output writes atomic (write a temporary file then rename it) so interrupted jobs do not leave a partial image that looks successful.
The direct costs are your compute, browser deployment, network use, database, and image storage. Puppeteer itself does not set a price per screenshot. Full-page captures and high concurrency typically require more resources than a viewport capture, so control page length, output format, worker count, and retention.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
Could not find Chrome |
The install script did not download a browser, or the runtime cache differs from the install environment. | Install the compatible browser during deployment with Puppeteer’s browser install command, or configure the managed executable. See the installation guide. |
| Navigation timeout | The site is slow, unreachable, or never reaches the chosen readiness condition. | Use a suitable timeout and readiness event; avoid network idle for pages with ongoing requests. Record and retry only if policy allows. |
| Screenshot is blank or incomplete | The app renders after document load, needs authentication, or depends on delayed/lazy content. | Wait for a known selector or application-ready signal, establish required cookies/session state, or scroll to trigger lazy content. |
| Many rows fail together | Concurrency may exceed machine capacity, the target may throttle requests, or shared network/database access is failing. | Reduce workers, inspect the per-row error records, and retry a smaller batch after correcting the shared cause. |
| Images overwrite one another | The output filename is based on URL basename or another non-unique value. | Use a unique database ID or generated key in every path. |
| Invalid protocol or malformed URL | The database value is incomplete, malformed, or uses an unsupported scheme. | Validate during ingestion and accept only the protocols your application intends to capture. |
| Huge screenshot files or memory spikes | Full-page output is very tall, pages are resource-heavy, or too many pages are active. | Use viewport or clipped capture where suitable, reduce concurrency, choose a lossy format if acceptable, and set storage limits. |
| Browser process remains after an error | Browser cleanup was skipped on an exceptional path. | Put browser closure in a top-level finally block, as in the example, and close each page similarly. |
8. Or skip the browser setup
If you want a screenshot endpoint instead of installing and operating Chrome, ScreenshotNeo accepts one GET request per URL and returns an image or PDF. It can fit a database batch by calling the API for each row and saving each response under that row’s stable ID; see the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
In a database loop, substitute each validated row URL and write the response to a unique ID-based filename. ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, and failed loads are never billed; its MCP server lets AI agents take screenshots; and 1,000 screenshots per month are free with no card, with paid plans starting at $5 for 3,000. Sign up for the free plan.
9. Frequently asked questions
Can Puppeteer capture URLs from any database?
Yes. The database driver supplies the rows; Puppeteer receives the URL string. Keep the query and browser capture as separate steps so each can be changed independently.
Should I use one browser or one browser per URL?
For a batch, reuse a browser and bound the number of active pages. A new browser per URL repeats startup work and makes resource use harder to control.
How do I match a screenshot to its database record?
Use the database row’s stable ID in the image filename and in the result record. A URL can appear more than once and may not be safe as a filename.
Can a batch continue after one URL fails?
Yes. Catch errors per row, record the failed row, and allow workers to continue. Reserve batch-level failure for shared problems such as browser startup or database access.
Does fullPage: true guarantee all dynamic content is captured?
No. It captures the page’s full dimensions at capture time; it does not guarantee that deferred, lazy, authenticated, or continually changing content has finished rendering.


