ScreenshotNeo

BlogHow-to

How to bulk screenshot URLs from a Notion database

Export a Notion database to CSV, capture each URL with Playwright, and save results with a repeatable naming scheme. Includes a hosted API option.

By the ScreenshotNeo team4 October 202611 min read

To bulk screenshot URLs in a Notion database, export the database as CSV, read the URL column, and loop through valid addresses with a browser automation tool such as Playwright. Save each screenshot using a stable row identifier or sanitized page title, and record failures for review. This workflow keeps capture files local; putting screenshots back into Notion is a separate API task that requires checking Notion’s current upload and page-property documentation.

1. Prepare and export the Notion database

  1. Make sure each row has a URL property containing the page address you want to capture. Notion’s Web Clipper can add the original webpage URL to a database entry, but it does not perform bulk screenshots.
  2. Use a full-page database and choose Export, then export as Markdown & CSV. Notion documents that full-page databases export as a CSV, with Markdown files for subpages. [Notion: Export your content]
  3. Use a Table view when exporting a database with a Form view: Notion says Form views cannot currently be exported and recommends exporting the questions and responses from Table view. Guests need Full access to export, and workspace or teamspace settings may disable export.
  4. Find the CSV and note the exact name of the URL column. The examples below assume it is named URL.

If the database is inline, export the page that contains it and confirm the resulting CSV includes the intended rows and URL property before running a capture batch.

2. Install Playwright and a browser

Playwright’s documented Page example navigates to a URL and saves a screenshot file. That makes it suitable for iterating over exported rows. [Playwright Page API]

Python setup

python -m venv .venv
# macOS or Linux:
source .venv/bin/activate
# Windows PowerShell:
# .venv\Scripts\Activate.ps1

python -m pip install playwright
python -m playwright install chromium

Node.js setup

npm init -y
npm install playwright
npx playwright install chromium

Choose one implementation below. Both use Chromium, a fixed viewport, per-page timeouts, sanitized filenames, and a CSV manifest that records successful captures and failures. They continue after an individual URL fails.

3. Capture the URLs from CSV

Python: complete CSV-to-screenshot script

Save as capture_notion_csv.py. It uses Python’s built-in CSV module, so no CSV package is needed.

import csv
import re
import sys
from pathlib import Path
from urllib.parse import urlparse
from playwright.sync_api import sync_playwright

CSV_PATH = Path(sys.argv[1] if len(sys.argv) > 1 else "notion-export.csv")
URL_COLUMN = "URL"  # Change to the exact CSV column name.
NAME_COLUMN = "Name"  # Optional: change to your title column, or leave as-is.
OUTPUT_DIR = Path("screenshots")
MANIFEST_PATH = OUTPUT_DIR / "manifest.csv"
TIMEOUT_MS = 45_000
VIEWPORT = {"width": 1440, "height": 1000}


def safe_name(value: str, fallback: str) -> str:
    value = re.sub(r"[^A-Za-z0-9._-]+", "-", (value or "").strip())
    value = value.strip("-._")[:100]
    return value or fallback


def valid_http_url(value: str) -> bool:
    try:
        parsed = urlparse((value or "").strip())
        return parsed.scheme in ("http", "https") and bool(parsed.netloc)
    except ValueError:
        return False


def main() -> None:
    if not CSV_PATH.is_file():
        raise SystemExit(f"CSV not found: {CSV_PATH}")

    OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
    results = []
    with CSV_PATH.open("r", encoding="utf-8-sig", newline="") as f:
        reader = csv.DictReader(f)
        if not reader.fieldnames or URL_COLUMN not in reader.fieldnames:
            raise SystemExit(
                f"Missing URL column {URL_COLUMN!r}. CSV columns: {reader.fieldnames}"
            )
        for row_number, row in enumerate(reader, start=1):
            url = (row.get(URL_COLUMN) or "").strip()
            label = row.get(NAME_COLUMN, "") if NAME_COLUMN else ""
            base = safe_name(label, f"row-{row_number:04d}")
            # Include the source row number to avoid collisions between duplicate titles.
            filename = f"{row_number:04d}-{base}.png"
            result = {
                "row": row_number,
                "name": label,
                "url": url,
                "file": str(OUTPUT_DIR / filename),
                "status": "skipped",
                "error": "",
            }

            if not valid_http_url(url):
                result["error"] = "Empty or invalid HTTP(S) URL"
                results.append(result)
                continue

            try:
                with sync_playwright() as p:
                    browser = p.chromium.launch(headless=True)
                    page = browser.new_page(viewport=VIEWPORT)
                    page.goto(url, wait_until="domcontentloaded", timeout=TIMEOUT_MS)
                    page.screenshot(path=result["file"], full_page=True)
                    browser.close()
                result["status"] = "ok"
            except Exception as exc:
                result["status"] = "error"
                result["error"] = str(exc)[:1000]
            results.append(result)
            print(f"{result['status']}: {url} -> {result['file']}")

    with MANIFEST_PATH.open("w", encoding="utf-8", newline="") as f:
        writer = csv.DictWriter(f, fieldnames=["row", "name", "url", "file", "status", "error"])
        writer.writeheader()
        writer.writerows(results)
    print(f"Manifest: {MANIFEST_PATH}")


if __name__ == "__main__":
    main()

Run it with the exported CSV path:

python capture_notion_csv.py notion-export.csv

Node.js: complete CSV-to-screenshot script

Save as capture-notion-csv.mjs. Install the small CSV reader with npm install csv-parse.

import fs from 'node:fs';
import path from 'node:path';
import { parse } from 'csv-parse/sync';
import { chromium } from 'playwright';

const csvPath = process.argv[2] ?? 'notion-export.csv';
const urlColumn = 'URL'; // Change to the exact CSV column name.
const nameColumn = 'Name';
const outputDir = 'screenshots';
const timeoutMs = 45_000;
const viewport = { width: 1440, height: 1000 };

function safeName(value, fallback) {
  const cleaned = String(value ?? '')
    .trim()
    .replace(/[^A-Za-z0-9._-]+/g, '-')
    .replace(/^[-._]+|[-._]+$/g, '')
    .slice(0, 100);
  return cleaned || fallback;
}

function validHttpUrl(value) {
  try {
    const parsed = new URL(String(value ?? '').trim());
    return ['http:', 'https:'].includes(parsed.protocol) && Boolean(parsed.hostname);
  } catch {
    return false;
  }
}

if (!fs.existsSync(csvPath)) throw new Error(`CSV not found: ${csvPath}`);
fs.mkdirSync(outputDir, { recursive: true });
const text = fs.readFileSync(csvPath, 'utf8').replace(/^\uFEFF/, '');
const rows = parse(text, { columns: true, skip_empty_lines: true, bom: true });
if (rows.length && !(urlColumn in rows[0])) {
  throw new Error(`Missing URL column ${JSON.stringify(urlColumn)}. Found: ${Object.keys(rows[0]).join(', ')}`);
}

const browser = await chromium.launch({ headless: true });
const results = [];
try {
  for (const [index, row] of rows.entries()) {
    const rowNumber = index + 1;
    const url = String(row[urlColumn] ?? '').trim();
    const label = String(row[nameColumn] ?? '');
    const filename = `${String(rowNumber).padStart(4, '0')}-${safeName(label, `row-${rowNumber}`)}.png`;
    const file = path.join(outputDir, filename);
    const result = { row: rowNumber, name: label, url, file, status: 'skipped', error: '' };
    if (!validHttpUrl(url)) {
      result.error = 'Empty or invalid HTTP(S) URL';
      results.push(result);
      continue;
    }
    const page = await browser.newPage({ viewport });
    try {
      await page.goto(url, { waitUntil: 'domcontentloaded', timeout: timeoutMs });
      await page.screenshot({ path: file, fullPage: true });
      result.status = 'ok';
    } catch (error) {
      result.status = 'error';
      result.error = String(error).slice(0, 1000);
    } finally {
      await page.close();
    }
    results.push(result);
    console.log(`${result.status}: ${url} -> ${file}`);
  }
} finally {
  await browser.close();
}

const csvEscape = (value) => `"${String(value ?? '').replaceAll('"', '""')}"`;
const fields = ['row', 'name', 'url', 'file', 'status', 'error'];
const manifest = [fields.join(','), ...results.map((r) => fields.map((f) => csvEscape(r[f])).join(','))].join('\n');
fs.writeFileSync(path.join(outputDir, 'manifest.csv'), manifest + '\n', 'utf8');
console.log(`Manifest: ${path.join(outputDir, 'manifest.csv')}`);

Run it with:

node capture-notion-csv.mjs notion-export.csv

4. Choose capture behavior and naming

  • Viewport or full page: full_page=True in Python and fullPage: true in Node capture the full scrollable page. Set these to false or omit them for a viewport shot. Very long pages can produce large images; use viewport captures when consistent dimensions matter.
  • Wait condition: The examples wait for domcontentloaded, which returns before every external image or script necessarily finishes. Use load when the page’s load event matters, or wait for a known selector that indicates the content is ready. A fixed delay can help with a known delayed render, but increases batch time and is not a guarantee. Avoid treating network idle as universally reliable: analytics, chat, and streaming requests can keep a page active.
  • Viewport and scale: Change the width and height to match the view you need. A larger device scale factor can produce sharper output and larger files; check the current Playwright browser context options before adding it.
  • Duplicate and unsafe names: The scripts prefix filenames with the CSV row number, so repeated titles do not overwrite each other. They remove path separators and other unsafe characters. Keep the manifest as the mapping back to the original title and URL.
  • Redirects and authentication: A URL may redirect, require login, or render different content depending on cookies or headers. The sample scripts do not sign in to sites. Only automate access you are authorized to use, and supply site-specific session state through an appropriately protected setup if needed.
  • Dynamic or lazy content: Full-page screenshot capture may trigger layout or lazy-loading behavior, but do not assume every site has finished loading its content. Wait for a page-specific selector or add a bounded delay when required, then inspect representative captures.

5. Run and review the batch

  1. Try a small CSV with two or three rows first, including one page with images or dynamic content.
  2. Open the output files and check dimensions, page completeness, and whether the captured content matches your intended view.
  3. Inspect screenshots/manifest.csv. Fix invalid URLs and investigate timeouts, redirects, blocked pages, and authentication walls separately.
  4. Rerun only failed rows after editing the CSV or adapting the script. For large batches, add a retry count with a delay and save results after every row so an interrupted run does not lose its log.

6. Throughput, reliability, and cost

Browser capture has no per-screenshot API charge in this local workflow, but it uses your computer or server’s CPU, memory, disk, and network. Each page takes as long as its navigation and rendering require. Reusing one browser, as the Node script does, avoids launching a browser process for every row; the Python example is intentionally straightforward but launches one per row. For larger exports, improve throughput by reusing a browser and limiting concurrent pages to a small number that the host can handle. Excessive parallelism can exhaust memory or cause target sites to throttle or block requests.

For repeatable runs, keep the exported CSV, script version, viewport, wait rules, and manifest together. Add bounded retries for transient failures, preserve partial results, and record the final URL if redirect diagnosis matters. A browser workflow is sensitive to page changes, network conditions, bot checks, and browser updates, so review failures rather than treating a missing file as a valid capture.

For a hosted screenshot API, compare authentication support, wait controls, full-page behavior, output format, data handling, rate limits, failure reporting, and cost. ScreenshotNeo is a website screenshot API and MCP server; it accepts a URL in one GET request and returns an image or PDF. Its product-specific options and pricing are described below.

7. Troubleshooting

Symptom Likely cause Fix
Script reports a missing URL column The exported CSV uses a different property name, capitalization, or whitespace. Open the CSV header row and set URL_COLUMN in Python or urlColumn in Node to match it exactly.
Many rows are skipped as invalid Blank cells, non-HTTP links, or URLs stored in another property. Correct the URL property in Notion or map the right CSV column. Confirm addresses begin with http:// or https://.
Browser executable is missing Playwright package is installed but its browser has not been downloaded. Run python -m playwright install chromium or npx playwright install chromium for the implementation in use.
Navigation timeout Slow site, stalled request, or navigation condition not reached. Check the address manually, raise the bounded timeout for known slow sites, and consider domcontentloaded or a page-specific readiness check.
Screenshot is blank or incomplete Capture happened before client-side content rendered, or the page blocks automated browsing. Wait for a content selector or a short bounded delay and inspect the page outcome. Some bot checks and access restrictions cannot be solved by waiting.
Images or page sections are missing Lazy loading, third-party resources, or delayed rendering. Use full-page capture where appropriate, wait for the relevant content, and validate samples. Site behavior differs, so tune per site when necessary.
Files overwrite each other Names were generated only from titles that are duplicated or normalize to the same string. Retain the row-number prefix or use a stable unique database identifier in filenames.
Export is unavailable The current view or access level does not permit export, or workspace settings disable it. Use Table view for Form data, confirm your access level, and ask a workspace administrator whether export is disabled.

8. Optionally return screenshots to Notion

The workflow above saves files locally. Returning images to Notion requires another pipeline step. The research for this guide does not verify Notion’s current API version, database query and pagination behavior, file-upload route, upload restrictions, or the payload for attaching an image to a page or database property. Check the current official API documentation for those details before implementing a write-back integration. Do not assume a local screenshot path can be assigned directly to a Notion image property.

Or skip the browser setup

ScreenshotNeo accepts a URL and returns a screenshot in one API request. See the ScreenshotNeo API documentation for the available options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners are accepted and removed, and known consent platforms, newsletter popups, and chat widgets can be removed before capture; each of those steps can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers report the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
  • The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan.

ScreenshotNeo also supports full-page capture, CSS-selector element capture, device presets and custom viewports, retina scale, PDF options, HTML-to-image, custom CSS and JavaScript, selector clicks and waits, request blocking, headers, cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed public image links, async jobs with signed webhooks, batches of up to 100 URLs per call, a usage API, and an OpenAPI spec. Parameter names used by other screenshot APIs also work. See ScreenshotNeo for the product and plan details.

Sign up for 1,000 free screenshots a month with no card.

FAQ

Can I capture every URL in a Notion database automatically?

Yes. Export the database to CSV and automate the URL column, or build a Notion API workflow after verifying its current query and pagination documentation.

Does Notion’s Web Clipper take screenshots?

No. It can save a page to a database and preserve the original address in a URL property; screenshot capture is a separate step.

Can I save screenshots back into Notion?

That needs an upload and attachment workflow. Confirm the current Notion API requirements before building it; the local scripts here only save files and a manifest.

Should I capture the viewport or full page?

Use viewport capture for a consistent visible-screen preview. Use full-page capture when the entire document is needed, and review especially long or dynamic pages for completeness.