ScreenshotNeo

BlogHow-to

How to Process a Large URL List into Screenshots with n8n in Batches

Build an n8n workflow that splits URL lists, captures pages in controlled batches, saves images safely, and retries failures without loading every screenshot into memory.

By the ScreenshotNeo team4 October 20269 min read

To process a large URL list into screenshots with n8n, turn each URL into its own item, call a screenshot endpoint for each item, and save each image as soon as it returns. Use Loop Over Items when you need a controlled batch size, a delay, or multi-step work per batch. n8n already runs ordinary nodes across incoming items, so a loop is not required just to capture one item at a time.

For large jobs, keep image binaries out of the long-running workflow data: write files to durable storage and pass only their paths or object keys downstream. Start with a small pilot, then tune batch size and delays against the endpoint’s documented limits and your n8n memory.

1. Prepare the URL items

Inspect the incoming data before building the capture loop. The loop should receive one item per URL, ideally with a stable identifier alongside it.

Input shape What to do
One item per URL, such as { "url": "https://example.com" } Pass the items into the capture path. n8n nodes commonly process incoming items automatically.
One item containing an array, such as { "urls": ["https://a.example", "https://b.example"] } Use the Split Out node on urls to make separate items before capture.

Keep any record ID or source metadata on each item. It helps match a returned image to its original URL and makes failed records straightforward to retry. Check item counts immediately after splitting; an array left intact may be treated as one item and yield only one capture.

2. Choose automatic item processing or an explicit loop

If the workflow simply needs one HTTP Request and one storage action per URL, connect those nodes directly. n8n processes the input items through the path. Add Loop Over Items when you need a deliberate batch size, a wait between batches, or a sequence that must complete for each group before continuing.

To wire an explicit loop, send its loop output through screenshot capture and storage, then connect the last node in that processing path back to Loop Over Items. Connect the loop node’s done output to reporting or a final summary step. The n8n documentation notes that setting Batch Size to 1 processes items individually; larger values let a batch proceed together.

  1. Add Loop Over Items after the normalized URL items.
  2. Set a conservative batch size for the pilot. Screenshot payloads are much larger than small text records, so do not copy a general small-record batch recommendation as a screenshot-specific setting.
  3. Connect the loop output to the HTTP Request node, then to file or object storage.
  4. Return the storage node’s output to the loop node to continue.
  5. Use the done output for a summary, notification, or aggregate of file references.

When the endpoint supports a documented bulk request, you can instead send one request per batch. Confirm its batch contract first: a normal one-item-per-URL request inside the loop is not the same as one array request per batch.

3. Configure the screenshot request

A hosted screenshot API keeps browser installation and runtime management outside n8n. Browserless documents a POST request to its screenshot endpoint, with a page URL and optional screenshot settings, returning an image. Its documentation also provides an n8n HTTP Request integration example. Store the Browserless token in n8n Credentials and follow its current API documentation for the exact endpoint, authentication, options, and limits.

In the HTTP Request node, map the current item’s URL into the request body as required by the provider. Configure the response as a file/binary response when the node and endpoint support it. Keep the incoming URL and ID available so the binary result can be paired with the correct record.

Choose capture settings based on the output you need. Browserless documents PNG, JPEG, and WebP output and full-page capture. Use full-page mode when the entire document is needed; viewport capture is usually smaller and fits cases where only the visible screen matters. A full-page setting does not guarantee that lazy-loaded sections have rendered: pages may need scrolling or other page-specific loading behavior before capture.

For custom browser interactions or rendering logic beyond the hosted API’s options, a self-managed browser automation runtime using Puppeteer is another route. Puppeteer’s Page.screenshot() can save a screenshot, but you own deployment, browser availability, runtime maintenance, and output handling. See its screenshot guide for supported behavior.

4. Store each image as the loop runs

Do not accumulate all screenshot binaries into one final JSON object. Image data can dominate workflow payload size and memory use. Write each result to durable storage during the loop, then carry a filename, object key, or storage reference with the source URL and capture status.

  • Use a deterministic filename or object key based on a record ID, not an unescaped URL. URLs can contain characters that are awkward in paths and may expose query data.
  • Record the original URL, storage reference, output format, and success or error status for each item.
  • Retain enough metadata to retry failures without reprocessing successful captures.
  • Keep binary data only as long as the next node needs it. Configure HTTP Request and storage nodes for file handling where appropriate.

5. Tune batch size, pace, and memory

There is no universal screenshot batch size or interval. n8n’s general large-workflow guidance gives around 100–200 small records as a starting point, while explicitly making the right size dependent on payload size, node behavior, memory, and external service limits. Screenshot files are comparatively large, so treat that figure as context only, not a recommended screenshot setting.

  1. Run a small sample and inspect execution time, binary size, memory pressure, and provider responses.
  2. Increase batch size gradually only while the workflow and endpoint remain stable.
  3. Use a Wait node between batches if the endpoint’s documented rate limits or observed responses call for pacing.
  4. Store files as you go and avoid aggregating binary output for the entire list.
  5. For long jobs, make progress observable with per-item statuses and a final count of succeeded, failed, and skipped records.

More concurrency can shorten elapsed time, but it also increases simultaneous browser work, response payloads, and pressure on n8n and the screenshot service. Set concurrency and waits from the provider’s documented limits and your own measured workload; the research does not establish universal throughput, a success rate, or a best retry policy.

6. Handle failures and retries

Web pages can fail to load or render as expected because of site scripts, authentication requirements, consent screens, timing, or service controls. Keep the failed URL and error information instead of dropping the item. Route errors into a failure branch or record them for a later retry, while allowing successful files to remain stored.

Make retries selective. Retry transient network or service failures according to the provider’s documented behavior; correct malformed URLs and authentication problems before retrying. Avoid retrying every item indefinitely, and avoid overwriting a good stored result with a failed attempt. The source material does not establish a universal retry count or delay.

7. Verify the pilot before scale-up

  • Confirm the item count after splitting matches the number of intended URLs.
  • Check that each request uses the URL from its own item and that output references retain the matching ID.
  • Verify the expected output format and whether the capture is viewport or full-page.
  • Open a sample of saved files to confirm they are readable and stored at the expected location.
  • Test a known failing URL and verify that its error status is preserved without stopping unrelated records.
  • Check execution memory and duration before raising batch size.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. A single GET request captures a URL as PNG, JPEG, WebP, or PDF. Use the URL from the current n8n item in the url parameter and store the response as a file. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

In n8n, set the HTTP Request method to GET, use the URL expression from the current item, and provide the access key through n8n Credentials. Configure the response for file handling and connect it to storage. ScreenshotNeo accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.

Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

Troubleshooting

Symptom Likely cause Fix
Only one screenshot is produced The URLs remain in one array-valued item. Split the array into individual items before the request or loop, then verify the item count.
The workflow makes too many requests at once Items are being processed automatically or the explicit batch is too large for the provider or available memory. Add Loop Over Items for controlled batching, lower the batch size, and use Wait when the provider limits or observed responses require it.
Execution runs out of memory or becomes slow Binary screenshots are retained across many items or collected for a final aggregation. Write each image to durable storage promptly and pass references downstream; reduce batch size and inspect binary retention settings.
The output is blank, incomplete, or shows a consent screen The site has not finished loading, lazy content has not appeared, or the page requires authentication or consent interaction. Check the target page and provider options, allow appropriate loading time, and scroll or interact when needed for lazy content. Results can differ by site behavior.
The saved file is not an image The response was handled as JSON/text, or the endpoint returned an error body. Set the node response format to file/binary where supported; inspect status and response headers before saving.
Files are associated with the wrong URLs URL metadata was discarded or binary output was merged without a stable item identifier. Preserve URL and record ID through capture and storage, and test mapping with a small sample.
Some URLs fail while others succeed Individual pages may time out, block automated access, require credentials, or encounter transient service failures. Record per-item errors, check access and provider guidance, and retry only eligible failures.

Performance, reliability, and cost notes

  • Performance: Full-page files and complex pages take more resources than viewport captures. Batch size, wait intervals, page load behavior, and storage latency all affect total job time.
  • Reliability: Capture output is affected by the target site and its scripts, authentication, consent behavior, and controls. Preserve item-level status and design retries around documented provider behavior.
  • Cost: The consulted n8n, Browserless, and Puppeteer sources do not establish a comparable total cost or throughput. Check current provider pricing and limits before choosing a hosted service; self-managed Puppeteer also requires you to operate its runtime.
  • Data handling: URLs may contain sensitive query parameters and screenshots may contain private page content. Use appropriate credential storage and access controls for workflow data and saved files.

FAQ

Do I need Loop Over Items for every screenshot workflow?

No. Ordinary n8n nodes process incoming items. Use the loop when you need deliberate batches, waits, or explicit multi-step iteration.

Should a batch contain 100 or 200 URLs?

There is no screenshot-specific universal value. The cited general guidance applies to small records; pilot with a smaller workload and tune based on binary size, memory, node behavior, and API limits.

Can I send the whole URL list in one request?

Only when the selected screenshot endpoint documents a bulk request and its request and response format. Otherwise send one URL per item through the capture path.

What should the workflow keep after saving an image?

Keep a durable file reference, source URL or record ID, and capture status. This supports reporting and selective retries without retaining every binary in memory.

Sources