ScreenshotNeo

BlogHow-to

How to Scrape Multiple URLs with a Web Scraping API

Submit a URL list, collect asynchronous results, handle per-page failures, and choose batch settings that fit your scraping workload.

By the ScreenshotNeo team29 September 20269 min read

How to Scrape Multiple URLs with a Web Scraping API

To scrape multiple known URLs with a web scraping API, send the URL list to a batch endpoint. For a small job, a synchronous endpoint may return results in the same request. For longer or larger jobs, submit asynchronously, save the job or task IDs, then poll for results or receive a webhook. Track status per URL: a batch can contain both successful and failed pages.

A batch is for URLs you already know. If you want the service to discover and traverse links from a starting page, use a crawl feature instead. Firecrawl documents this distinction in its batch scrape documentation.

1. Prepare the URL list and decide how results should arrive

Start with a stable input list and the output your application needs. Keep each submitted URL alongside its result identifier and final status. This lets you reconcile results even if the provider returns them in a different order or individual pages fail.

A batch starts with known URLs; each page has its own status and result.
A batch starts with known URLs; each page has its own status and result.
  • Use synchronous batch processing when the list is small and your client can wait for the full response.
  • Use an asynchronous job when work may take longer than a normal HTTP request or needs to continue after the submitting process exits.
  • Use webhooks or callbacks when you want the provider to notify your application as results finish.
  • Use polling when a job is occasional or you do not have a webhook receiver.

Batch request bodies and authentication are provider-specific. Do not copy a URL, key field, or response shape from one vendor into another API. The example below uses ScraperAPI’s documented asynchronous batch endpoint, not a generic standard.

2. Submit an asynchronous batch with ScraperAPI

ScraperAPI documents a POST to https://async.scraperapi.com/batchjobs with a JSON object containing apiKey and a urls array. The response contains a separate record for each URL, including an ID, status, status URL, and URL. See the ScraperAPI batch request documentation for its current request and response fields.

Set the API key in an environment variable rather than committing it to source control. The following examples submit the same two URLs. They print or save the returned records so you can process each job later.

cURL

export SCRAPERAPI_KEY='YOUR_API_KEY'
curl --fail-with-body --silent --show-error \
  -X POST 'https://async.scraperapi.com/batchjobs' \
  -H 'Content-Type: application/json' \
  -d "{\"apiKey\":\"$SCRAPERAPI_KEY\",\"urls\":[\"https://example.com/\",\"https://example.org/\"]}"

Keep the response body. Extract each returned job ID and status URL; do not assume the response is one atomic batch result.

Python

import json
import os
import requests

api_key = os.environ["SCRAPERAPI_KEY"]
urls = ["https://example.com/", "https://example.org/"]

response = requests.post(
    "https://async.scraperapi.com/batchjobs",
    json={"apiKey": api_key, "urls": urls},
    timeout=30,
)
response.raise_for_status()
jobs = response.json()

with open("submitted-jobs.json", "w", encoding="utf-8") as f:
    json.dump(jobs, f, indent=2)

for job in jobs:
    print(job.get("url"), job.get("id"), job.get("status"), job.get("statusUrl"))

Node.js

const apiKey = process.env.SCRAPERAPI_KEY;
if (!apiKey) throw new Error('Set SCRAPERAPI_KEY');

const urls = ['https://example.com/', 'https://example.org/'];
const response = await fetch('https://async.scraperapi.com/batchjobs', {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({ apiKey, urls }),
  signal: AbortSignal.timeout(30_000),
});

if (!response.ok) {
  throw new Error(`Batch submission failed: ${response.status} ${await response.text()}`);
}
const jobs = await response.json();
console.log(jobs);

3. Collect results without losing the input-to-output mapping

After submission, persist the returned IDs and associate each one with its original URL. A useful record includes the input URL, provider task ID, current status, attempt count, timestamps, and final result location. Keep secrets out of logs, and store response data that your application needs in your own storage.

For polling, request each provider’s documented status URL and stop when the individual job reaches a terminal state. Do not poll every few milliseconds. Scrape.do recommends exponential backoff for status checks, documents 429 rate limiting, and advises checking each task’s status. Its async API guide shows a create-job, get-job, and get-task workflow.

delay = 2 seconds
while any task is pending:
    wait(delay)
    fetch status for pending tasks
    record status and any result/error
    delay = min(delay * 2, 60 seconds)

This is pseudocode: use the exact status endpoint and terminal status names documented by your provider. Add a maximum elapsed time and a way to resume polling after a process restart. For frequent or production workloads, webhooks avoid continuous status requests. Validate webhook signatures if the provider supports them; Firecrawl, for example, documents HMAC-SHA256 verification using the X-Firecrawl-Signature header.

4. Handle partial failures and retry selectively

Batch completion does not guarantee that every page succeeded. A provider may report an overall job state while exposing individual URL or task errors. Inspect each item. Save successes, isolate failed URLs, and retry only those that are appropriate to retry. This selective retry approach follows from the per-URL status and error information documented by batch APIs.

  1. Classify each result as success, retryable failure, or permanent/unknown failure.
  2. Record the provider status, error message, attempt number, and time.
  3. Retry transient failures with a delay and a bounded attempt count.
  4. Do not repeatedly resubmit successful URLs unless your data policy requires a fresh scrape.
  5. Send persistent failures to a review queue or report rather than silently dropping them.

Firecrawl documents an operation to inspect errors for failed URLs and per-page webhook events. Scrape.do likewise advises checking task-level status and handling errors. Keep your internal outcome model independent of a vendor’s exact status strings so you can change providers without rewriting downstream processing.

5. Set batch size, concurrency, and delivery mode

A batch endpoint is not unlimited parallelism. The provider may cap URLs per submission, concurrent browser sessions, or submission rates by plan. Check the current account limits before sizing jobs, and keep your own concurrency bounded when submitting multiple batches.

Provider example Documented behavior What to verify
Firecrawl Explicit-list batches can be synchronous or asynchronous; per-job maxConcurrency is configurable. Team browser concurrency, chosen job limit, and result retention.
ScraperAPI Async batch returns one job record per URL; its documentation states up to 50,000 URLs per batch job. Current batch cap, plan limits, and acceptable submission rate.
Oxylabs Push-Pull Async batch supports up to 5,000 URL or query values per POST; results can use callbacks or cloud storage. Plan-dependent job submission limits and retention.
Scrape.do Async jobs use separate create, job-status, and task-result operations. Current plan concurrency, task expiry, and rate limits.

These are vendor-specific values from documentation accessed in 2026, not general API limits; product limits can change. Scrape.do’s listed async concurrency is plan-specific, and should be checked against the current account documentation. Do not choose a provider by maximum batch size alone. Compare explicit URL-list support, sync and async modes, per-page errors, concurrency controls, callbacks, output shape, retention, and whether the service returns raw HTML or structured data.

6. Store results before they expire

Asynchronous APIs are job processors, not necessarily archives. Scrape.do warns that task results are temporary and should be fetched before ExpiresAt. Firecrawl documents API availability for batch results for 24 hours after completion, followed by activity-log availability. Oxylabs says Push-Pull results remain available for at least 24 hours. Treat every retention period as provider-specific: download needed data and persist it in storage you control.

For large jobs, write each result as it arrives rather than waiting for the entire set to complete. This reduces the amount of work lost if a worker stops and makes partial progress visible. Make writes idempotent using the provider task ID or a stable key derived from your input URL and run ID.

7. Operational concerns: performance, reliability, and cost

Performance

Throughput depends on provider concurrency, page load behavior, response size, and target-site behavior. A larger batch can reduce client coordination but does not necessarily make each page finish faster. Start with a modest per-job concurrency, observe completion time and failure rates, then adjust within your provider’s documented limit. Avoid launching many full-concurrency jobs at once.

Reliability

Use request timeouts for submission, durable job records, bounded retries, and a restartable collector. Make webhook handlers idempotent because delivery systems can retry notifications. Keep a dead-letter path for malformed or repeatedly failing records. If polling gets a rate-limit response such as 429, back off and honor any retry guidance in the response rather than immediately retrying.

Cost

Pricing is provider-specific and depends on plan, successful requests, proxy or browser options, and sometimes the target or output mode. The research documentation does not establish a comparable per-URL price across these vendors, so calculate cost from the selected provider’s current billing rules. Track submitted, successful, failed, and retried pages separately; this reveals waste from retry loops and helps estimate future jobs. A batch size limit is not a cost estimate.

8. Troubleshooting common batch-scraping errors

Symptom Likely cause Fix
Submission returns an authentication error Wrong key field, invalid key, or credentials sent to another vendor’s endpoint. Match the endpoint and JSON shape to that provider’s current docs; rotate exposed keys and load secrets from the environment.
Request rejected as too large The URL array exceeds the provider’s batch cap or request body limit. Split the list into smaller batches according to the provider’s documented limit.
Some pages fail while the job continues Per-URL fetch or target-page failure; batches may expose mixed outcomes. Inspect task-level status and errors, preserve successful results, and retry only suitable failed URLs.
Status checks return 429 Polling or submission exceeded a provider rate limit. Increase delays with backoff, reduce concurrent checks, and follow provider response guidance.
Results are missing when fetched later Temporary result retention expired. Fetch and persist results before the documented expiry; do not rely on the API as long-term storage.
Webhook event is processed twice Notification delivery may be retried or duplicated. Make the handler idempotent using event or task identifiers, and verify signatures when supported.
Job appears stalled Pages may be slow, concurrency may be constrained, or the collector may be checking the wrong status resource. Confirm the documented job/task flow, inspect per-task states, and use a bounded timeout and escalation path.

9. Or skip the browser setup

If your goal is to capture screenshots of many pages rather than extract their text or structured fields, ScreenshotNeo provides a screenshot API and MCP server. A single GET request captures one URL; its bulk capture option accepts up to 100 URLs per call. The API returns PNG, JPEG, WebP, or PDF, and the ScreenshotNeo API documentation describes the request options.

Visual capture can remove common overlays before saving the page image.
Visual capture can remove common overlays before saving the page image.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

Cookie banners, popups, and chat widgets are removed before the screenshot. Bot checks, blank pages, and failed loads are never billed, and response headers report the page verdict and billing status. An MCP server lets AI agents use screenshot tools. The free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 screenshots. ScreenshotNeo is for visual capture, not a substitute for an API that extracts page text or structured records.

Sign up for 1,000 free screenshots a month, with no card.

FAQ

Is scraping a list the same as crawling a site?

No. A batch processes URLs you supply; a crawl discovers or follows links from a starting point.

Can one batch have both successful and failed URLs?

Yes. Inspect each URL or task outcome and preserve partial successes.

Should I use a webhook or polling?

Polling is simple for occasional work. Webhooks or callbacks reduce repeated status requests for ongoing workloads.

Can I assume submitted results remain available indefinitely?

No. Retention differs by provider. Retrieve and store any result your application needs.

Can I reuse one provider’s batch code with another?

No. Authentication fields, endpoints, response records, status names, and retention rules vary by API.