ScreenshotNeo

BlogHow-to

How to Crawl Lists of URLs Efficiently

Crawl a known URL list efficiently with bounded concurrency, careful deduplication, polite per-host limits, resumable state, and fewer repeat downloads.

By the ScreenshotNeo team4 October 202611 min read

Direct answer: To crawl a list of URLs efficiently, start with the smallest reliable URL inventory, deduplicate and bound it, then fetch with limited concurrency per host. Increase concurrency gradually while watching response codes and latency. Save progress and response validators so interrupted runs can resume and unchanged pages do not need to be downloaded again.

There is no universally safe or fastest concurrency setting. Useful concurrency depends on the target site’s tolerance, network conditions, and your crawler’s own scheduling and processing capacity. Raising the worker count blindly can increase throttling, local queue pressure, or processing delays—and can make the crawl slower.

1. Define the crawl scope

Before sending requests, decide what the crawler is allowed and expected to fetch:

  • Allowed hosts: List the hostnames you intend to access. Decide whether redirects may lead outside them.
  • URL patterns: Include only paths and query variants relevant to the task.
  • Maximum inventory size: Set a limit so a malformed source or discovery loop cannot create an unbounded crawl.
  • Freshness: Decide whether this is a one-time snapshot or a recurring crawl, and how old a cached response may be.
  • Access rules: Check the target’s terms and applicable restrictions. Robots rules are not authentication and do not grant permission to access a site.

Robots.txt rules apply to the host, protocol, and port where that file is served. Implementations may interpret directives differently. Google supports directives such as user-agent, allow, disallow, and sitemap, but not crawl-delay. Scrapy’s robots middleware does not automatically act on Crawl-delay or Request-rate. Verify the behavior of your particular crawler rather than assuming a directive is enforced. See the Google robots.txt guidance and RFC 9309.

2. Get the URL inventory from the most direct source

If you already have a file or database of URLs, use it. Otherwise, see whether the site provides a sitemap, documented API, bulk export, or search endpoint that supplies the needed inventory. Scrapy notes that an API, bulk export, or search endpoint can be faster for the operator and cheaper for the website than crawling its pages. Compare coverage, freshness, permissions, quotas, and total requests before choosing.

A URL list improves utilization over serial link discovery: workers can fetch known pages while other requests are in flight. However, producing every request at once may consume scheduler memory or disk. Keep the queue bounded, especially for large inventories.

3. Normalize and deduplicate conservatively

Two URL strings can lead to equivalent content, but superficially similar URLs can also represent genuinely different pages. Normalize only what you understand:

  • Remove fragments (#section) for HTTP fetching; fragments are not sent to the server.
  • Normalize scheme and host casing, and remove default ports where appropriate.
  • Do not reorder or drop query parameters unless the target treats them as equivalent.
  • Remove tracking or session parameters only when they do not affect content, authorization, or access semantics.
  • Keep meaningful pagination, locale, sorting, and filter parameters.
  • Block unbounded patterns such as calendar links that create a new URL for every future or past date unless that is specifically in scope.

Google’s crawl budget guidance and URL structure guidance discuss duplicate and parameter-driven URL expansion. These are useful cautions for crawler design; Google-specific recommendations should not be mistaken for guarantees about every custom crawler.

4. Fetch with bounded, per-host concurrency

Partition the queue by host, then apply a per-host concurrency cap and request pacing. A global worker limit alone can accidentally send most of the work to one domain. Start conservatively, record outcomes, and raise the limit in small steps only while latency and error rates remain acceptable.

In the following runnable Python example, the input file contains one URL per line. It uses a fixed global worker pool, a per-host semaphore, bounded request timeouts, retry delays for temporary overload responses, redirect host checks, and a JSON Lines output file. Adjust the limits for the target and deployment; the starting values are examples, not universal recommendations.

import asyncio
import json
import time
from collections import defaultdict
from urllib.parse import urlsplit

import aiohttp

INPUT_FILE = "urls.txt"
OUTPUT_FILE = "results.jsonl"
WORKERS = 12
PER_HOST = 2
MAX_URLS = 100_000
TIMEOUT_SECONDS = 30
USER_AGENT = "ExampleResearchCrawler/1.0 (+contact: ops@example.com)"

host_limits = defaultdict(lambda: asyncio.Semaphore(PER_HOST))


def allowed_host(url):
    parts = urlsplit(url)
    return parts.scheme in {"http", "https"} and bool(parts.hostname)


async def fetch_one(session, url):
    if not allowed_host(url):
        return {"url": url, "error": "unsupported or invalid URL"}

    host = urlsplit(url).hostname.lower()
    async with host_limits[host]:
        for attempt in range(4):
            try:
                async with session.get(url, allow_redirects=True) as response:
                    final_url = str(response.url)
                    if urlsplit(final_url).hostname.lower() != host:
                        return {"url": url, "error": "redirect left the original host", "final_url": final_url}
                    body = await response.read()
                    return {
                        "url": url,
                        "final_url": final_url,
                        "status": response.status,
                        "content_type": response.headers.get("Content-Type"),
                        "etag": response.headers.get("ETag"),
                        "last_modified": response.headers.get("Last-Modified"),
                        "bytes": len(body),
                        "body": body.decode(response.charset or "utf-8", errors="replace")
                    }
            except (aiohttp.ClientError, asyncio.TimeoutError) as exc:
                if attempt == 3:
                    return {"url": url, "error": type(exc).__name__ + ": " + str(exc)}
            await asyncio.sleep(min(2 ** attempt, 10))


async def run():
    with open(INPUT_FILE, encoding="utf-8") as source:
        # Conservative baseline: exact string de-duplication. Apply any URL
        # canonicalization only when its effect on target semantics is known.
        urls = list(dict.fromkeys(line.strip() for line in source if line.strip()))
    if len(urls) > MAX_URLS:
        raise ValueError(f"URL inventory exceeds configured cap of {MAX_URLS}")

    timeout = aiohttp.ClientTimeout(total=TIMEOUT_SECONDS)
    connector = aiohttp.TCPConnector(limit=WORKERS, limit_per_host=PER_HOST)
    headers = {"User-Agent": USER_AGENT}
    queue = asyncio.Queue()
    for url in urls:
        queue.put_nowait(url)

    async with aiohttp.ClientSession(timeout=timeout, connector=connector, headers=headers) as session:
        async def worker():
            while True:
                try:
                    url = queue.get_nowait()
                except asyncio.QueueEmpty:
                    return
                result = await fetch_one(session, url)
                result["fetched_at"] = time.time()
                with open(OUTPUT_FILE, "a", encoding="utf-8") as out:
                    out.write(json.dumps(result, ensure_ascii=False) + "\n")
                queue.task_done()

        await asyncio.gather(*(worker() for _ in range(WORKERS)))


if __name__ == "__main__":
    asyncio.run(run())

Install the dependency with python -m pip install aiohttp. This example stores response bodies in the JSON Lines output for clarity; for large pages or large crawls, write body bytes to separate files or object storage and keep metadata in the index. To make it restartable, load already completed URLs from the output before enqueueing work, and write each completed result atomically or append one complete JSON record at a time.

5. Measure the crawl and tune it gradually

Track at least:

  • Requests per host per unit of time and active requests per host.
  • Status codes, especially 429 and 503, plus retry counts and any detected ban or challenge pages.
  • Response latency percentiles, timeout rate, and bytes transferred.
  • Queue depth, memory and disk use, and local callback or parsing time.
  • Completion rate and whether each required URL has a saved result.

When the target responds quickly but the crawl is slow, inspect local work before adding network concurrency. Scrapy documents that callbacks, middleware, and item pipelines share a thread with the event loop; CPU-heavy processing can delay request and response handling. Move expensive parsing or transformation to suitable worker processes or queues, and keep callbacks from blocking.

Scrapy also explains the scheduling tradeoff: queuing known requests early can keep the downloader busy, while too many queued requests consume memory or disk. Use a bounded producer-consumer queue and size it for available resources.

6. Avoid downloading unchanged pages again

For repeat crawls, store validators returned by the server, particularly ETag and Last-Modified. On the next run, send If-None-Match or If-Modified-Since as applicable. A 304 Not Modified response tells a conditional client to reuse its cached representation rather than download the body again. Persist the validator alongside the cached body and URL.

Also save crawl state, result metadata, and retry status so a process restart does not discard completed work. Avoid long redirect chains where you control the target; for Google Search crawling specifically, Google recommends current sitemaps, faster server responses, and using 304 where conditional request semantics allow. See Google’s crawl budget guidance.

7. cURL, Python, and Node.js for small lists

For a handful of URLs, a shell loop is sufficient. This deliberately runs one request at a time. Do not use it as a high-throughput crawler without adding a bounded queue, per-host limits, response recording, and retry policy.

while IFS= read -r url; do
  [ -z "$url" ] && continue
  curl --fail --location --max-time 30 --output /dev/null \
    --write-out '%{http_code} %{time_total} %{url_effective}\n' \
    "$url"
done < urls.txt

A concise sequential Python alternative using the standard library:

from urllib.request import Request, urlopen

with open("urls.txt", encoding="utf-8") as f:
    urls = list(dict.fromkeys(line.strip() for line in f if line.strip()))

for url in urls:
    request = Request(url, headers={"User-Agent": "ExampleResearchCrawler/1.0"})
    try:
        with urlopen(request, timeout=30) as response:
            body = response.read()
            print(response.status, response.url, len(body))
    except Exception as exc:
        print("ERROR", url, type(exc).__name__, str(exc))

A sequential Node.js alternative using the built-in fetch API (Node.js 18 or newer):

import { readFile } from "node:fs/promises";

const urls = [...new Set((await readFile("urls.txt", "utf8"))
  .split(/\r?\n/).map(s => s.trim()).filter(Boolean))];

for (const url of urls) {
  try {
    const response = await fetch(url, {
      headers: { "User-Agent": "ExampleResearchCrawler/1.0" },
      signal: AbortSignal.timeout(30_000),
      redirect: "follow"
    });
    const body = await response.arrayBuffer();
    console.log(response.status, response.url, body.byteLength);
  } catch (error) {
    console.error("ERROR", url, error.name, error.message);
  }
}

These small-list examples report outcomes but do not implement the complete state, robots handling, host partitioning, or conditional caching workflow. Add those controls before using them for recurring or large jobs.

8. Troubleshooting

Symptom Likely cause What to do
Many 429 responses Request rate exceeds the host’s limit. Reduce per-host concurrency and pacing, honor any applicable response guidance, and retry with backoff. Do not immediately resend every failed request.
503 responses or rising latency Temporary server load, overload, or an intermediary rejecting traffic. Slow down or pause that host, use bounded retries with exponential backoff and jitter, then resume gradually.
Ban, CAPTCHA, or challenge pages The site’s protection classified the traffic as automated or disallowed. Stop or reduce requests and review permission and access rules. Do not attempt to bypass access controls.
High concurrency but low throughput Throttling, slow targets, connection limits, or local parsing/event-loop bottlenecks. Compare per-host latency and errors, inspect local processing time and queue use, and lower concurrency if overload signals are present.
Memory or disk grows unexpectedly Too many requests or response bodies are queued or retained. Bound the input queue, stream or separately store large bodies, and persist results incrementally.
Repeated pages appear in results Duplicate URLs, redirects, or query variants produce equivalent content. Record the final URL and apply conservative canonicalization based on the site’s actual semantics.
Unexpectedly missing pages Input truncation, redirect policy, timeout, access restriction, or overly aggressive URL filtering. Compare input and completed counts, record every error and final URL, and inspect filters and redirect handling.
Runs refetch everything after restart Completion state and response validators were not persisted. Store completed URL keys and cached validators durably; resume only incomplete work and use conditional requests when supported.

9. Performance, reliability, and cost

Performance: Throughput is constrained by the slowest part of the pipeline: target response time, per-host tolerance, network bandwidth, scheduler, or local parsing. A sitemap or bulk endpoint may reduce discovery requests. Bounded concurrency avoids an unmanageable backlog, and conditional requests reduce transferred response bodies for unchanged pages.

Reliability: Use explicit timeouts, bounded retries with backoff, durable progress, and clear per-URL outcomes. Retry transient transport failures and temporary server responses selectively; do not retry every status indefinitely. Make redirect policy explicit, especially when URLs may leave the allowed host set.

Cost: The direct costs are compute, bandwidth, storage, and any API or infrastructure charges in your environment. Excess retries, duplicate variants, unnecessarily large bodies, and repeated full downloads increase those costs. The target site also bears request and processing load, which is another reason to use direct inventory sources and conservative limits.

Or skip the browser setup

If the URL list is intended for visual snapshots rather than HTML extraction, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. A single GET request captures a URL as PNG, JPEG, WebP, or PDF. Its API accepts bulk capture of up to 100 URLs per call, and the same capture options include full-page screenshots, element selectors, device presets, custom waits, and more. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

Cookie banners, newsletter popups, and chat widgets are removed before the shot, and each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers report the page verdict and billing status. The MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Every feature is on every plan.

Sign up for 1,000 free screenshots a month, with no card required.

FAQ

How many concurrent requests should my crawler make?

There is no universal number. Start with a low per-host limit, measure latency and overload responses, and tune for the target and your implementation.

Use the source that most directly supplies the inventory you need. A sitemap, API, export, or search endpoint can avoid spending requests discovering URLs, but check freshness, coverage, permissions, and quotas.

Does robots.txt tell me that I have permission to crawl?

No. Robots rules communicate crawler preferences; they are not authentication or legal authorization. Check the actual site terms and applicable constraints.

Can a 304 response be stored as an empty page?

No. It indicates that the cached representation can be reused. Keep the prior body and update its validation metadata as appropriate.

Sources