ScreenshotNeo

BlogHow-to

How to Download Images in Bulk from a List of URLs

Download images from a URL list with repeatable shell, Python, Node.js, cURL, and browser workflows, plus retries, naming, and troubleshooting.

By the ScreenshotNeo team1 October 20269 min read

Direct answer: Put one direct image URL on each line of a text file, then fetch those URLs with a script or command-line tool into a dedicated folder. For repeatable batches, control timeouts, concurrency, filenames, retries, and a cache. If you prefer a graphical workflow, a Chrome extension can ingest pasted URLs or CSV/TXT files and package results into a ZIP.

What you need before starting

  • A plain-text file such as urls.txt, with one URL per line.
  • Direct image URLs. A page URL such as https://example.com/gallery/123 may return HTML rather than an image.
  • A destination folder with enough disk space.
  • Permission to download and use each image. Check the image license and the rules that apply to your use and jurisdiction.

Start by inspecting a few entries. Remove blank lines, comments, tracking URLs you do not need, and duplicates. A minimal file looks like this:

https://images.example.com/catalog/red-shirt.jpg
https://cdn.example.org/photos/city.webp
https://static.example.net/assets/hero.png

Route 1: repeatable downloads with Python

Python’s standard library provides urllib.request.urlopen(), which accepts a URL string or a Request object and supports a timeout for blocking HTTP and HTTPS operations. The script below streams each response to disk, creates deterministic sequence-based names, skips duplicate URLs, and records failures.

Complete standard-library script

#!/usr/bin/env python3
from pathlib import Path
from urllib.parse import urlparse
from urllib.request import Request, urlopen
from urllib.error import HTTPError, URLError
import mimetypes
import re
import time

INPUT = Path("urls.txt")
OUTPUT = Path("downloaded-images")
TIMEOUT_SECONDS = 30
WAIT_SECONDS = 0.2

EXTENSIONS = {
    "image/jpeg": ".jpg",
    "image/png": ".png",
    "image/webp": ".webp",
    "image/gif": ".gif",
    "image/tiff": ".tif",
    "image/avif": ".avif",
}


def safe_name(value: str) -> str:
    value = re.sub(r"[^A-Za-z0-9._-]+", "_", value).strip("._")
    return value or "image"


def extension_for(url: str, content_type: str) -> str:
    media_type = content_type.split(";", 1)[0].lower()
    if media_type in EXTENSIONS:
        return EXTENSIONS[media_type]
    suffix = Path(urlparse(url).path).suffix.lower()
    return suffix if suffix in {".jpg", ".jpeg", ".png", ".webp", ".gif", ".tif", ".tiff", ".avif"} else ".bin"


def read_urls(path: Path) -> list[str]:
    seen = set()
    result = []
    for line in path.read_text(encoding="utf-8").splitlines():
        url = line.strip()
        if not url or url.startswith("#") or url in seen:
            continue
        seen.add(url)
        result.append(url)
    return result


def download(url: str, number: int) -> tuple[bool, str]:
    request = Request(url, headers={"User-Agent": "bulk-image-downloader/1.0"})
    try:
        with urlopen(request, timeout=TIMEOUT_SECONDS) as response:
            content_type = response.headers.get("Content-Type", "")
            extension = extension_for(url, content_type)
            destination = OUTPUT / f"{number:05d}{extension}"
            with destination.open("wb") as output:
                while chunk := response.read(1024 * 1024):
                    output.write(chunk)
        return True, str(destination)
    except (HTTPError, URLError, TimeoutError, OSError) as error:
        return False, f"{url}\t{error}"


OUTPUT.mkdir(parents=True, exist_ok=True)
urls = read_urls(INPUT)
failed = []
for number, url in enumerate(urls, start=1):
    ok, message = download(url, number)
    print(("OK   " if ok else "FAIL ") + message)
    if not ok:
        failed.append(message)
    time.sleep(WAIT_SECONDS)

if failed:
    (OUTPUT / "failed.txt").write_text("\n".join(failed) + "\n", encoding="utf-8")
    print(f"Completed with {len(failed)} failures; see {OUTPUT / 'failed.txt'}")
else:
    print(f"Downloaded {len(urls)} images to {OUTPUT}")

Run it

python3 bulk_download.py

The script names files by sequence, so two different URLs that end in image.jpg cannot overwrite one another. It uses the response’s content type when available and falls back to the URL suffix.

Retries and resumability

For a large or unreliable batch, add a retry loop around urlopen() with a short exponential backoff (for example, 1, 2, and 4 seconds), and write successful URLs to a manifest. On a rerun, skip entries already present in that manifest. Do not retry permanent errors such as repeated 404 responses.

Route 2: the imgdl command-line and Python API

The imgdl project documents a command-line utility that accepts a URL text file and a Python API that accepts a list or another iterator. Its documented options include an output folder, worker count, timeout, optional thumbnails, wait intervals, and a persistent cache that skips files already present unless forced.

A typical workflow is:

  1. Install the package using the installation command in its current README.
  2. Pass urls.txt and an output directory to the CLI.
  3. Set a timeout and a conservative worker count.
  4. Enable its cache for reruns, and inspect the documented failure result (a failed item can return a None path).

The README shows 50 workers as an example configuration. Treat that as an example, not a universal recommendation. Host limits, bandwidth, latency, and rate limiting determine a sensible value.

Using its Python interface

The documented API accepts an iterable, so a generator can supply URLs without loading a very large file into memory. Use the exact function and option names from the version you install because package interfaces can change.

from pathlib import Path


def urls(path):
    for line in Path(path).read_text(encoding="utf-8").splitlines():
        value = line.strip()
        if value and not value.startswith("#"):
            yield value

# Consult the installed imgdl README for the current import and call signature.
# The documented options include output folder, workers, timeout, wait, thumbnails,
# and persistent caching.
# results = imgdl.download(urls("urls.txt"), output="downloaded-images", workers=8, timeout=30)

Route 3: cURL for a small or scripted batch

cURL is useful when you want a simple shell loop and already know the filenames. The -f flag makes HTTP errors fail, -L follows redirects, and --retry retries transient failures.

mkdir -p downloaded-images
n=0
while IFS= read -r url; do
  [ -z "$url" ] && continue
  case "$url" in \#*) continue ;; esac
  n=$((n + 1))
  printf -v file 'downloaded-images/%05d' "$n"
  curl --fail --location --retry 3 --retry-delay 1 \
       --connect-timeout 10 --max-time 60 \
       --remote-header-name --remote-name \
       --output "$file" "$url" || echo "$url" >> failed.txt
  sleep 0.2
done < urls.txt

When you need reliable extensions and collision-proof names, use the Python script instead of depending on remote filenames.

Route 4: Node.js

Modern Node.js includes fetch. This example downloads sequentially, follows the URL’s path suffix when possible, and records failures.

import { readFile, mkdir, writeFile } from 'node:fs/promises';
import path from 'node:path';

const urls = (await readFile('urls.txt', 'utf8'))
  .split(/\r?\n/)
  .map(s => s.trim())
  .filter(s => s && !s.startsWith('#'));
await mkdir('downloaded-images', { recursive: true });
const failures = [];

for (let i = 0; i < urls.length; i++) {
  const url = urls[i];
  try {
    const response = await fetch(url, { signal: AbortSignal.timeout(30000) });
    if (!response.ok) throw new Error(`HTTP ${response.status}`);
    const extension = path.extname(new URL(url).pathname) || '.bin';
    const output = path.join('downloaded-images', `${String(i + 1).padStart(5, '0')}${extension}`);
    await writeFile(output, Buffer.from(await response.arrayBuffer()));
    console.log(`OK   ${output}`);
  } catch (error) {
    failures.push(`${url}\t${error.message}`);
    console.error(`FAIL ${url}: ${error.message}`);
  }
}
if (failures.length) await writeFile('downloaded-images/failed.txt', failures.join('\n') + '\n');

Browser extension workflow

If you prefer a graphical interface, the Chrome Web Store listing for Bulk Image Downloader From URL List describes pasted URL lists plus CSV/TXT ingestion, filtering, filename controls, parallel downloads, and ZIP packaging. These are publisher-listed features, not an independent speed test.

  1. Install the extension from the Chrome Web Store.
  2. Paste the URL list or import the documented CSV/TXT format.
  3. Filter entries and configure filename and destination options.
  4. Choose parallel downloads or ZIP output if those options fit your batch.
  5. Review the output and retry only failed entries.

Choosing filenames, folders, and formats

  • Dedicated folder: Keep each batch separate so you can count and review files.
  • Duplicate paths: Two URLs can share a basename. Use sequence numbers, URL hashes, or another deterministic key.
  • Extensions: Prefer the HTTP Content-Type header; a URL ending in .jpg can still return HTML or a different format.
  • Integrity: Compare the number of successful responses with the number of input URLs and open a sample from the beginning, middle, and end.

Concurrency, timeouts, and reliability

More workers can reduce elapsed time when the network and hosts allow it, but no researched source establishes a best worker count or comparative download speed. Start conservatively, then adjust while watching timeouts, HTTP 429 responses, connection errors, and the remote site’s terms.

Control Practical starting point Why it matters
Connect/read timeout 30–60 seconds Prevents one stalled host from blocking the batch.
Workers Small number such as 4–8 Limits pressure on your network and the origin.
Wait interval 100–500 ms Spreads requests when a tool supports pauses.
Retries 2–3 with backoff Helps transient failures without hammering a host.
Cache/manifest Enabled for reruns Avoids downloading completed files again.

Troubleshooting

Symptom Likely cause Fix
Downloaded file is HTML The URL points to a page, login screen, or error document. Open the URL directly, inspect Content-Type, and obtain the actual image URL.
404 or 410 The resource was removed or the path is wrong. Verify the URL; do not repeatedly retry a permanent error.
403 Hotlink protection, authentication, or access rules. Use an authorized request with the required headers or obtain permission; do not bypass access controls.
429 Rate limiting. Reduce workers, add delays, and honor the site’s retry guidance.
Timeouts Slow origin, large file, or network problem. Increase the timeout moderately, retry with backoff, and keep a failed-URL log.
Files overwrite each other Names were derived only from the basename. Use sequence numbers, URL hashes, or collision checks.
Some results are missing Blank rows, duplicate filtering, or unrecorded failures. Count inputs, successes, and failures; preserve a manifest and review failed.txt.
Images look corrupted Partial transfer or an error response saved as an image. Use streaming writes, check HTTP status and content type, and redownload the affected file.

Cost and performance notes

Your main costs are bandwidth, storage, and any service or browser resources used to process the batch. The supplied research describes tool controls but contains no independent benchmark, so measure a representative sample in your own environment. For very large lists, process chunks, persist a manifest, and keep failed URLs separate from successful output.

Or skip the browser setup

When the task is capturing pages or assets as clean screenshots rather than downloading original image files, ScreenshotNeo provides a single HTTP request. Its API can return PNG, JPEG, WebP, or PDF, and the documentation lists the available options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners, newsletter popups, and chat widgets are removed before the shot.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify the page verdict and billing status.
  • An MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf.
  • Free includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account to try it with 1,000 screenshots a month and no card.

FAQ

Can I put page URLs in the list?

Only if the tool or script is designed to extract images from pages. A bulk downloader expects direct image resources; page URLs commonly return HTML.

Should I use 50 workers?

No universal value is established. The imgdl README uses 50 as an example. Begin with a smaller value and increase only when the hosts and your network handle it reliably.

How do I resume a failed batch?

Keep a success manifest or enable the tool’s persistent cache, then retry only failed URLs. Confirm that existing files are complete before skipping them.

Can I download any image I can reach?

Reachability does not establish permission. Check the image’s license and applicable rules before downloading or redistributing it.