How to Download All Images from an HTML File
Learn how to extract and download images from a local HTML file or remote page with Wget, Python, Node.js, cURL, and browser rendering.
Short answer: If you have a remote webpage, start with GNU Wget’s --page-requisites option. If you have a local .html file or need precise filtering and filenames, parse the markup, resolve each image URL against the page URL, and download the files with an HTTP client. If JavaScript inserts the images, render the page in a browser before extracting URLs.
“All images” needs a definition. A basic script can collect <img src> values, while a more complete process also considers srcset, <picture> sources, CSS backgrounds, lazy-loading attributes, data URLs, and images revealed after scrolling or interaction.
1. Choose the method that matches your HTML
| Source and goal | Recommended method | What it finds | Main limitation |
|---|---|---|---|
| One remote page and its display assets | GNU Wget -p |
Resources referenced by the page’s HTML and CSS, including inline images and referenced stylesheets | It does not promise to discover every image revealed by runtime interaction |
| Local HTML file or custom filenames | Beautiful Soup plus an HTTP client | Markup-selected URLs with your own filtering, deduplication, and naming rules | Static parsing does not execute JavaScript |
| JavaScript-populated page | Browser rendering, then extraction | Image elements present after scripts run, scrolling, or other scripted actions | Requires a browser and site-specific handling |
| One known image URL | cURL | The URL you explicitly provide | cURL does not parse HTML and discover image URLs by itself |
2. Download assets from a remote page with Wget
GNU Wget documents --page-requisites (short form -p) for retrieving files needed to properly display a given HTML page, including inline images and referenced stylesheets. For a single page, begin without recursive crawling:
wget -p "https://example.com/page.html"
The command saves the page and its required resources in a local directory structure. It is useful when the page is remote and you want the browser-visible assets referenced by the initial HTML or CSS. See the GNU Wget manual for the documented option behavior.
Save a locally viewable copy
When you also want link conversion and a self-contained local rendering, Wget documents this combination:
wget -E -H -k -K -p "https://example.com/page.html"
-p is the relevant asset-download option. The additional flags change how the saved page and links are written; they do not turn the command into a recursive gallery crawler.
3. Download images referenced by a local HTML file with Python
Beautiful Soup can parse an open local file and expose tags and attributes. The following script handles regular src values, common lazy-loading attributes, srcset, <picture> sources, URL resolution, duplicate URLs, sensible filenames, and per-file failures.
Install dependencies
python -m pip install beautifulsoup4 requests
Complete script
from pathlib import Path
from urllib.parse import urljoin, urlparse
import mimetypes
import re
import requests
from bs4 import BeautifulSoup
HTML_FILE = Path("page.html")
OUTPUT_DIR = Path("downloaded-images")
# Use the URL that originally served the saved HTML. It is required for
# resolving relative paths such as /images/photo.jpg or ../hero.webp.
BASE_URL = "https://example.com/path/page.html"
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
with HTML_FILE.open("rb") as handle:
soup = BeautifulSoup(handle, "html.parser")
# Respect an HTML base element when one exists.
base_tag = soup.find("base", href=True)
base_url = urljoin(BASE_URL, base_tag["href"]) if base_tag else BASE_URL
candidates = []
for tag in soup.find_all(["img", "source"]):
for attribute in ("src", "data-src", "data-lazy-src", "data-original"):
value = tag.get(attribute)
if value:
candidates.append(value.strip())
break
srcset = tag.get("srcset")
if srcset:
for item in srcset.split(","):
url = item.strip().split()[0]
if url:
candidates.append(url)
seen = set()
for index, raw_url in enumerate(candidates, start=1):
if raw_url.startswith("data:"):
print(f"Skipping data URL #{index}")
continue
image_url = urljoin(base_url, raw_url)
if image_url in seen:
continue
seen.add(image_url)
try:
response = requests.get(image_url, timeout=30)
response.raise_for_status()
except requests.RequestException as error:
print(f"FAILED {image_url}: {error}")
continue
parsed = urlparse(image_url)
name = Path(parsed.path).name
name = re.sub(r"[^A-Za-z0-9._-]", "_", name)
if not name or name in {".", ".."}:
name = f"image-{index}"
# Add an extension when the URL has none and the server supplies a type.
if "." not in name:
extension = mimetypes.guess_extension(
response.headers.get("Content-Type", "").split(";", 1)[0]
)
if extension:
name += extension
destination = OUTPUT_DIR / f"{index:04d}-{name}"
destination.write_bytes(response.content)
print(f"Saved {destination} <- {image_url}")
print(f"Downloaded {len(seen)} unique image URLs; files are in {OUTPUT_DIR}")
The script is intentionally explicit about its scope. It does not claim to find CSS background-image values, images created only after JavaScript runs, or every possible custom lazy-loading attribute. Add those rules when the target site requires them.
Why URL resolution matters
A value such as /images/photo.jpg is a path, not a complete address. Resolve it with the page URL and any applicable HTML <base> element. The requests-html documentation describes absolute-link and base-URL handling that illustrates this distinction.
4. Download selected images with Node.js
Node.js does not include an HTML parser, so use a parser package for reliable markup handling. This example uses Cheerio and the built-in fetch available in modern Node.js versions.
Install and run
npm install cheerio
node download-images.mjs page.html https://example.com/path/page.html
import fs from "node:fs/promises";
import path from "node:path";
import { fileURLToPath } from "node:url";
import * as cheerio from "cheerio";
const [, , htmlPath = "page.html", suppliedBaseUrl] = process.argv;
const html = await fs.readFile(htmlPath, "utf8");
const $ = cheerio.load(html);
const baseHref = $("base[href]").first().attr("href");
const baseUrl = new URL(baseHref || suppliedBaseUrl || "file:///" , suppliedBaseUrl).href;
const outputDir = "downloaded-images";
await fs.mkdir(outputDir, { recursive: true });
const urls = new Set();
$("img, source").each((_, element) => {
const node = $(element);
for (const attribute of ["src", "data-src", "data-lazy-src", "data-original"]) {
const value = node.attr(attribute);
if (value) urls.add(new URL(value.trim(), baseUrl).href);
}
const srcset = node.attr("srcset");
if (srcset) {
for (const item of srcset.split(",")) {
const value = item.trim().split(/\s+/)[0];
if (value) urls.add(new URL(value, baseUrl).href);
}
}
});
let index = 0;
for (const imageUrl of urls) {
if (imageUrl.startsWith("data:")) continue;
index += 1;
try {
const response = await fetch(imageUrl);
if (!response.ok) throw new Error(`${response.status} ${response.statusText}`);
const bytes = Buffer.from(await response.arrayBuffer());
const parsed = new URL(imageUrl);
let filename = path.basename(parsed.pathname) || `image-${index}`;
filename = filename.replace(/[^A-Za-z0-9._-]/g, "_");
await fs.writeFile(path.join(outputDir, `${String(index).padStart(4, "0")}-${filename}`), bytes);
console.log(`Saved ${imageUrl}`);
} catch (error) {
console.error(`Failed ${imageUrl}: ${error.message}`);
}
}
console.log(`Processed ${urls.size} unique URLs`);
For production use, add retry policy, concurrency limits, response-size limits, content-type checks, and authentication headers appropriate to the site.
5. Fetch a known image with cURL
cURL is suitable when you already know the image URL. Its tutorial documents saving a specified URL with -o or using the remote name with -O; it does not discover image URLs from an HTML document.
curl -L "https://example.com/images/photo.jpg" -o photo.jpg
curl -L -O "https://example.com/images/photo.jpg"
6. Handle JavaScript-generated images
Parsing the downloaded HTML only sees the initial markup. A site may create image elements after JavaScript runs, replace placeholders while scrolling, or request URLs through an API. In that case, render the page in a browser-capable tool and inspect the resulting DOM.
The requests-html documentation describes render(), which reloads a response in Chromium, executes JavaScript, and supports options such as waiting and scrolling. A minimal outline is:
from requests_html import HTMLSession
session = HTMLSession()
response = session.get("https://example.com/page.html", timeout=30)
response.html.render(timeout=30, wait=2, scrolldown=5)
for image in response.html.find("img"):
print(image.attrs.get("src"))
Rendering adds a browser dependency and still may not reproduce every interaction, login state, consent flow, or anti-bot challenge. Render only when static extraction is insufficient.
7. Decide what “all images” includes
- Regular images:
<img src>values. - Responsive images: every candidate in
srcsetand<picture><source>. Downloading every candidate can create several sizes of the same image. - Lazy images: site-specific attributes such as
data-src,data-original, or placeholders that are replaced on scroll. - CSS backgrounds: URLs in stylesheets, inline
styleattributes, and pseudo-elements. Wget can retrieve CSS-referenced resources; a simple HTML parser will not. - Data URLs: embedded images beginning with
data:. They do not need an HTTP request, but require decoding and a generated extension. - Runtime or interaction-only images: content that appears after JavaScript, scrolling, clicking, login, or an API call. Use a browser workflow and define the interaction sequence.
8. Common errors and fixes
| Error or symptom | Likely cause | Fix |
|---|---|---|
| Every download returns 404 | A relative path was requested as though it were absolute | Resolve each value with urljoin or new URL(value, baseUrl); honor <base>. |
| No images are found | Images are inserted by JavaScript or stored in lazy attributes | Inspect data-* attributes, then render the page and extract the post-render DOM. |
| Only low-resolution files download | The parser selected src while the desired candidate is in srcset |
Parse srcset and choose a candidate based on the target width, or intentionally download all candidates. |
| 403 or 401 responses | The server requires authentication, cookies, a user agent, or blocks automated requests | Use authorized credentials and appropriate headers, slow the request rate, and follow the site’s access rules. |
| Files overwrite each other | Different URLs share the same basename | Prefix names with an index or hash and keep the URL in a manifest. |
| HTML is parsed incorrectly | Malformed markup or an unsuitable parser | Try Beautiful Soup’s html5lib for browser-like leniency or lxml for speed, accepting their dependency trade-offs. |
| Downloads hang | No timeout, stalled server, or very large response | Set connect/read timeouts, cap response sizes, and retry only transient failures. |
| Wget downloads fewer files than expected | The missing images are runtime-created, interaction-gated, or outside recognized HTML/CSS references | Use browser rendering or inspect the site’s network/API requests. |
9. Performance, reliability, and cost considerations
- Bound concurrency: downloading hundreds of files in parallel can overload your connection or the origin. Use a small worker pool and honor rate limits.
- Reuse connections: a persistent HTTP session reduces connection setup overhead.
- Retry selectively: retry timeouts and temporary 5xx responses with backoff; do not blindly retry permanent 4xx errors.
- Validate responses: check status, content type, and a maximum size before writing files. Some servers return an HTML error page with a successful status.
- Keep a manifest: record source URL, local filename, status, content type, and failure reason so a later run can resume.
- Deduplicate deliberately: URL deduplication avoids repeated requests; content hashing finds identical files served from different URLs.
- Respect access and rights: download only material you are entitled to keep and follow the target site’s access rules.
- Cost: local Wget, Python, Node.js, and cURL use your own compute and network resources. Browser rendering generally consumes more memory and startup time than static parsing.
10. Or skip the browser setup
If your actual goal is a reliable image of the page rather than extracting each original image file, ScreenshotNeo provides a website screenshot API. It accepts one GET request and returns a PNG, JPEG, WebP, or PDF. The API can wait for a selector, delay, or network idle; load lazy images with full-page capture; use custom headers, cookies, user agents, timezone, and geolocation; block ads, trackers, requests, or resource types; and capture one CSS-selected element.
See the ScreenshotNeo API documentation for request options. A minimal call is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing result. An MCP server lets Claude, Cursor, and other MCP clients use take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.
Create a free ScreenshotNeo account and start with 1,000 screenshots per month at no charge.
11. FAQ
Can Wget download every image visible in a browser?
No. Its documented page-requisites mode follows recognized HTML and CSS references needed to display a page. Runtime-created or interaction-only images may require rendering.
Should I download every URL in a srcset?
Only if you need every responsive variant. Otherwise choose the candidate that matches your target display width to avoid duplicate sizes.
Why does a local HTML file need a base URL?
Relative paths have meaning only in the context of the page that referenced them. Supply the original page URL or an appropriate <base> value.
Can cURL discover images from HTML?
No. cURL transfers a URL you provide. Pair it with a parser or use Wget for a remote page’s recognized display resources.
What should I do when the page requires login?
Use an authorized session with the required cookies or headers, and make sure your download process complies with the site’s rules and your access rights.


