How to Extract Images from an HTML File
Extract images from HTML with Python: handle src, srcset, picture, Base64 data, relative URLs, local files, and JavaScript-rendered pages.
Direct answer: Parse the HTML, inspect img and picture elements, collect src, srcset, and source URLs, resolve relative references, decode data: URIs, and save or download the resulting bytes. Static parsing cannot find images added only after JavaScript runs; render the page first in that case.
1. Choose the extraction method
Your method depends on where the image reference exists:
| Source | What to do |
|---|---|
| Local HTML file | Read it from disk and resolve local relative paths against the file directory. |
| Downloaded HTML | Use the page URL as the base for relative URLs. |
img src or srcset |
Collect every candidate URL, then deduplicate. |
picture |
Inspect each source srcset and the fallback img. |
| Base64 or percent-encoded data | Decode the data: URI instead of making an HTTP request. |
| JavaScript-created images | Use a browser renderer, save the post-render DOM, then parse it. |
Beautiful Soup documentation describes Beautiful Soup as a Python library for pulling data out of HTML and XML. Its parser choices include the built-in html.parser, lxml, and browser-like html5lib; malformed markup can produce different trees with different parsers.
2. Extract and download images from a static HTML file
Install the dependencies:
python -m pip install beautifulsoup4 requests
Save this as extract_images.py:
from pathlib import Path
from urllib.parse import urljoin, urlparse
from base64 import b64decode
import base64
import mimetypes
import re
import requests
from bs4 import BeautifulSoup
HTML_PATH = Path("page.html")
# Set this when page.html came from a website. Leave it None for a local archive.
BASE_URL = "https://example.com/articles/page.html"
OUTPUT = Path("extracted-images")
OUTPUT.mkdir(exist_ok=True)
html = HTML_PATH.read_text(encoding="utf-8")
soup = BeautifulSoup(html, "html.parser")
refs = []
def add_srcset(value):
if not value:
return
for candidate in value.split(","):
url = candidate.strip().split()[0]
if url:
refs.append(url)
for img in soup.find_all("img"):
if img.get("src"):
refs.append(img["src"])
add_srcset(img.get("srcset"))
for source in soup.select("picture source"):
add_srcset(source.get("srcset"))
if source.get("src"):
refs.append(source["src"])
# Keep order while removing duplicates.
refs = list(dict.fromkeys(refs))
SAFE_SCHEMES = {"http", "https", "file", "data"}
def extension_from_content_type(content_type):
media_type = content_type.split(";", 1)[0].strip().lower()
return mimetypes.guess_extension(media_type) or ".bin"
def save_data_uri(uri, index):
header, payload = uri.split(",", 1)
media_type = header[5:].split(";", 1)[0] or "application/octet-stream"
if ";base64" in header.lower():
data = b64decode(payload, validate=False)
else:
from urllib.parse import unquote_to_bytes
data = unquote_to_bytes(payload)
suffix = mimetypes.guess_extension(media_type) or ".bin"
path = OUTPUT / f"image-{index:04d}{suffix}"
path.write_bytes(data)
return path
for index, ref in enumerate(refs, 1):
if ref.startswith("data:"):
print(f"saved {save_data_uri(ref, index)}")
continue
if BASE_URL:
absolute = urljoin(BASE_URL, ref)
else:
absolute = (HTML_PATH.parent / ref).resolve().as_uri()
parsed = urlparse(absolute)
if parsed.scheme not in SAFE_SCHEMES:
print(f"skipped unsupported scheme: {ref}")
continue
if parsed.scheme == "file":
source_path = Path(parsed.path)
if source_path.is_file():
suffix = source_path.suffix or ".bin"
path = OUTPUT / f"image-{index:04d}{suffix}"
path.write_bytes(source_path.read_bytes())
print(f"saved {path}")
else:
print(f"missing local file: {source_path}")
continue
try:
response = requests.get(absolute, timeout=30)
response.raise_for_status()
content_type = response.headers.get("Content-Type", "")
suffix = extension_from_content_type(content_type)
path = OUTPUT / f"image-{index:04d}{suffix}"
path.write_bytes(response.content)
print(f"saved {path} ({len(response.content)} bytes)")
except requests.RequestException as exc:
print(f"failed {absolute}: {exc}")
Run it with python extract_images.py. The script handles src, responsive candidates, picture, inline data, relative URLs, local files, duplicate references, unsupported schemes, response status errors, and content-type-based extensions.
3. Extract URLs without downloading bytes
Sometimes you only need an inventory for auditing, migration, or later processing:
from bs4 import BeautifulSoup
from pathlib import Path
from urllib.parse import urljoin
html = Path("page.html").read_text(encoding="utf-8")
soup = BeautifulSoup(html, "html.parser")
base = "https://example.com/page.html"
for img in soup.find_all("img"):
values = [img.get("src", "")]
values += [part.strip().split()[0] for part in img.get("srcset", "").split(",") if part.strip()]
for value in dict.fromkeys(values):
if value:
print(urljoin(base, value))
for source in soup.select("picture source[srcset]"):
for part in source["srcset"].split(","):
print(urljoin(base, part.strip().split()[0]))
4. Understand srcset and picture
An img element normally uses src as its fallback and may provide several responsive alternatives through srcset. A candidate can have a width descriptor such as 800w or a density descriptor such as 2x. The extraction task is to collect the URL before the descriptor; choosing which candidate a browser would display requires viewport, density, and media-condition logic.
picture groups alternate resources. Each source can have media, type, and srcset; the nested img is the required fallback. If your goal is archival completeness, save every candidate. If your goal is the displayed image, reproduce the target viewport and browser selection rules.
Google’s image guidance documents responsive images and the img fallback, and describes common formats including BMP, GIF, JPEG, PNG, WebP, SVG, and AVIF.
5. Decode Base64 and other data URIs
A data: URI contains its bytes in the attribute itself. Its header identifies the media type and may include ;base64. Base64 payloads should be decoded with Base64 decoding; non-Base64 payloads should be percent-decoded. Do not send a data URI to requests.get.
from base64 import b64decode
from urllib.parse import unquote_to_bytes
uri = "data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAEAAAAB"
header, payload = uri.split(",", 1)
if ";base64" in header.lower():
image_bytes = b64decode(payload)
else:
image_bytes = unquote_to_bytes(payload)
open("inline-image.bin", "wb").write(image_bytes)
Validate the decoded bytes before trusting the extension. A filename or URL suffix is not authoritative; the HTTP Content-Type and file signature are better evidence.
6. Handle JavaScript-rendered pages
Python’s standard HTML parser exposes attributes, but content inside script and style is returned as-is rather than parsed as HTML. An image inserted after page load therefore will not appear in the original source. Use a browser automation tool to wait for rendering, then parse the resulting DOM or observe image network requests.
python -m pip install playwright beautifulsoup4
python -m playwright install chromium
import asyncio
from pathlib import Path
from urllib.parse import urljoin
from bs4 import BeautifulSoup
from playwright.async_api import async_playwright
async def main():
async with async_playwright() as p:
browser = await p.chromium.launch()
page = await browser.new_page()
await page.goto("https://example.com", wait_until="networkidle")
html = await page.content()
Path("rendered.html").write_text(html, encoding="utf-8")
soup = BeautifulSoup(html, "html.parser")
for img in soup.find_all("img"):
if img.get("src"):
print(urljoin(page.url, img["src"]))
await browser.close()
asyncio.run(main())
For lazy-loaded images, scroll or trigger the relevant section before calling page.content(). For a complete byte archive, capture the network responses while the page loads; the rendered DOM may contain a URL but not the response body.
7. Parser and file-handling choices
html.parser: no external parser dependency and a sensible starting point.lxml: generally useful when parsing speed matters and the dependency is acceptable.html5lib: browser-like error recovery for severely malformed HTML.
Different parsers can construct different trees from invalid markup. Keep the parser fixed for repeatable jobs, and record the original URL, extracted reference, resolved URL, response status, media type, byte length, and checksum.
8. Edge cases checklist
- Resolve relative paths against the document URL, including
../, root-relative, query-only, and fragment-bearing references. - Ignore fragments when downloading; they identify a document location, not different image bytes.
- Deduplicate after URL resolution, not only before it.
- Handle protocol-relative URLs such as
//cdn.example/image.webp. - Expect SVG, AVIF, WebP, GIF, and images served without a useful filename suffix.
- Do not follow unsafe schemes such as
javascript:. - Preserve query strings because CDNs often use them for image transforms or authorization.
- Use unique output names; never let a remote basename overwrite an earlier file.
- Respect access controls, robots policies, site terms, and image licenses. Extraction mechanics do not grant reuse rights.
- Some sites expose image URLs in custom attributes such as
data-srcordata-srcset; add those only when the site uses them and document the rule.
9. Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| No images found | Images are injected by JavaScript or stored in custom attributes. | Render the page, inspect the post-render DOM, and check lazy-load attributes. |
| 404 or wrong host | Relative URL was joined without the real document URL. | Pass the page URL as BASE_URL; do not use an invented base. |
| Only one responsive image saved | Only src was inspected. |
Parse every srcset candidate and each picture source. |
| Base64 download fails | A data URI was treated as HTTP. | Split at the first comma and decode the payload. |
| Files have wrong extensions | URL suffix is missing or misleading. | Use Content-Type and, for strict pipelines, inspect file signatures. |
| Parser misses malformed markup | Parser error recovery differs. | Compare html.parser, lxml, and html5lib. |
| 403 or login page saved | The server requires authentication or rejects the client. | Use authorized credentials and headers, or stop rather than bypassing access controls. |
| Timeouts and partial results | Slow origin, large files, or unstable network. | Set connect/read timeouts, retry only idempotent requests, and log failures for replay. |
10. Performance, reliability, and cost
For large files, stream responses to disk instead of holding every image in memory. Use a session to reuse HTTP connections, cap concurrency so you do not overload the origin, and write a manifest as each file completes. Cache by resolved URL and checksum when repeated pages are expected. Browser rendering costs more CPU and time than static parsing, so reserve it for pages whose image references are not present in the source.
Retries should distinguish transient network errors from permanent 4xx responses. Keep timeouts finite, record status and exception details, and make output names deterministic so a failed run can resume without overwriting valid files. Downloading an image may incur bandwidth or service charges on the source; this workflow has no universal cost estimate.
11. Or skip the browser setup
If you need a clean screenshot of a page rather than a local archive of every source image, ScreenshotNeo returns a PNG, JPEG, WebP, or PDF from one GET request. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.
See the ScreenshotNeo API documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
12. FAQ
Can I extract images from an HTML string?
Yes. Pass the string directly to Beautiful Soup instead of reading a file, then apply the same attribute and data-URI handling.
Should I save every srcset candidate?
Save every candidate for a complete asset inventory. Select one only when you intentionally reproduce a browser viewport and device density.
Does extracting an image allow me to republish it?
No. Check the image license, site terms, and applicable law before reuse.
Why is the downloaded file larger than expected?
The URL may point to an original asset while the browser displays a transformed CDN variant. Preserve the full query string and compare response headers.


