Image Extractor from HTML
Extract image URLs from HTML, including srcset and picture sources, with runnable Python, Node.js, cURL, browser, and troubleshooting guides.
Direct answer: parse every <img>, every <source> inside <picture>, and preserve each srcset candidate with its descriptor. Treat the result as a URL inventory, not proof of which responsive image the browser displayed. An img-only pass also misses CSS background images.
Responsive candidates are conditional. A browser can choose a different file according to viewport width, device pixel ratio, media conditions, and supported image type. The WHATWG HTML Images specification defines the markup model.
1. Decide what you want to extract
| Scope | Collected data | Use case |
|---|---|---|
| Markup inventory | img[src], img[srcset], and picture>source |
Audits and migrations |
| Browser choice | The resource selected for one viewport and DPR | Reproducing what a user sees |
| Loaded resources | Requests observed during rendering | Network analysis |
| Relevant images | A filtered subset such as article figures | Feeds and content extraction |
“All images” can mean different scopes. Keep every srcset URL when building an inventory. Use a browser when you need the currently selected resource.
2. Python extractor
python3 -m pip install beautifulsoup4
from __future__ import annotations
import json, sys
from urllib.parse import urljoin
from bs4 import BeautifulSoup
def split_srcset(value):
result = []
for part in (value or '').split(','):
bits = part.strip().split()
if bits:
result.append({'url': bits[0], 'descriptor': ' '.join(bits[1:])})
return result
def extract(html, page_url=None):
soup = BeautifulSoup(html, 'html.parser')
absolute = lambda value: urljoin(page_url, value) if page_url and value else value
records = []
for picture in soup.find_all('picture'):
sources = []
for source in picture.find_all('source', recursive=False):
sources.append({
'srcset': [{**x, 'url': absolute(x['url'])} for x in split_srcset(source.get('srcset'))],
'media': source.get('media'), 'type': source.get('type'), 'sizes': source.get('sizes')})
img = picture.find('img')
if img:
records.append({'tag': 'img', 'in_picture': True,
'src': absolute(img.get('src')), 'srcset': [{**x, 'url': absolute(x['url'])} for x in split_srcset(img.get('srcset'))],
'sizes': img.get('sizes'), 'alt': img.get('alt'), 'picture_sources': sources})
for img in soup.find_all('img'):
if img.find_parent('picture'):
continue
records.append({'tag': 'img', 'in_picture': False,
'src': absolute(img.get('src')), 'srcset': [{**x, 'url': absolute(x['url'])} for x in split_srcset(img.get('srcset'))],
'sizes': img.get('sizes'), 'alt': img.get('alt')})
return records
if __name__ == '__main__':
html = sys.stdin.read()
print(json.dumps(extract(html, sys.argv[1] if len(sys.argv) > 1 else None), indent=2))
curl -L https://example.com/article -o page.html
python3 extract_images.py https://example.com/article < page.html
The extractor resolves relative URLs when given the document URL, retains every candidate, and avoids counting a picture fallback twice.
3. Node.js extractor
npm install cheerio
import fs from 'node:fs';
import * as cheerio from 'cheerio';
const $ = cheerio.load(fs.readFileSync(0, 'utf8'));
const pageUrl = process.argv[2];
const absolute = value => pageUrl && value ? new URL(value, pageUrl).href : value;
const split = value => (value || '').split(',').map(x => x.trim()).filter(Boolean).map(x => { const [url, ...d] = x.split(/\s+/); return {url: absolute(url), descriptor: d.join(' ')}; });
const out = [];
$('picture').each((_, p) => {
const sources = [];
$(p).children('source').each((__, s) => { const e = $(s); sources.push({srcset: split(e.attr('srcset')), media: e.attr('media') || null, type: e.attr('type') || null, sizes: e.attr('sizes') || null}); });
const e = $(p).children('img').first();
if (e.length) out.push({tag: 'img', in_picture: true, src: absolute(e.attr('src')), srcset: split(e.attr('srcset')), sizes: e.attr('sizes') || null, alt: e.attr('alt') || null, picture_sources: sources});
});
$('img').each((_, node) => { if ($(node).parents('picture').length) return; const e = $(node); out.push({tag: 'img', in_picture: false, src: absolute(e.attr('src')), srcset: split(e.attr('srcset')), sizes: e.attr('sizes') || null, alt: e.attr('alt') || null}); });
console.log(JSON.stringify(out, null, 2));
curl -L https://example.com/article | node extract-images.mjs https://example.com/article
4. cURL and quick inspection
curl -L --fail --compressed https://example.com/article -o page.html
rg -o '<(img|source)\b[^>]*(src|srcset)\s*=\s*["'"'][^"'"']+' page.html
Regular expressions are suitable for triage only. Use an HTML parser for malformed markup, entities, nested elements, and production processing.
5. Find the browser-selected image
Read currentSrc after rendering at the viewport and device pixel ratio you care about:
python3 -m pip install playwright
playwright install chromium
from playwright.sync_api import sync_playwright
import json, sys
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page(viewport={'width': 1280, 'height': 800}, device_scale_factor=1)
page.goto(sys.argv[1], wait_until='networkidle')
print(json.dumps(page.eval_on_selector_all('img', "els => els.map(x => ({src:x.src, currentSrc:x.currentSrc, alt:x.alt}))"), indent=2))
browser.close()
currentSrc reports the selected resource in that browser context. CSS background-image values are outside an img-only parser. Inspect computed styles or network events if CSS coverage is required; external stylesheets, pseudo-elements, canvas, blob URLs, and JavaScript-created resources need site-specific handling.
6. Important extraction options
| Option | Recommended handling |
|---|---|
| Relative URLs | Resolve with the final document URL, while retaining the original attribute. |
srcset |
Store every URL and its w or x descriptor. |
picture |
Store each source’s media, type, and sizes, plus fallback img. |
| Duplicates | Deduplicate in a separate URL view; retain element order for audits. |
| Relevance | Filter by article container, figure, or site-specific rules after extraction. |
| CSS | Label whether CSS was ignored, parsed, or observed in a browser. |
7. Edge cases
- Lazy loading: markup may contain URLs before the browser requests them; browser observation depends on scrolling and timing.
- Art direction: picture sources can provide different crops or formats, so keep all alternatives.
- Missing
src: recordsrcsetand custom lazy-load attributes without inventing a URL. - Authentication or challenges: an HTTP fetch may return a login page or denial instead of the intended document.
- Data and blob URLs: preserve the scheme, but handle them separately from downloadable HTTP files.
- Canvas output: it may have no image URL in HTML and requires browser instrumentation.
8. Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| No images | Redirect, login page, or client-rendered shell. | Inspect curl -L -i output and use browser rendering for JavaScript content. |
| Only one responsive URL | srcset or picture was ignored. |
Parse all candidates and source elements. |
| Downloads fail | Relative references were not resolved. | Resolve against the final page URL. |
| Wrong image for a viewport | An inventory was mistaken for the browser selection. | Render at the target viewport and read currentSrc. |
| Backgrounds missing | They are CSS rather than img elements. |
Inspect computed styles or network requests. |
| Duplicate records | The picture fallback was counted twice. | Skip images whose ancestor is picture after emitting the picture record. |
| Navigation hangs | Slow origin or blocked resource. | Set explicit HTTP and browser timeouts and log failed URLs. |
9. Performance, reliability, and cost
- Static HTML parsing is cheaper and faster than browser automation for markup inventories.
- Use a browser only for selected candidates, JavaScript-created nodes, lazy loading, or computed CSS.
- Record final URL, status, retrieval time, parser version, and whether the result came from markup or browser observation.
- Cache documents and deduplicate downloads, while retaining source-element provenance.
- Only download assets when you have permission and need the files; an inventory often requires no asset download.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a clean PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API docs.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are never billed, and response headers identify the page verdict and billing result. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Create a free ScreenshotNeo account.
FAQ
Does src equal the displayed file?
Not always. Responsive selection can use a srcset candidate or picture source. Use currentSrc in a matching browser context.
How do I extract article images only?
Collect markup records first, then filter by the article container, figure elements, or other site-specific rules.
Can an HTML parser find every requested image?
No. CSS, JavaScript, canvas, blob URLs, and resources requested only during rendering require browser or network observation.
Should I download every extracted URL?
Only when permitted and necessary. A URL inventory is often sufficient for audits.


