ScreenshotNeo

BlogHow-to

Image Extractor from HTML

Extract image URLs from HTML, including srcset and picture sources, with runnable Python, Node.js, cURL, browser, and troubleshooting guides.

By the ScreenshotNeo team1 October 20266 min read

Direct answer: parse every <img>, every <source> inside <picture>, and preserve each srcset candidate with its descriptor. Treat the result as a URL inventory, not proof of which responsive image the browser displayed. An img-only pass also misses CSS background images.

Responsive candidates are conditional. A browser can choose a different file according to viewport width, device pixel ratio, media conditions, and supported image type. The WHATWG HTML Images specification defines the markup model.

1. Decide what you want to extract

Scope Collected data Use case
Markup inventory img[src], img[srcset], and picture>source Audits and migrations
Browser choice The resource selected for one viewport and DPR Reproducing what a user sees
Loaded resources Requests observed during rendering Network analysis
Relevant images A filtered subset such as article figures Feeds and content extraction

“All images” can mean different scopes. Keep every srcset URL when building an inventory. Use a browser when you need the currently selected resource.

2. Python extractor

python3 -m pip install beautifulsoup4
from __future__ import annotations
import json, sys
from urllib.parse import urljoin
from bs4 import BeautifulSoup

def split_srcset(value):
    result = []
    for part in (value or '').split(','):
        bits = part.strip().split()
        if bits:
            result.append({'url': bits[0], 'descriptor': ' '.join(bits[1:])})
    return result

def extract(html, page_url=None):
    soup = BeautifulSoup(html, 'html.parser')
    absolute = lambda value: urljoin(page_url, value) if page_url and value else value
    records = []
    for picture in soup.find_all('picture'):
        sources = []
        for source in picture.find_all('source', recursive=False):
            sources.append({
                'srcset': [{**x, 'url': absolute(x['url'])} for x in split_srcset(source.get('srcset'))],
                'media': source.get('media'), 'type': source.get('type'), 'sizes': source.get('sizes')})
        img = picture.find('img')
        if img:
            records.append({'tag': 'img', 'in_picture': True,
                'src': absolute(img.get('src')), 'srcset': [{**x, 'url': absolute(x['url'])} for x in split_srcset(img.get('srcset'))],
                'sizes': img.get('sizes'), 'alt': img.get('alt'), 'picture_sources': sources})
    for img in soup.find_all('img'):
        if img.find_parent('picture'):
            continue
        records.append({'tag': 'img', 'in_picture': False,
            'src': absolute(img.get('src')), 'srcset': [{**x, 'url': absolute(x['url'])} for x in split_srcset(img.get('srcset'))],
            'sizes': img.get('sizes'), 'alt': img.get('alt')})
    return records

if __name__ == '__main__':
    html = sys.stdin.read()
    print(json.dumps(extract(html, sys.argv[1] if len(sys.argv) > 1 else None), indent=2))
curl -L https://example.com/article -o page.html
python3 extract_images.py https://example.com/article < page.html

The extractor resolves relative URLs when given the document URL, retains every candidate, and avoids counting a picture fallback twice.

3. Node.js extractor

npm install cheerio
import fs from 'node:fs';
import * as cheerio from 'cheerio';
const $ = cheerio.load(fs.readFileSync(0, 'utf8'));
const pageUrl = process.argv[2];
const absolute = value => pageUrl && value ? new URL(value, pageUrl).href : value;
const split = value => (value || '').split(',').map(x => x.trim()).filter(Boolean).map(x => { const [url, ...d] = x.split(/\s+/); return {url: absolute(url), descriptor: d.join(' ')}; });
const out = [];
$('picture').each((_, p) => {
  const sources = [];
  $(p).children('source').each((__, s) => { const e = $(s); sources.push({srcset: split(e.attr('srcset')), media: e.attr('media') || null, type: e.attr('type') || null, sizes: e.attr('sizes') || null}); });
  const e = $(p).children('img').first();
  if (e.length) out.push({tag: 'img', in_picture: true, src: absolute(e.attr('src')), srcset: split(e.attr('srcset')), sizes: e.attr('sizes') || null, alt: e.attr('alt') || null, picture_sources: sources});
});
$('img').each((_, node) => { if ($(node).parents('picture').length) return; const e = $(node); out.push({tag: 'img', in_picture: false, src: absolute(e.attr('src')), srcset: split(e.attr('srcset')), sizes: e.attr('sizes') || null, alt: e.attr('alt') || null}); });
console.log(JSON.stringify(out, null, 2));
curl -L https://example.com/article | node extract-images.mjs https://example.com/article

4. cURL and quick inspection

curl -L --fail --compressed https://example.com/article -o page.html
rg -o '<(img|source)\b[^>]*(src|srcset)\s*=\s*["'"'][^"'"']+' page.html

Regular expressions are suitable for triage only. Use an HTML parser for malformed markup, entities, nested elements, and production processing.

5. Find the browser-selected image

Read currentSrc after rendering at the viewport and device pixel ratio you care about:

python3 -m pip install playwright
playwright install chromium
from playwright.sync_api import sync_playwright
import json, sys
with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page(viewport={'width': 1280, 'height': 800}, device_scale_factor=1)
    page.goto(sys.argv[1], wait_until='networkidle')
    print(json.dumps(page.eval_on_selector_all('img', "els => els.map(x => ({src:x.src, currentSrc:x.currentSrc, alt:x.alt}))"), indent=2))
    browser.close()

currentSrc reports the selected resource in that browser context. CSS background-image values are outside an img-only parser. Inspect computed styles or network events if CSS coverage is required; external stylesheets, pseudo-elements, canvas, blob URLs, and JavaScript-created resources need site-specific handling.

6. Important extraction options

Option Recommended handling
Relative URLs Resolve with the final document URL, while retaining the original attribute.
srcset Store every URL and its w or x descriptor.
picture Store each source’s media, type, and sizes, plus fallback img.
Duplicates Deduplicate in a separate URL view; retain element order for audits.
Relevance Filter by article container, figure, or site-specific rules after extraction.
CSS Label whether CSS was ignored, parsed, or observed in a browser.

7. Edge cases

  • Lazy loading: markup may contain URLs before the browser requests them; browser observation depends on scrolling and timing.
  • Art direction: picture sources can provide different crops or formats, so keep all alternatives.
  • Missing src: record srcset and custom lazy-load attributes without inventing a URL.
  • Authentication or challenges: an HTTP fetch may return a login page or denial instead of the intended document.
  • Data and blob URLs: preserve the scheme, but handle them separately from downloadable HTTP files.
  • Canvas output: it may have no image URL in HTML and requires browser instrumentation.

8. Troubleshooting

Symptom Cause Fix
No images Redirect, login page, or client-rendered shell. Inspect curl -L -i output and use browser rendering for JavaScript content.
Only one responsive URL srcset or picture was ignored. Parse all candidates and source elements.
Downloads fail Relative references were not resolved. Resolve against the final page URL.
Wrong image for a viewport An inventory was mistaken for the browser selection. Render at the target viewport and read currentSrc.
Backgrounds missing They are CSS rather than img elements. Inspect computed styles or network requests.
Duplicate records The picture fallback was counted twice. Skip images whose ancestor is picture after emitting the picture record.
Navigation hangs Slow origin or blocked resource. Set explicit HTTP and browser timeouts and log failed URLs.

9. Performance, reliability, and cost

  • Static HTML parsing is cheaper and faster than browser automation for markup inventories.
  • Use a browser only for selected candidates, JavaScript-created nodes, lazy loading, or computed CSS.
  • Record final URL, status, retrieval time, parser version, and whether the result came from markup or browser observation.
  • Cache documents and deduplicate downloads, while retaining source-element provenance.
  • Only download assets when you have permission and need the files; an inventory often requires no asset download.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a clean PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API docs.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are never billed, and response headers identify the page verdict and billing result. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Create a free ScreenshotNeo account.

FAQ

Does src equal the displayed file?

Not always. Responsive selection can use a srcset candidate or picture source. Use currentSrc in a matching browser context.

How do I extract article images only?

Collect markup records first, then filter by the article container, figure elements, or other site-specific rules.

Can an HTML parser find every requested image?

No. CSS, JavaScript, canvas, blob URLs, and resources requested only during rendering require browser or network observation.

Should I download every extracted URL?

Only when permitted and necessary. A URL inventory is often sufficient for audits.