ScreenshotNeo

BlogHTML to image & PDF

How to Extract Images from PDFs with an API

Learn how to extract embedded PDF images with Adobe PDF Extract or PyMuPDF, handle duplicates and masks, and choose the right workflow.

By the ScreenshotNeo team1 October 20269 min read

Direct answer: Use a hosted extraction API when you need structured document data and managed processing. Adobe PDF Extract provides structured JSON with extracted figures as PNG files, or PDF-to-Markdown with figures embedded as base64. Use PyMuPDF when you want local processing, direct access to image bytes and metadata, and control over document handling.

The reliable workflow is:

  1. Keep credentials on a server, never in browser code.
  2. Upload the PDF or open it locally.
  3. Run extraction.
  4. Save the returned image bytes using the reported format.
  5. Deduplicate image references when one embedded image appears on multiple pages.
  6. Handle stencil masks when transparency matters.

Choose a hosted API or local extraction

Requirement Hosted extraction API PyMuPDF
Structured document elements Adobe returns structured JSON and image elements. You assemble the structure in application code.
LLM or Markdown pipeline Adobe PDF-to-Markdown can embed figures as base64. Generate Markdown yourself after extraction.
Data handling The documented flow uploads the PDF to a cloud service. Can run in a local application workflow.
Image bytes and metadata Returned in the extraction result. Available directly through page blocks or image xrefs.
Implementation languages Adobe lists Node.js, Python, .NET and Java SDKs. The examples below use Python.
Operational work Credentials, upload, asynchronous job, polling or webhook, download. Install a library and process the file in your application.

There is no independent accuracy benchmark in the supplied research. Choose based on output structure, deployment, privacy requirements, language integration and how much duplicate or mask handling your application needs.

Hosted extraction with Adobe PDF Extract

Adobe’s documented REST sequence has five stages:

  1. Create API credentials and obtain an access token.
  2. Request an asset upload URI.
  3. Upload the PDF and retain its asset ID.
  4. Submit an Extract PDF operation.
  5. Poll the operation location, or configure a completion webhook, then download the result.

Keep the client secret and access token in a server-side secret store. The documented service is intended for server-based use cases; do not place credentials in a browser or mobile client.

Select the output that matches your consumer

  • Structured JSON / Extract PDF: use this when downstream code needs element types, page relationships and separate image files. Adobe documents extracted figures as PNG files in this mode.
  • PDF-to-Markdown: use this when Markdown is the next stage, such as an LLM ingestion pipeline. Figures can be embedded as base64 data. A consumer that expects standalone files must decode those values first.

Command-line workflow shape

The exact Adobe endpoint paths and request fields depend on the current PDF Services API documentation and your credential type. The following shell sequence shows the required stages without putting a secret in source control:

export ADOBE_ACCESS_TOKEN='replace-on-the-server'
export PDF_FILE='./input.pdf'

# 1. Request an upload URI and asset identifier using your Adobe credentials.
# 2. Upload the PDF bytes to that URI.
# 3. Submit an Extract PDF job for the returned asset identifier.
# 4. Poll the operation location until the job is complete.
# 5. Download the result URI and unpack the extracted images.

# Keep the token in an environment variable or secret manager.
# Never commit it or send it to browser JavaScript.

Adobe advertises a free tier of 500 Document Transactions per month on its PDF Extract API product page. Treat that allowance as a changeable vendor term and verify it before estimating production cost.

Extract embedded images locally with PyMuPDF

PyMuPDF exposes two useful approaches:

  • page.get_text("dict") returns page blocks. Blocks with type == 1 contain image bytes, dimensions and an extension.
  • page.get_images() returns image references. Pass each xref to doc.extract_image(xref) to obtain the original bytes and reported extension.

Complete page-oriented extractor

#!/usr/bin/env python3
"""Extract images from every page of a PDF."""
from pathlib import Path
import sys
import fitz  # PyMuPDF


def extract_page_images(pdf_path: str, output_dir: str) -> int:
    source = Path(pdf_path)
    destination = Path(output_dir)
    destination.mkdir(parents=True, exist_ok=True)

    count = 0
    with fitz.open(source) as document:
        for page_number, page in enumerate(document, start=1):
            blocks = page.get_text("dict").get("blocks", [])
            for block_number, block in enumerate(blocks, start=1):
                if block.get("type") != 1:
                    continue

                image_bytes = block.get("image")
                extension = block.get("ext") or "bin"
                if not image_bytes:
                    continue

                output = destination / f"page-{page_number:04d}-image-{block_number:04d}.{extension}"
                output.write_bytes(image_bytes)
                count += 1
                print(f"wrote {output} ({block.get('width')}x{block.get('height')})")

    return count


if __name__ == "__main__":
    if len(sys.argv) != 3:
        raise SystemExit("usage: extract_pdf_images.py input.pdf output-directory")
    total = extract_page_images(sys.argv[1], sys.argv[2])
    print(f"extracted {total} page image block(s)")

Run it with:

python -m pip install PyMuPDF
python extract_pdf_images.py input.pdf extracted-images

Use the extension returned by PyMuPDF. Do not label every output as PNG: embedded data may be JPEG, PNG, BMP, TIFF or another supported format.

Extract each underlying image object once

A PDF can reference the same image object on several pages. Deduplicate by xref when you want one file per underlying image rather than one file per page occurrence.

from pathlib import Path
import fitz


def extract_unique_images(pdf_path: str, output_dir: str) -> int:
    destination = Path(output_dir)
    destination.mkdir(parents=True, exist_ok=True)
    seen = set()
    written = 0

    with fitz.open(pdf_path) as document:
        for page_number, page in enumerate(document, start=1):
            for image_index, image_info in enumerate(page.get_images(full=True), start=1):
                xref = image_info[0]
                if xref in seen:
                    continue
                seen.add(xref)

                extracted = document.extract_image(xref)
                extension = extracted.get("ext", "bin")
                output = destination / f"xref-{xref}.{extension}"
                output.write_bytes(extracted["image"])
                written += 1
                print(f"page {page_number}, image {image_index}: {output}")

    return written


if __name__ == "__main__":
    print(extract_unique_images("input.pdf", "unique-images"))

Important PDF image edge cases

Repeated references

The same xref can occur on multiple pages. Track xrefs in a set when deduplication is required. Do not deduplicate by filename alone because different xrefs can have the same dimensions or extension.

Stencil masks and transparency

A stencil mask may store transparency separately from the base image. Extracting only the base image can produce an incorrect result, such as a black or opaque background. If transparent artwork matters, inspect the image metadata and combine the mask with the base image using an image-processing library.

Rendered drawings are not embedded images

Charts, text and vector illustrations may not exist as raster image objects. An embedded-image extractor will not return them as separate PNG files. Render the relevant page to an image when the requirement is a visual snapshot of everything displayed on the page.

Encrypted or damaged PDFs

A password-protected document may require authentication before pages or objects can be read. A damaged or malformed file can fail during opening or while reading a particular page. Preserve the original file, report the page or xref that failed, and decide whether your application should stop or continue with other pages.

Build a production extraction pipeline

  1. Validate the input: check file type, size limits and whether the PDF can be opened.
  2. Choose an output contract: define whether consumers receive one image per page occurrence, one per xref, or structured elements.
  3. Use deterministic names: include page number, block number or xref and preserve the reported extension.
  4. Record metadata: store page, xref, width, height, extension and byte count alongside each file.
  5. Make jobs restartable: write to a temporary location, then atomically move completed files into the final directory.
  6. Validate outputs: open the written bytes with an image decoder and compare the number of extracted records with your expected count.
  7. Retain provenance: keep the source document hash and extraction version so results can be reproduced.

Performance, reliability and cost

  • Hosted jobs: account for upload time, asynchronous processing and result download. Poll with backoff and stop after a deadline; use a webhook when your deployment can receive one reliably.
  • Local processing: avoids upload and network latency but uses your own CPU, memory and disk. Process large documents page by page and avoid holding every image in memory.
  • Duplicate work: xref deduplication can reduce output size when logos or repeated figures appear throughout a document.
  • Retries: retry transient network failures and status checks, but do not blindly resubmit a completed extraction job.
  • Security: treat PDFs and extracted images as untrusted input. Store credentials outside source code and apply access controls to both originals and outputs.
  • Cost: a hosted provider charges according to its current transaction terms. Local extraction has no API transaction fee but still consumes compute, storage and operational time.

Troubleshooting

Symptom Likely cause Fix
No images are returned The PDF contains vectors or text rather than embedded raster images. Render pages when you need a visual result, or inspect the document structure before choosing extraction.
The same image appears many times One xref is referenced from multiple pages. Deduplicate by xref.
Transparent artwork has a solid background A stencil mask was not combined with its base image. Extract and apply the mask with an image-processing step.
Every file is saved as PNG The output filename ignored the reported extension. Use the extension returned by the API or PyMuPDF.
Hosted job never finishes Polling stopped too early, the operation failed, or the network request timed out. Inspect the operation status and error payload, use bounded exponential backoff, and support the documented webhook option.
Unauthorized API response Expired token, wrong credential, or a secret sent to the wrong endpoint. Refresh credentials on the server and check the authorization header and environment configuration.
Extraction works locally but fails in production Missing native dependency, file permission, memory limit or different library version. Pin the dependency, log the library version, check writable storage and test with representative PDFs.

Or skip the browser setup

ScreenshotNeo is useful when your actual requirement is a clean image of a PDF or document page rendered at a URL, rather than the original embedded image objects. One GET request returns PNG, JPEG, WebP or PDF output. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/report.pdf -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com/report.pdf"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/report.pdf' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', buffer);

Before capture, cookie banners, newsletter popups and chat widgets are removed. Bot checks, blank pages and failed loads are never billed, and response headers identify the page verdict and billing result. ScreenshotNeo also provides an MCP server for AI agents, including Claude and Cursor. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Does PDF extraction return the original image quality?

PyMuPDF’s object extraction returns the embedded bytes where available. A hosted service returns the formats and structures documented by that service. Rendering a page creates a new image and can change resolution.

Should I extract by page block or xref?

Use page blocks when page placement matters. Use xrefs when you want each underlying image object once and can handle repeated references yourself.

Can an API extract images from scanned PDFs?

A scan may contain one page-sized raster image per page. If the content is only pixels, extraction returns that page image; it cannot recover separate original photos that were flattened into the scan.

When should I choose Markdown output?

Choose PDF-to-Markdown when the next consumer expects Markdown and can decode embedded base64 figures. Choose structured JSON when you need standalone files and element metadata.

Is ScreenshotNeo an embedded-image extraction API?

No. ScreenshotNeo captures a rendered URL. Use PyMuPDF or a PDF extraction service for embedded object bytes; use ScreenshotNeo when a clean rendered page image is the desired output.