ScreenshotNeo

BlogHTML to image & PDF

How to Export Specific Pages from a PDF in Python

Select, reorder, validate, and save PDF pages in Python with PyMuPDF or pypdf, plus fixes for indexing, links, and output errors.

By the ScreenshotNeo team1 October 20268 min read

To export specific pages from a PDF in Python, open the source document, select the zero-based page indexes you need, and write a new PDF. PyMuPDF offers the shortest approach with Document.select(). The pypdf reader/writer pattern is useful when you want to add pages one by one, build output from several files, or make the destination document explicit.

This guide shows both methods, converts human page numbers safely, supports ranges, reordering, and duplicates, and explains validation, links, annotations, bookmarks, troubleshooting, performance, and cost. It uses the current APIs documented by PyMuPDF and pypdf.

Choose the right Python approach

Use case Recommended API
Select pages from one document with minimal code PyMuPDF doc.select(indexes)
Construct an output by adding pages explicitly pypdf PdfReader + PdfWriter.add_page()
Reorder or repeat pages Either; pass the desired index sequence
Combine pages from multiple PDFs pypdf writer pattern

Export pages with PyMuPDF

Install the package (the import name is pymupdf):

python -m pip install PyMuPDF

The indexes are zero-based: index 0 is the first physical page. The official tutorial describes select() as shrinking a PDF to selected pages. The sequence determines output order and may contain repeated indexes.

import pymupdf

source_path = "input.pdf"
output_path = "selected-pages.pdf"
indexes = [0, 2, 4]  # first, third, and fifth pages

doc = pymupdf.open(source_path)
try:
    if not indexes:
        raise ValueError("Select at least one page")
    page_count = doc.page_count
    invalid = [i for i in indexes if i < 0 or i >= page_count]
    if invalid:
        raise ValueError(f"Invalid page indexes {invalid}; document has {page_count} pages")

    doc.select(indexes)
    doc.save(output_path)
finally:
    doc.close()

# Optional verification
check = pymupdf.open(output_path)
try:
    if check.page_count != len(indexes):
        raise RuntimeError("Output page count does not match the selection")
finally:
    check.close()

For the first and second pages only, use:

import pymupdf

doc = pymupdf.open("input.pdf")
doc.select([0, 1])
doc.save("selected-pages.pdf")
doc.close()

That is the concise form shown in the PyMuPDF basics documentation. The checked version is safer for scripts that receive page numbers from users or another system.

Accept human page numbers

People normally say “pages 1, 3, and 5,” while Python libraries expect 0, 2, 4. Convert explicitly and validate before changing the document:

import pymupdf


def export_human_pages(source, destination, requested_pages):
    """requested_pages uses one-based page numbers, such as [1, 3, 5]."""
    doc = pymupdf.open(source)
    try:
        total = doc.page_count
        if not requested_pages:
            raise ValueError("At least one page is required")
        if any(not isinstance(n, int) for n in requested_pages):
            raise TypeError("Page numbers must be integers")
        invalid = [n for n in requested_pages if n < 1 or n > total]
        if invalid:
            raise ValueError(f"Pages {invalid} are outside 1..{total}")

        indexes = [n - 1 for n in requested_pages]
        doc.select(indexes)
        doc.save(destination)
    finally:
        doc.close()


export_human_pages("input.pdf", "pages-1-3-5.pdf", [1, 3, 5])

Ranges, order, and repeated pages

import pymupdf

doc = pymupdf.open("input.pdf")
try:
    # Pages 2 through 5, then page 1, then page 5 again (one-based input).
    requested = [2, 3, 4, 5, 1, 5]
    indexes = [page - 1 for page in requested]
    if any(i < 0 or i >= doc.page_count for i in indexes):
        raise ValueError("A requested page is outside the document")
    doc.select(indexes)
    doc.save("reordered-and-repeated.pdf")
finally:
    doc.close()

Do not confuse printed labels with physical indexes. A PDF can display a page label such as “iii” or “1” while the library still addresses physical pages from zero. If labels matter, inspect the document’s page-label configuration and define the mapping your application expects.

Export pages with pypdf

Install pypdf:

python -m pip install pypdf

Use PdfReader.pages[index] and add each chosen page to a PdfWriter:

from pypdf import PdfReader, PdfWriter

reader = PdfReader("input.pdf")
writer = PdfWriter()
indexes = [0, 2, 4]

if not indexes:
    raise ValueError("Select at least one page")
if any(i < 0 or i >= len(reader.pages) for i in indexes):
    raise ValueError(f"Valid indexes are 0 through {len(reader.pages) - 1}")

for index in indexes:
    writer.add_page(reader.pages[index])

with open("selected-pages.pdf", "wb") as output:
    writer.write(output)

The pypdf documentation defines page access as zero-based and documents add_page(). Its merging guide also shows selecting indexes while appending pages; range-related conveniences can vary by installed version, so check the versioned API before using them.

Copy a contiguous range

from pypdf import PdfReader, PdfWriter

reader = PdfReader("input.pdf")
writer = PdfWriter()
start, stop = 1, 5  # zero-based, stop is exclusive: pages 2-5

if not (0 <= start < stop <= len(reader.pages)):
    raise ValueError("Invalid range")

for index in range(start, stop):
    writer.add_page(reader.pages[index])

with open("range.pdf", "wb") as output:
    writer.write(output)

Combine selected pages from multiple PDFs

from pypdf import PdfReader, PdfWriter

jobs = [
    ("chapter-a.pdf", [0, 2]),
    ("chapter-b.pdf", [1, 3]),
]
writer = PdfWriter()

for path, indexes in jobs:
    reader = PdfReader(path)
    for index in indexes:
        if index < 0 or index >= len(reader.pages):
            raise ValueError(f"Invalid index {index} for {path}")
        writer.add_page(reader.pages[index])

with open("combined-selection.pdf", "wb") as output:
    writer.write(output)

Validation checklist before and after saving

  1. Confirm the source path exists and is readable.
  2. Open the PDF and read its page count.
  3. Convert one-based input to zero-based indexes exactly once.
  4. Reject an empty selection when your workflow requires output pages.
  5. Reject every index outside 0 <= index < page_count.
  6. Write to a distinct destination so the source is not overwritten accidentally.
  7. Reopen the result and verify its page count.
  8. Open the result in a PDF viewer and inspect important links, annotations, and bookmarks.

PyMuPDF documents that selected-page output retains links, annotations, and bookmarks that remain valid when they point to a selected page or an external resource. Removing destination pages can still make internal references meaningless, so inspect navigation in outputs that depend on cross-page links.

Common errors and fixes

Error or symptom Cause Fix
ValueError from select() An index is negative, outside the page count, or the sequence is invalid. Check doc.page_count and validate every index before selection.
Wrong pages exported One-based human numbers were passed directly to a zero-based API. Use [n - 1 for n in requested_pages] once.
Output has zero pages The selection list was empty. Reject empty input or define an explicit empty-document policy.
FileNotFoundError Relative path resolves from an unexpected working directory. Use an absolute path or log Path.cwd(); check existence first.
ModuleNotFoundError The package is not installed in the active virtual environment. Run python -m pip install PyMuPDF or python -m pip install pypdf with that interpreter.
Encrypted PDF cannot be read The source requires a password. Use the library’s password/decryption support only when you are authorized to access the file, then verify the output.
Bookmarks or links point nowhere Their destination page was omitted. Inspect navigation and rebuild or remove references as your application requires.
Output is unexpectedly large Pages may contain large images or embedded resources. Measure output size, avoid unnecessary duplicate pages, and apply an appropriate PDF optimization workflow after selection.
Corrupt or unusual source PDF The file may be truncated, malformed, or only partially supported. Open it in a desktop viewer, obtain a complete copy, and try the other library to identify whether the issue is parser-specific.

Performance, reliability, and cost notes

  • Memory: both workflows must parse the source and construct an output document. For very large PDFs, process jobs individually, close documents promptly, and monitor process memory.
  • Speed: the research sources do not establish a universal performance winner. Choose the API that matches your document workflow and measure with your own files.
  • Reliability: validate indexes before mutating the document, save to a new path, reopen the result, and keep the original until verification succeeds.
  • Repeatability: pin library versions in your project and review the installed documentation, especially for version-specific pypdf append/range examples.
  • Cost: PyMuPDF and pypdf perform local processing; your direct cost is normally compute and storage. A hosted capture service is useful when the input is a web page rather than an existing PDF.

Or skip the browser setup

If your workflow starts with a URL and you need a PDF or screenshot before selecting pages, ScreenshotNeo provides a single API endpoint and an MCP server for AI agents. It captures a page as PNG, JPEG, WebP, or PDF; PDF options include paper size, margins, landscape mode, and page ranges. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

See the ScreenshotNeo API documentation for all options. A direct PDF request looks like this:

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -d format=pdf \
  -o page.pdf
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={
        "access_key": "YOUR_API_KEY",
        "url": "https://stripe.com",
        "format": "pdf",
    },
    timeout=90,
)
r.raise_for_status()
open("page.pdf", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com',
  format: 'pdf'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('page.pdf', buffer);

ScreenshotNeo also supports custom CSS and JavaScript, waiting for selectors or network idle, element capture, full-page capture with lazy images loaded, headers, cookies, user agents, authorization, timezone, geolocation, caching with a chosen TTL, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, signed links, and usage reporting. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. There are 1,000 free screenshots each month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

FAQ

Can I export pages without rendering them?

Yes. PyMuPDF and pypdf copy selected PDF pages; they do not need a browser. Use ScreenshotNeo when you first need to turn a web page URL into a PDF.

Can I preserve the original page order?

Yes. Pass indexes in ascending order. To reorder or repeat pages, pass the desired sequence explicitly.

Why do examples use zero-based indexes?

Both documented Python APIs address the first physical page as index zero. Convert user-facing page numbers before calling the library.

Which library should a new project choose?

Use PyMuPDF for concise in-place selection on one document; use pypdf when explicit page-by-page assembly or multi-file composition fits your design. The sources do not support a universal quality or speed winner.

Should I overwrite the input file?

No. Save to a separate path, verify the result, and replace the original only as a deliberate second step.