ScreenshotNeo

BlogHTML to image & PDF

Export Specific Pages from a Generated PDF in Java

Copy selected pages from a generated PDF into a new file with PDFBox or iText, including ranges, non-contiguous pages, validation, and common fixes.

By the ScreenshotNeo team1 October 20267 min read

To export specific pages from a generated PDF in Java, create a destination PDF and copy only the requested pages into it. For one contiguous range, Apache PDFBox’s PageExtractor is direct; iText 7 uses copyPagesTo. For pages such as 1, 3, and 7, use a page-copy workflow or iText 5’s documented page-selection API.

Choose the extraction method

Need Recommended API Selection model
One inclusive range PDFBox PageExtractor One-based start and end pages
One inclusive range in an iText project iText 7 copyPagesTo One-based start and end pages
Non-contiguous pages iText 5 selectPages, or a page-copy loop List or range expression

Keep the same PDF library when possible. It avoids conversion steps and keeps behavior consistent with the code that generated the source document.

Apache PDFBox: extract a contiguous range

PDFBox’s PageExtractor takes a source PDDocument, a start page, and an end page, then returns a new document. Both endpoints are inclusive. Values below 1 are clamped to page 1; an end page beyond the source runs through the final page. An invalid range can produce a blank document. See the PDFBox PageExtractor API.

import java.io.IOException;
import java.nio.file.Path;
import org.apache.pdfbox.Loader;
import org.apache.pdfbox.multipdf.PageExtractor;
import org.apache.pdfbox.pdmodel.PDDocument;

public final class ExtractRange {
    public static void main(String[] args) throws IOException {
        Path input = Path.of("generated.pdf");
        Path output = Path.of("pages-3-to-6.pdf");
        int startPage = 3;
        int endPage = 6;

        if (startPage < 1 || endPage < startPage) {
            throw new IllegalArgumentException("Use a one-based range with start <= end");
        }

        // PDFBox 3.x loading API. Use the loader appropriate for your PDFBox version.
        try (PDDocument source = Loader.loadPDF(input.toFile())) {
            if (startPage > source.getNumberOfPages()) {
                throw new IllegalArgumentException("Start page is outside the source PDF");
            }

            PageExtractor extractor = new PageExtractor(source, startPage, endPage);
            try (PDDocument selected = extractor.extract()) {
                selected.save(output.toFile());
            }
        }
    }
}

PDFBox’s command-line documentation describes the same one-based, inclusive behavior for startPage and endPage. If your project uses PDFBox 2.x, adapt the loading call to that version’s API while keeping the extraction semantics the same.

Validate the range yourself

Although PageExtractor defines clamping behavior, application code should validate user input before extraction. Decide whether an out-of-range end page should mean “through the end” or should be rejected, then test that policy explicitly. Reject an empty or reversed range rather than silently writing a blank PDF.

iText 7: copy an inclusive page range

In an iText 7 project, open the generated PDF with a PdfReader, create a destination with a PdfWriter, and call copyPagesTo(pageFrom, pageTo, destination). The destination must be closed so the writer can finish the file. Check the iText version and its licensing terms before shipping.

import com.itextpdf.kernel.pdf.PdfDocument;
import com.itextpdf.kernel.pdf.PdfReader;
import com.itextpdf.kernel.pdf.PdfWriter;

public final class ITextRange {
    public static void main(String[] args) throws Exception {
        String input = "generated.pdf";
        String output = "pages-3-to-6.pdf";
        int pageFrom = 3;
        int pageTo = 6;

        try (PdfDocument source = new PdfDocument(new PdfReader(input));
             PdfDocument destination = new PdfDocument(new PdfWriter(output))) {
            if (pageFrom < 1 || pageTo < pageFrom || pageFrom > source.getNumberOfPages()) {
                throw new IllegalArgumentException("Invalid page range");
            }
            pageTo = Math.min(pageTo, source.getNumberOfPages());
            source.copyPagesTo(pageFrom, pageTo, destination);
        }
    }
}

Copy non-contiguous pages such as 1, 3, and 7

A contiguous-range helper cannot express a set with gaps. Build a destination document and copy each requested page in order with the page-copy API available in your chosen library. Validate every page number and decide whether output order should follow the request.

For iText 5, PdfReader.selectPages accepts a comma-separated range expression or a List<Integer>. The selected pages are retained, may be reordered, and cannot be repeated.

import com.itextpdf.text.Document;
import com.itextpdf.text.pdf.PdfCopy;
import com.itextpdf.text.pdf.PdfReader;
import com.itextpdf.text.pdf.PdfImportedPage;
import java.io.FileOutputStream;

public final class IText5Pages {
    public static void main(String[] args) throws Exception {
        PdfReader reader = new PdfReader("generated.pdf");
        Document output = new Document();
        PdfCopy copy = new PdfCopy(output, new FileOutputStream("pages-1-3-7.pdf"));
        output.open();

        int[] requested = {1, 3, 7};
        for (int page : requested) {
            if (page < 1 || page > reader.getNumberOfPages()) {
                throw new IllegalArgumentException("Page outside source PDF: " + page);
            }
            PdfImportedPage imported = copy.getImportedPage(reader, page);
            copy.addPage(imported);
        }

        output.close();
        reader.close();
    }
}

Extract pages immediately after generating the PDF

Finish and serialize the generated document before importing pages whenever possible. PDFBox documents warn that importing a page from an unfinished generated document can encounter incomplete structures, including font-subsetting information. A safe sequence is:

  1. Finish writing the generator’s document.
  2. Close it or save it to a completed file.
  3. Reopen that file for extraction.
  4. Select and save the requested pages.
  5. Close both source and destination documents with try-with-resources.

Also inspect annotations that point to pages outside the selected set. Those references can make the destination larger than expected. Verify forms, outlines, metadata, encryption, annotations, embedded files, and external references if they matter to your application; page-copy support differs by library and workflow.

Preserve output fidelity

  • Annotations: confirm that links, comments, and page destinations still resolve in the new document.
  • Forms: test AcroForm fields, especially when several selected pages share field names.
  • Outlines: bookmarks may refer to pages that were not copied and may need rebuilding.
  • Metadata: copy or set title, author, subject, and custom properties deliberately.
  • Fonts and resources: reopen a completed generated file to avoid unfinished subset data.
  • Encryption: supply the correct password and check whether the destination should be encrypted again.

Troubleshooting

Symptom Likely cause Fix
Output is blank Start is after end, or the selected range is invalid. Validate one-based values and reject reversed ranges.
Only one page appears The same page was copied repeatedly, or the loop exits early. Log each requested page and add each imported page once.
FileNotFoundException The generator has not saved the source path or the process lacks access. Close/save the generator first and verify the absolute path and permissions.
Missing fonts or malformed resources Pages were imported from an unfinished generated document. Close the generator, reopen the serialized PDF, then extract.
Output unexpectedly grows Annotations or shared resources reference pages and objects outside the selection. Inspect annotations and remove or rebuild references that are not needed.
Bookmarks point nowhere Outlines target pages that were not selected. Rebuild outlines for the destination or omit invalid entries.
Encrypted source cannot open A password is required. Open with the library’s password-aware reader and follow your security policy.
iText code fails at compile time Examples target a different major version. Pin the dependency and use that version’s API; do not mix iText 5 and 7 classes.

Performance, reliability, and cost considerations

  • Extraction still reads the source PDF and may copy shared resources; memory and time depend on document structure, page content, and annotations.
  • For large files, process one extraction job at a time or apply a queue and memory limit rather than loading many documents concurrently.
  • Write to a temporary destination and atomically move it into place after a successful close, so readers never see a partially written PDF.
  • Use deterministic page-range validation and record the source identifier, requested pages, library version, and output checksum for reproducibility.
  • Benchmark your own representative PDFs if latency or memory limits matter; the research does not provide a general benchmark.
  • PDFBox is Apache-licensed. iText licensing depends on the distribution and version; review the applicable terms before deployment.

Or skip the browser setup

If your pipeline first needs screenshots of web pages before assembling or reviewing a PDF, ScreenshotNeo provides a single HTTP request for a clean PNG, JPEG, WebP, or PDF. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result.

It also offers an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000 shots. See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Create a free ScreenshotNeo account with 1,000 screenshots each month and no card.

FAQ

Are PDF page numbers zero-based?

No. The APIs described here use one-based page numbers. Page 1 is the first page.

Can I preserve the original page order?

Yes. Copy pages in ascending order for the original order, or in the requested order when producing a rearranged document.

Should I use PDFBox or iText?

Use the library already used to generate the source when it supports the required selection. Choose based on contiguous versus non-contiguous selection, required fidelity, version compatibility, and licensing.

Can I extract pages without writing an intermediate source file?

You can use an in-memory document when the generator has fully completed it, but reopening a completed serialized PDF is safer for generated fonts and other unfinished structures.