How to Export Selected Pages from a PDF in Node.js
Extract pages in any order into a new PDF with pdf-lib, validate input safely, and learn when qpdf is a better fit.
Use pdf-lib when you want a pure-JavaScript Node.js solution. Load the source PDF, convert the requested one-based page numbers to zero-based indices, call copyPages, append the returned pages in the requested order, and save the destination document.
The example below exports pages 1, 3, and 5 to selected-pages.pdf:
import { readFile, writeFile } from 'node:fs/promises'
import { PDFDocument } from 'pdf-lib'
const input = await readFile('input.pdf')
const source = await PDFDocument.load(input)
const output = await PDFDocument.create()
// pdf-lib uses zero-based indices: pages 1, 3, and 5 are 0, 2, and 4.
const selected = await output.copyPages(source, [0, 2, 4])
for (const page of selected) output.addPage(page)
const bytes = await output.save()
await writeFile('selected-pages.pdf', bytes)
The PDFDocument API documents copyPages(srcDoc, indices); the project documentation supports Node.js and other JavaScript runtimes.
Install pdf-lib
mkdir pdf-page-export
cd pdf-page-export
npm init -y
npm install pdf-lib
# package.json: add "type": "module" if you use the import syntax above
Put input.pdf beside your script, then run it with a current Node.js release:
node export-pages.js
Export pages chosen at runtime
Applications usually receive page numbers from a request, form, or queue. Validate the one-based values before converting them. This version accepts a comma-separated list such as 1,3,5, rejects invalid values, and preserves the requested order.
import { readFile, writeFile } from 'node:fs/promises'
import { PDFDocument } from 'pdf-lib'
function parsePageNumbers(value, pageCount) {
if (!value || typeof value !== 'string') {
throw new Error('pages must be a comma-separated list such as 1,3,5')
}
const oneBased = value.split(',').map((part) => {
const trimmed = part.trim()
if (!/^\d+$/.test(trimmed)) {
throw new Error(`Invalid page number: ${part}`)
}
const number = Number(trimmed)
if (!Number.isSafeInteger(number) || number < 1 || number > pageCount) {
throw new Error(`Page ${number} is outside 1-${pageCount}`)
}
return number
})
return oneBased.map((number) => number - 1)
}
const source = await PDFDocument.load(await readFile('input.pdf'))
const indices = parsePageNumbers(process.argv[2] ?? '1,3,5', source.getPageCount())
const output = await PDFDocument.create()
const pages = await output.copyPages(source, indices)
for (const page of pages) output.addPage(page)
await writeFile('selected-pages.pdf', await output.save())
console.log(`Wrote ${pages.length} page(s) to selected-pages.pdf`)
Run it with:
node export-pages.js 1,3,5
Keep a custom order or repeat a page
copyPages returns pages in the same order as the index array. For pages 5, 2, and 5:
const selected = await output.copyPages(source, [4, 1, 4])
for (const page of selected) output.addPage(page)
Duplicate indices intentionally repeat a page. Decide whether your product should allow that; reject duplicates if each source page must appear only once.
Export a contiguous range
Convert a one-based inclusive range to indices with a small helper:
function rangeIndices(firstPage, lastPage, pageCount) {
if (!Number.isInteger(firstPage) || !Number.isInteger(lastPage)) {
throw new Error('Range endpoints must be integers')
}
if (firstPage < 1 || lastPage < firstPage || lastPage > pageCount) {
throw new Error(`Range must be within 1-${pageCount}`)
}
return Array.from(
{ length: lastPage - firstPage + 1 },
(_, offset) => firstPage - 1 + offset
)
}
const indices = rangeIndices(4, 7, source.getPageCount()) // [3, 4, 5, 6]
Insert pages at a specific position
Appending is simplest, but the returned PDFPage objects can also be positioned with insertPage. When you need a fixed layout, construct the destination in the required sequence and insert each page at its destination index.
const pages = await output.copyPages(source, [4, 1, 3])
for (let index = 0; index < pages.length; index++) {
output.insertPage(index, pages[index])
}
What gets preserved, and what to verify
Page content is copied into a different PDFDocument. Treat document-level behavior as something to verify for your files. Test forms, annotations, outlines, metadata, encryption, embedded files, and unusual fonts before promising exact fidelity. The API documentation does not promise that every document-level feature transfers identically.
- Forms: check field names, values, and appearance streams.
- Annotations: open links and interactive annotations in a viewer.
- Outlines: confirm bookmarks point to valid destination pages.
- Metadata: set destination metadata explicitly if required.
- Encrypted PDFs: supply the supported load options and verify the document can be opened before copying.
Native alternative: qpdf
qpdf’s --pages option uses CLI page-range syntax, supports reverse ordering, and can select pages from multiple input files. For a simple extraction:
qpdf input.pdf --pages . 1,3,5 -- selected-pages.pdf
qpdf is useful when your server image already includes native PDF tooling or when you need its established page-range and multi-file syntax. It requires an installed executable, process startup, and careful argument validation.
Invoke qpdf safely from Node.js
import { execFile } from 'node:child_process'
import { promisify } from 'node:util'
const execFileAsync = promisify(execFile)
// Keep arguments as an array. Do not build a shell command by concatenating user input.
await execFileAsync('qpdf', [
'input.pdf',
'--pages', '.', '1,3,5', '--',
'selected-pages.pdf'
])
Use an allowlist for input and output paths, impose process timeouts, and capture stderr for diagnostics. qpdf’s document-level behavior differs between normal mode and --empty; choose deliberately when metadata from the primary input matters.
Choosing between pdf-lib and qpdf
| Concern | pdf-lib | qpdf |
|---|---|---|
| Deployment | Pure JavaScript dependency | Native executable required |
| Selection | Zero-based JavaScript index array | CLI ranges such as 1,3,5 |
| Multiple files | Load documents and copy pages in code | Explicit cross-file page selection syntax |
| Execution | In process | Child process with executable discovery |
| Fidelity | Verify forms, annotations, outlines, metadata, and encryption | Verify the same requirements and qpdf mode |
Performance and reliability checklist
- Read the source once and call
source.getPageCount()before validating requests. - Reject impossible or excessively large selections before allocating the output document.
- For concurrent jobs, cap parallel PDF loads and saves to avoid memory pressure.
- Write to a temporary file, then rename it after
save()completes so readers never see a partial output. - Keep the original page order in request logs so an output can be reproduced.
- Use deterministic filenames or job IDs and clean temporary files after success or failure.
- Measure memory with representative PDFs; image-heavy pages can make output and heap usage much larger than page count suggests.
Troubleshooting
“Cannot read properties…” or an invalid PDF error
The input may be HTML, truncated bytes, or an encrypted document that needs a password. Confirm the download status and content type, preserve the original bytes, and load the file before attempting page selection.
“Page is outside…”
The request used one-based numbers while the API expects zero-based indices, or the number exceeds source.getPageCount(). Validate one-based input, then subtract one exactly once.
The pages appear in the wrong order
The index array controls output order. Pass indices in the desired order and append the returned pages sequentially.
The output opens but interactive features changed
Page copying does not guarantee identical document-level features. Test forms, annotations, outlines, metadata, and encryption with the actual PDFs you support; use qpdf or a specialized PDF workflow when its fidelity is required.
The Node process runs out of memory
Large, scanned, or image-heavy PDFs can consume substantial memory during load and save. Limit concurrent jobs, process files in a worker, enforce upload limits, and benchmark with production-shaped documents.
PDFKit is for creating PDFs
PDFKit’s getting-started guide focuses on generating a new document and piping it to a writable stream. It is appropriate when you are drawing a new PDF, but it is not the default choice for copying selected pages from an existing PDF.
Or skip the browser setup
If your workflow starts with web pages that must become PDFs or images, ScreenshotNeo provides a single request to capture a URL. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. An MCP server lets Claude, Cursor, and other MCP clients call screenshot and PDF tools.
See the ScreenshotNeo API documentation for all options, including PDF paper size, margins, landscape mode, and page ranges.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
There is a free plan with 1,000 screenshots per month and no card. Paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Can I export pages in reverse order?
Yes. Pass the indices in reverse order, such as [4, 3, 2, 1, 0], or use qpdf’s reverse page-range syntax.
Can I combine pages from two PDFs?
Yes. Load both source documents, copy the required pages into the same destination document, and append each returned page in the required sequence. qpdf also documents cross-file selection.
Does save() return a file path?
No. It returns PDF bytes. Write those bytes to a file, object storage, or an HTTP response according to your application.
Should page numbers in an API be zero-based?
User interfaces are usually clearer with one-based page numbers. Convert and validate at the boundary, then use zero-based indices internally.


