How to Parse PDFs in Node.js with pdf-parse
Use the current pdf-parse v2 class API to extract PDF text in Node.js, handle passwords and errors, clean up resources, and avoid v1 examples.
Use the current pdf-parse v2 class API: install the package, create a PDFParse instance with a PDF URL, call getText(), read the returned text field, and always call destroy() in a finally block.
The most common mistake is copying an older v1 example such as pdf(buffer).then(...) and combining it with the v2 class API. Choose one major-version API and keep its loading and result handling consistent.
Install pdf-parse
npm install pdf-parse
The npm listing identified version 2.4.5 as the latest tag when this guide was researched. Check the current npm package page before pinning a version, because releases and tags can change.
Parse a PDF URL with the current v2 API
This complete CommonJS example follows the current project README’s URL-based flow. It prints the extracted text and releases parser resources whether parsing succeeds or fails.
const { PDFParse } = require('pdf-parse');
async function main() {
const parser = new PDFParse({
url: 'https://bitcoin.org/bitcoin.pdf'
});
try {
const result = await parser.getText();
console.log(result.text);
} catch (error) {
console.error('PDF parsing failed:', error);
process.exitCode = 1;
} finally {
await parser.destroy();
}
}
main();
Save it as parse-pdf.cjs and run:
node parse-pdf.cjs
The documented result exposes extracted content through result.text. A PDF with unusual fonts, scanned pages, columns, or complex layout may produce text that needs additional cleanup; the package documentation describes capabilities, not perfect extraction for every file.
Use ESM instead of CommonJS
If your project uses ESM, use the named import shown by the current README:
import { PDFParse } from 'pdf-parse';
async function main() {
const parser = new PDFParse({
url: 'https://bitcoin.org/bitcoin.pdf'
});
try {
const result = await parser.getText();
console.log(result.text);
} finally {
await parser.destroy();
}
}
main().catch((error) => {
console.error('PDF parsing failed:', error);
process.exitCode = 1;
});
Use a package configuration that enables ESM, or save the file with an .mjs extension.
How the v2 API works
| Step | Code | Purpose |
|---|---|---|
| Construct | new PDFParse({ url }) |
Creates a parser for the PDF input. |
| Extract | await parser.getText() |
Reads text from the document. |
| Consume | result.text |
Returns the extracted text string. |
| Release | await parser.destroy() |
Frees parser resources after success or failure. |
The project describes additional operations for document information, header validation, page screenshots, embedded image extraction, and table extraction. Treat those as separate documented capabilities and check the README for the exact method names and options in the version you installed.
Passwords and protected PDFs
The current API documents a password load parameter. Supply it when the PDF requires a password, then handle password-specific failures separately from network or malformed-file errors.
const { PDFParse } = require('pdf-parse');
async function parseProtectedPdf() {
const parser = new PDFParse({
url: 'https://example.com/protected.pdf',
password: process.env.PDF_PASSWORD
});
try {
const result = await parser.getText();
return result.text;
} finally {
await parser.destroy();
}
}
parseProtectedPdf()
.then((text) => console.log(text))
.catch((error) => {
console.error('Could not parse protected PDF:', error);
process.exitCode = 1;
});
Do not hard-code passwords in source control. Read them from a secret manager or environment variable. If the password is wrong, changing parser options will not make the document readable; obtain the correct password or an unrestricted copy.
v1 versus v2: do not mix the examples
Older v1 documentation used a function-style call resembling:
pdf(buffer).then((result) => {
console.log(result.text);
});
The current README presents PDFParse as the v2 API. The constructor, loading options, method calls, result shape, and cleanup requirements can differ across major versions. If an old snippet fails, first check the installed version and then use the matching README rather than patching v1 code into a v2 example.
Node.js versions and project setup
The project documentation lists Node.js 20 (at least 20.16.0), 22 (at least 22.3.0), 23 (at least 23.0.0), and 24 (at least 24.0.0) as supported at research time. It lists 19 and earlier and 21 as unsupported. Verify the README before publishing or upgrading because runtime support is a project fact that can change.
node --version
npm install pdf-parse
Pin the dependency in a lockfile for repeatable deployments, and recheck the package documentation when changing major versions.
Downloading a PDF before parsing
For a URL input, let pdf-parse load the document as shown above. If your application must download files itself, use a separate HTTP client and then follow the installed version’s documented local or buffer loading syntax. The exact local-file and Buffer form is version-sensitive; do not assume a v1 Buffer example remains valid unchanged in v2.
For a quick connectivity check, cURL can download the source PDF:
curl -L 'https://bitcoin.org/bitcoin.pdf' -o bitcoin.pdf
That command only downloads the file. Parsing still happens in Node.js with the API documented for your installed pdf-parse version.
What extraction can and cannot guarantee
- Text PDFs: selectable text is generally the intended input for text extraction.
- Scanned PDFs: image-only pages require OCR; a PDF text parser cannot create text that is not present as a text layer.
- Columns and reading order: extracted strings may not follow the visual order a person sees.
- Tables: the project documents table extraction, but inspect the output for merged cells, spacing, and page breaks.
- Fonts and encodings: unusual encodings can yield missing or garbled characters.
- Images: image extraction is a separate documented capability from text extraction.
Keep the original PDF when the extracted text is used for indexing, compliance, or audit work so a reviewer can compare the source with the parsed output.
Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
PDFParse is not a constructor |
v1 and v2 code are mixed, or the import does not match the module format. | Check the installed version, use the v2 named export, and choose CommonJS or ESM consistently. |
| Parser resolves but text is empty | The document may be scanned, have an unusual text layer, or contain unsupported encoding. | Inspect the PDF visually, try a text-based copy, or add OCR for image-only pages. |
| Password exception | The file is encrypted or the supplied password is wrong. | Pass the documented password parameter and verify the secret. |
| Invalid PDF error | The URL returned HTML, a partial download, or a corrupt file. | Open the URL with curl -L, check the response and file, and confirm it is a real PDF. |
| Response or network error | The server is unavailable, redirects unexpectedly, or blocks the request. | Check the URL, redirects, network access, and server response; retry with an application-level timeout policy. |
| Memory grows during batch jobs | Parser instances are not destroyed. | Put every parse in try/finally and call await parser.destroy(). |
| Works locally but not in production | Unsupported Node.js version, dependency drift, or restricted outbound access. | Use a documented supported Node line, lock dependencies, and verify network policy. |
Performance, reliability, and cost considerations
Performance
- PDF size, page count, embedded images, and font complexity affect memory and elapsed time.
- For batches, limit concurrency so several large documents do not exhaust process memory.
- Release every parser promptly with
destroy(); do not retain full result objects longer than needed. - Store normalized text only when appropriate. Keep page boundaries or metadata if downstream search depends on them.
Reliability
- Validate that a remote response is actually a PDF before parsing when you control the download step.
- Handle password, invalid-file, and response exceptions separately so operators know whether to fix credentials, input, or connectivity.
- Record the source URL, parser version, Node.js version, and failure category for reproducibility.
- Retry transient network failures with bounded backoff, but do not repeatedly retry a wrong password or corrupt file.
Cost
pdf-parse is an npm package, so the parsing library itself does not introduce a per-document API charge. Your infrastructure can still incur storage, bandwidth, compute, and OCR costs. Large PDFs and high concurrency increase memory and CPU requirements.
Or skip the browser setup
If your workflow also needs clean screenshots or PDF capture of web pages, ScreenshotNeo provides a single HTTP endpoint and an MCP server for AI agents. It is separate from text extraction with pdf-parse, but useful when the deliverable is a visual capture rather than parsed text.
See the ScreenshotNeo API documentation for all options. A basic request is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/document.pdf -o shot.webp
Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Which API should a new project use?
Use the current v2 PDFParse class API documented by the installed release. Avoid copying v1 function-style examples into a v2 project.
Does pdf-parse perform OCR?
The documented package capabilities cover PDF parsing and extraction. An image-only scanned PDF needs an OCR step to create a text layer.
Why is destroy() necessary?
It releases parser resources. Calling it in finally prevents cleanup from being skipped when parsing throws.
Can I trust extracted text to preserve page layout?
No parser should be assumed to preserve every visual layout. Columns, tables, fonts, and reading order require validation against representative PDFs.


