How to Extract Text and Data From PDFs in n8n
Use n8n’s Extract From File node to read PDFs, handle binary input and OCR, then map the result into reliable structured data.
Use the Extract From File node with the Extract From PDF operation. Your upstream node must provide the PDF as binary data, usually in a property named data. After extraction, add a separate cleanup, parsing or AI step when you need fields such as an invoice number, date, vendor or total.
What you need
- An n8n workflow.
- A PDF-producing node, such as HTTP Request, Webhook, a local file source or a storage integration.
- The PDF available as binary data.
- The current Extract From File node. Older tutorials may call this node Read PDF; Extract From File replaced it from n8n 1.21.0 onward.
Build the basic PDF text workflow
- Add the node that obtains the PDF.
- Run that node once and inspect its output. Confirm that the Binary panel contains a PDF property. The default property name is commonly
data. - Add an Extract From File node and connect the file-producing node to it.
- Set the operation to Extract From PDF.
- Set the input binary field to the actual property name from the previous node. Leave it as
dataonly when that is the name you see. - Execute the node and inspect the JSON output.
- Send the extracted result to a Code node, database, spreadsheet, notification or AI model according to your use case.
The extraction node converts the binary document into workflow data. It does not know your business schema. Treat “get text out” and “turn text into validated fields” as two different stages.
Getting a PDF into n8n
HTTP Request node
Use an HTTP Request node when the PDF is available at a URL.
- Set the method to
GET. - Enter the PDF URL.
- Configure the response format as a file, or enable the option that downloads the response as binary data in your n8n version.
- Check the binary property name. Rename it or reference it in Extract From File if it is not
data.
Webhook uploads
For an uploaded PDF, connect a Webhook node to Extract From File. Enable the Webhook node’s Raw body option as required by the official extraction workflow. Then execute the Webhook once with a real upload and verify that the expected binary property exists before configuring the extraction node.
Storage integrations and local files
Google Drive, S3-compatible storage and other integrations can download a document into binary data. A local-file workflow can do the same on a self-hosted instance. The key requirement is unchanged: the connected node must emit the PDF as binary data, and the field name in Extract From File must match it.
Clean and format extracted text
PDF text often contains line breaks, repeated spaces, headers and footer fragments. Use a Code node for deterministic cleanup before sending the result elsewhere.
// Code node: normalize common whitespace noise
const input = $json.text ?? $json.data ?? $json.content ?? '';
const text = String(input)
.replace(/\\r\\n/g, '\\n')
.replace(/[ \\t]+/g, ' ')
.replace(/\\n{3,}/g, '\\n\\n')
.trim();
return [{ json: { text } }];
Inspect the actual output property in your n8n version before using this example. If the extraction node returns a differently named field, replace the fallback expression with that field.
Extract named fields from a PDF
For structured information, add a parsing stage after text extraction. A reliable workflow is:
- Extract the PDF text.
- Normalize whitespace and remove known boilerplate.
- Ask an AI or parser step for a fixed JSON schema.
- Validate required fields and data types.
- Route invalid or incomplete records for review.
Example schema for an invoice:
{
"invoice_number": "string|null",
"invoice_date": "YYYY-MM-DD|null",
"vendor": "string|null",
"currency": "string|null",
"subtotal": "number|null",
"tax": "number|null",
"total": "number|null"
}
Tell the parser to return only this schema, use null when a value is absent, and never infer a value that is not present. Validate the result in a Code node before writing it to your accounting or CRM system. Extraction and AI parsing can be incomplete, especially when the source layout is inconsistent, so retain the original text or file for review.
Scanned PDFs and OCR
A scanned PDF may contain page images rather than selectable characters. In that case, ordinary text extraction can return little or no useful text. The public n8n invoice workflow example recommends enabling OCR for scanned PDFs in the Extract From File node’s options. Option names can vary by n8n version, so confirm the setting in the node used by your deployment.
OCR quality depends on scan resolution, skew, handwriting, multi-column layouts and table structure. For sensitive records, validate totals and identifiers against the source document instead of accepting OCR output without checks.
Complete example: webhook to normalized invoice JSON
- Create a Webhook node that accepts a file upload.
- Enable Raw body.
- Connect it to Extract From File configured for Extract From PDF.
- Enable OCR when the incoming documents are scans.
- Connect a Code node that normalizes the extracted text.
- Connect your parser or AI node with the invoice schema.
- Add a validation Code node and route failures to a review branch.
// Code node: basic validation after your parser returns JSON
const item = $json;
const required = ['invoice_number', 'vendor', 'total'];
const missing = required.filter((key) => item[key] === null || item[key] === undefined || item[key] === '');
const totalIsNumber = item.total === null || typeof item.total === 'number';
return [{
json: {
...item,
valid: missing.length === 0 && totalIsNumber,
missing_fields: missing
}
}];
Calling a PDF source from code
These examples download a PDF. Configure the receiving n8n Webhook or storage step to preserve the response as binary data.
cURL
curl -L "https://example.com/invoice.pdf" -o invoice.pdf
Python
import requests
r = requests.get("https://example.com/invoice.pdf", timeout=90)
r.raise_for_status()
with open("invoice.pdf", "wb") as f:
f.write(r.content)
Node.js
const res = await fetch('https://example.com/invoice.pdf');
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('invoice.pdf', buffer));
Binary data checks
| Check | What it tells you | Fix |
|---|---|---|
| No Binary panel | The upstream node produced text or JSON, not a file. | Configure it to download or emit binary data. |
| Binary field mismatch | Extract From File is looking at the wrong property. | Set the input field to the upstream property name. |
| Empty extraction | The PDF may be scanned, encrypted or malformed. | Try OCR, verify the file opens, and check for a password requirement. |
| Webhook body is missing | The upload was not passed as the expected raw binary body. | Enable Raw body and send a new test upload. |
Troubleshooting
“Cannot find binary property”
Cause: the configured field, often data, does not match the actual binary property. Run the previous node, open Binary, and copy the exact property name into Extract From File.
The node returns no text
Cause: the PDF is image-only, protected, damaged or contains an encoding the extractor cannot interpret. Confirm that the file opens, enable OCR for scans, and test with a known text-based PDF.
The webhook receives JSON metadata but no file
Cause: the webhook configuration or client request did not transmit the file as raw body data. Enable Raw body, resend the upload, and inspect the binary output before continuing.
Fields are shifted or incorrect
Cause: extraction preserves text but not necessarily the visual meaning of columns, tables or nearby labels. Use explicit parsing rules, a fixed schema and validation. Keep the source document available for manual review.
An older tutorial mentions Read PDF
Use Extract From File and choose Extract From PDF in current n8n versions. The older Read PDF name refers to the retired integration.
Performance, reliability and cost considerations
- Large PDFs take longer to download, process and send to an AI model. Limit unnecessary pages before downstream processing when your source system supports it.
- OCR generally requires more processing than extracting an existing text layer. Use it only when scans require it.
- Keep binary-data storage in mind on self-hosted n8n. Storage configuration affects scaling and has security implications because document files may contain sensitive information.
- Design retries around the download and extraction boundaries. Avoid creating duplicate records by using an idempotency key such as a source file ID or hash.
- Log the source identifier, extraction status and validation result, while limiting sensitive document content in logs.
- Do not treat parser or OCR output as guaranteed correct. Route missing totals, conflicting values and low-confidence cases to review.
Or skip the browser setup
If the document or source page is online and you need a clean visual capture for an audit trail, handoff or downstream workflow, ScreenshotNeo provides a single screenshot request. It is separate from PDF text extraction, but can remove the browser automation work around capturing a source page.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Before capture, cookie banners, newsletter popups and chat widgets are removed. Bot checks, blank pages and failed loads are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. ScreenshotNeo also offers an MCP server so Claude, Cursor and other MCP clients can take screenshots. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. See the ScreenshotNeo documentation for options and request details, then create a free account.
FAQ
Which n8n node reads a PDF?
Use Extract From File with the Extract From PDF operation.
Does Extract From File create invoice fields automatically?
No. It extracts document content. Add a parsing, mapping or AI step and validate the resulting schema.
Do all PDFs need OCR?
No. PDFs with a usable text layer can be extracted directly. Scanned, image-only PDFs may need OCR.
Why does my workflow work with one PDF but fail with another?
PDFs differ in text layers, encryption, layout, tables, encoding and scan quality. Test representative files and keep a review path for exceptions.
Where should binary files be stored in self-hosted n8n?
Choose storage according to your deployment, scale and security requirements, and review n8n’s binary-data configuration for your version.


