How to Extract Data from PDFs with an API
Choose the right PDF extraction workflow for digital text, scans, tables, and structured output—with implementation examples and validation guidance.
PDF extraction starts with one question: does the file contain selectable text, or are its pages images? Digital PDFs can usually be parsed directly. Scanned PDFs require optical character recognition (OCR) before the text can be used. After that decision, choose the output your application needs: plain text, Markdown, structured JSON, tables, figures, or reading-order-aware blocks.
For content and layout extraction, Adobe PDF Extract documents text blocks, structure, reading order, tables, figures, and styling in JSON. Adobe also documents PDF-to-Markdown output. For image-based documents, Adobe OCR converts image text into searchable content, while Amazon Textract provides document text detection and analysis. Validate each option against representative files before committing to a provider.
1. Identify the PDF type
Open several representative files and try to select and copy text.
- Selectable text: use a text or structure extraction operation.
- Image-only pages: run OCR first, then parse the recognized text.
- Mixed PDFs: process each page according to its content; a single file may contain both digital and scanned pages.
Also record whether the files contain multiple columns, rotated pages, handwriting, tables, footnotes, figures, headers, or repeated page furniture. These characteristics affect reading order and field accuracy.
2. Define the output contract
| Application need | Useful output | What to verify |
|---|---|---|
| Search, indexing, summarization | Plain text or Markdown | Reading order, headings, page boundaries |
| Downstream data processing | Structured JSON | Block relationships, coordinates, confidence, missing content |
| Invoices, reports, schedules | Table extraction | Rows, columns, merged cells, numeric formatting |
| Forms and key-value fields | Provider-specific analysis features | Field names, values, checkboxes, confidence |
| Figures and diagrams | Figure metadata plus source regions | Placement, captions, extraction boundaries |
Do not convert structured output to plain text too early. Keep the original response and page references so a reviewer can trace every value back to its source page.
3. Choose an API workflow
Adobe PDF Extract
Adobe describes PDF Extract as a cloud service for extracting content and structural information from native or scanned PDFs. Its documented outputs include structured JSON, text, layout and reading order, table cells, figures, and styling. Adobe also documents a PDF-to-Markdown option for documentation and LLM workflows. The service provides SDKs for Node.js, Python, .NET, and Java, plus REST access. See the PDF Extract API overview and output details.
Adobe OCR
Use Adobe’s OCR operation when pages are images and you need searchable text before further processing. OCR quality depends on scan resolution, contrast, skew, language, and whether the document contains handwriting. Read the OCR PDF documentation.
Amazon Textract
Textract is an AWS API for document text detection and analysis. It is a practical choice when your application already uses AWS credentials, S3, asynchronous jobs, or AWS-native monitoring. Consult the Textract API reference and current pricing page for operation-specific limits and prices.
4. A repeatable extraction pipeline
- Store the original PDF with an immutable identifier.
- Inspect a sample of pages for selectable text and document complexity.
- Choose text extraction, OCR, table analysis, or a combination.
- Upload the document using the provider’s documented SDK or REST flow.
- Poll or await the operation result when the API is asynchronous.
- Persist the raw response, provider version, request options, and timestamp.
- Normalize the result into your application schema.
- Validate text, reading order, tables, figures, and page references against the source.
- Route low-confidence or structurally unusual documents for review.
5. Runnable AWS Textract examples
The following examples use Textract’s synchronous text detection operation for a local, image-based document. It accepts PNG or JPEG bytes; for multi-page PDFs, use the documented asynchronous S3 workflow instead.
Python
import boto3
from pathlib import Path
textract = boto3.client("textract", region_name="us-east-1")
image_bytes = Path("page.png").read_bytes()
response = textract.detect_document_text(Document={"Bytes": image_bytes})
lines = [
block["Text"]
for block in response["Blocks"]
if block["BlockType"] == "LINE"
]
print("\n".join(lines))
Node.js
import { TextractClient, DetectDocumentTextCommand } from "@aws-sdk/client-textract";
import { readFile } from "node:fs/promises";
const client = new TextractClient({ region: "us-east-1" });
const bytes = await readFile("page.png");
const result = await client.send(new DetectDocumentTextCommand({
Document: { Bytes: bytes }
}));
const lines = (result.Blocks ?? [])
.filter(block => block.BlockType === "LINE")
.map(block => block.Text);
console.log(lines.join("\n"));
cURL and REST integration
Most cloud PDF APIs require authentication headers, upload or asset-creation calls, and sometimes a second result request. Use the provider’s official REST documentation to generate the required authorization rather than copying unsigned examples. A generic integration shape is:
curl -X POST "$PDF_API_ENDPOINT" \
-H "Authorization: Bearer $PDF_API_TOKEN" \
-H "Content-Type: application/pdf" \
--data-binary @document.pdf
Replace the endpoint and authentication scheme with the exact operation documented by your provider. Do not send production documents to an endpoint until you have confirmed retention, region, encryption, and deletion behavior.
6. Handling digital PDFs
Digital PDFs often contain a text layer, but that layer may be fragmented, out of order, or positioned for visual layout rather than reading. Prefer a structure-aware response when you need columns, headings, tables, or page coordinates. Preserve page numbers and bounding boxes where available.
For an LLM or documentation system, Markdown can be a useful compact representation. For analytics or database ingestion, retain structured JSON and map blocks into an explicit schema such as:
{
"document_id": "invoice-042.pdf",
"pages": [
{
"page": 1,
"blocks": [
{"type": "heading", "text": "Invoice", "order": 0},
{"type": "table", "rows": [["Item", "Amount"], ["Service", "100.00"]]}
]
}
]
}
7. Handling scanned PDFs with OCR
OCR recognizes pixels; it does not recover information that is absent or unreadable in the scan. Before processing, deskew pages, improve contrast when appropriate, and avoid excessive compression. Test printed text, stamps, handwriting, rotated pages, and low-resolution pages separately.
Keep OCR confidence and page coordinates if the provider returns them. A low-confidence value should trigger review rather than silently entering a financial or legal workflow.
8. Tables, forms, and reading order
Tables
Validate every row and column against the rendered page. Common failures include merged headers, wrapped labels, totals placed outside the detected grid, and numbers split across lines. Store the original cell text before numeric conversion.
Forms
Map extracted fields to a schema with required, optional, and unknown states. An empty value can mean an unchecked box, a missing field, or a recognition failure; those cases should not be conflated.
Multi-column pages
Check that the result reads down the first column before moving to the next. Footnotes, sidebars, and headers often require explicit ordering rules.
9. Validation checklist
- Compare extracted headings and paragraphs with the source page.
- Check column order on two-column and three-column pages.
- Compare table row counts, headers, totals, and decimal separators.
- Verify page numbers, footnotes, captions, and repeated headers.
- Inspect rotated pages and pages containing figures.
- Measure OCR results separately for clean scans, skewed scans, and handwriting.
- Save failed examples and add them to regression tests.
Provider feature descriptions are not an accuracy benchmark. The research for this guide did not establish an independent head-to-head accuracy or throughput winner, so validate on your own document mix.
10. Errors and troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| No text returned | The PDF is image-only or the text layer is malformed. | Run OCR or render pages to images and process them. |
| Text is in the wrong order | Multi-column layout, positioned glyphs, or sidebars. | Use structure-aware output and validate block coordinates. |
| Tables lose columns | Merged cells, borders, or whitespace-based layout. | Use a table-capable operation and compare cell coordinates with the source. |
| Async job never completes | Incorrect polling, missing permissions, or an invalid object location. | Check operation status, IAM permissions, region, and provider limits; apply bounded retries. |
| Access denied | Expired credentials or insufficient API permissions. | Refresh credentials and grant only the documented operation and storage permissions. |
| Throttling or rate-limit errors | Concurrency exceeds the account or regional limit. | Use exponential backoff, a work queue, and a bounded concurrency limit. |
| Unexpected cost | Page rounding, analysis features, or retries increased transactions. | Count pages and selected features before production; inspect provider usage records. |
11. Performance, reliability, and cost
Performance
Separate upload, processing, and result retrieval timings. Parallelize independent documents within provider limits, but avoid unbounded concurrency. Cache results by a hash of the original bytes and extraction options.
Reliability
Make jobs idempotent. Persist the source identifier and operation ID, retry transient failures with exponential backoff, and send permanently failed documents to a review queue. Keep raw provider responses so you can reprocess after changing your normalization code.
Cost
Estimate cost from actual page volume, operation type, region, and selected analysis features. Adobe states that Extract PDF and PDF to Markdown page counts are rounded up in five-page increments for transaction calculations; confirm current terms in the Adobe licensing documentation. AWS publishes feature-based Textract pricing. Adobe’s overview currently lists 500 free Document Transactions per month, but offers and limits can change, so confirm before purchase.
12. Or skip the browser setup
If your workflow begins with a web page and you need a visual record before extracting or reviewing its PDF output, ScreenshotNeo provides a single-call capture API. It removes cookie banners, newsletter popups, and chat widgets before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server lets Claude, Cursor, and other MCP clients take screenshots. The free plan includes 1,000 screenshots each month without a card; paid plans start at $5 for 3,000 shots.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for capture options. Start with 1,000 free screenshots per month.
FAQ
Can an API extract text from every PDF?
No. Scans, handwriting, damaged text layers, unusual fonts, and complex layouts require OCR or specialized handling, and results must be validated.
Should I request Markdown or JSON?
Use Markdown for compact reading and LLM or documentation workflows. Use JSON when you need layout, relationships, tables, coordinates, or downstream field mapping.
Is OCR always needed for a scanned PDF?
Yes, if you need machine-readable text from image-only pages. A PDF viewer’s visual display does not imply that a text layer exists.
How do I compare providers fairly?
Create a representative corpus, run the same extraction goals, measure field and table correctness, record latency and failures, and calculate cost using current provider rules.
Can I extract tables and figures together?
Choose an operation that documents both capabilities, preserve the structured response, and verify table cells and figure boundaries against source pages.


