ScreenshotNeo

BlogHTML to image & PDF

How to Extract Structured Text from PDFs as JSON with an API

Learn how PDF extraction APIs return text, layout, tables, and reading order as JSON, with Adobe PDF Extract and Amazon Textract examples.

By the ScreenshotNeo team1 October 20269 min read

PDF extraction APIs can return more than a string of characters. Depending on the operation and provider, the response may include paragraphs, headings, lists, tables, forms, page numbers, reading order, coordinates, and relationships between detected elements. The correct workflow is to choose an operation for your document type, submit the PDF, map the provider’s schema into your own JSON model, and validate the result against the source pages.

A plain text response is sufficient when you only need searchable words. Use structured extraction when downstream code must preserve document meaning or location, such as headings, table cells, citations, page references, or reading order.

What structured PDF extraction returns

There is no universal PDF-to-JSON schema. Each service names and groups elements differently.

Need Typical structured data Validation concern
Searchable text Pages, lines, words, or text spans Reading order and repeated headers
Document hierarchy Headings, paragraphs, lists, footnotes, styles Provider-specific element labels
Tables Cells, rows, columns, spans, or table relationships Merged cells and visual boundaries
Forms Field names, values, keys, and geometry Unsupported form types and permissions
Citations or review Page numbers, bounding boxes, confidence, element IDs Coordinates must be checked against the page image

Scanned PDFs need text recognition. Results depend on scan quality, language, skew, contrast, layout, and file constraints. Always retain the source PDF and validate representative pages, especially multi-column layouts and complex tables.

A provider-neutral extraction workflow

  1. Classify the input. Identify native-text PDFs, image-only scans, forms, and table-heavy files. A native PDF may need layout extraction; a scan needs OCR; a form or table may require a separate analysis feature.
  2. Select the operation. Choose basic text detection for pages, lines, and words, or a layout/table/form analysis operation when structure is required.
  3. Upload or reference the PDF. Follow the provider’s documented upload, storage, authentication, and asynchronous-job process.
  4. Parse into an application-owned schema. Preserve page numbers, element types, order, geometry, and provider IDs when later review or citations depend on them.
  5. Validate against the source. Compare headings, columns, table boundaries, repeated headers, footers, and low-quality scans with the rendered page.
  6. Handle failures explicitly. Detect encrypted, password-protected, corrupt, unsupported, oversized, or overly complex documents and decide whether to reject, split, or route them for review.

Adobe PDF Extract API: structured JSON and renditions

Adobe documents a cloud PDF Extract API for native and scanned PDFs. Its JSON endpoint is designed for structured downstream processing and includes reading order and page layout. The documentation describes text grouped into paragraphs, headings, lists, and footnotes with styling information. Tables can include cell content and formatting, with optional CSV/XLSX output and PNG renditions; identified figures or images can also be returned as PNG files. Adobe provides Node.js, Python, .NET, and Java SDKs. See the official overview and the Extract API guide.

Adobe’s documented flow is: create an asset from the source PDF, configure extraction parameters, run the extract operation, then retrieve the JSON structure and optional renditions. The guide summarizes the purpose precisely: “The sample below extracts text element information from a PDF document and returns a JSON file.”

Mapping Adobe output into your schema

Do not make your application depend directly on Adobe’s response shape. Store the original response, then map it into a stable model such as:

{
  "document_id": "invoice-2026-001",
  "pages": [
    {
      "number": 1,
      "elements": [
        {
          "type": "heading",
          "text": "Invoice",
          "bbox": [72, 720, 180, 750],
          "order": 0
        },
        {
          "type": "table",
          "cells": [],
          "order": 1
        }
      ]
    }
  ]
}

Keep provider IDs and raw geometry when possible. A later reviewer can then locate a value on the original page even if your normalized schema changes.

Amazon Textract: blocks, tables, forms, and layout

Amazon Textract exposes more than one analysis path. DetectDocumentText returns JSON Block objects organized around pages, lines, and words. It has synchronous and asynchronous modes. AWS documents a 10 MB maximum for synchronous documents and a 500 MB maximum for asynchronous PDF files.

AnalyzeDocument accepts PDF input and supports feature selection such as TABLES, FORMS, QUERIES, SIGNATURES, and LAYOUT. Lines and words are included in the response. Textract Blocks are a provider schema, not automatically your business schema; use relationships and geometry to build your own representation.

Minimal Textract request examples

The following AWS CLI commands require configured AWS credentials and a supported document location. They show the operation shape; choose synchronous or asynchronous processing according to the file size and job duration.

aws textract detect-document-text \
  --document '{"S3Object":{"Bucket":"YOUR_BUCKET","Name":"documents/report.pdf"}}' \
  --output json > textract-text.json
aws textract analyze-document \
  --document '{"S3Object":{"Bucket":"YOUR_BUCKET","Name":"documents/report.pdf"}}' \
  --feature-types TABLES FORMS LAYOUT \
  --output json > textract-analysis.json

For long-running jobs, use Textract’s asynchronous APIs and poll or receive the completion notification according to AWS documentation. Do not assume that a synchronous request can handle every PDF.

Normalize provider JSON with Python

This standalone script converts a Textract response into page-indexed words and lines. It is deliberately small: add table, form, confidence, and geometry mappings that your application actually needs.

import json
from pathlib import Path

source = json.loads(Path("textract-text.json").read_text())
result = {"pages": []}
by_page = {}
for block in source.get("Blocks", []):
    page = block.get("Page", 1)
    by_page.setdefault(page, []).append(block)

for page_number in sorted(by_page):
    elements = []
    for block in by_page[page_number]:
        kind = block.get("BlockType")
        if kind not in {"LINE", "WORD"}:
            continue
        geometry = block.get("Geometry", {}).get("BoundingBox", {})
        elements.append({
            "type": kind.lower(),
            "text": block.get("Text", ""),
            "confidence": block.get("Confidence"),
            "bbox": geometry,
            "id": block.get("Id")
        })
    result["pages"].append({"number": page_number, "elements": elements})

Path("document.json").write_text(json.dumps(result, indent=2))
print("wrote document.json")

Normalize provider JSON with Node.js

import { readFile, writeFile } from 'node:fs/promises';

const source = JSON.parse(await readFile('textract-text.json', 'utf8'));
const pages = new Map();
for (const block of source.Blocks ?? []) {
  const page = block.Page ?? 1;
  if (!pages.has(page)) pages.set(page, []);
  if (block.BlockType !== 'LINE' && block.BlockType !== 'WORD') continue;
  const box = block.Geometry?.BoundingBox ?? {};
  pages.get(page).push({
    type: block.BlockType.toLowerCase(),
    text: block.Text ?? '',
    confidence: block.Confidence ?? null,
    bbox: box,
    id: block.Id ?? null
  });
}
const document = {
  pages: [...pages.entries()].sort(([a], [b]) => a - b)
    .map(([number, elements]) => ({ number, elements }))
};
await writeFile('document.json', JSON.stringify(document, null, 2));
console.log('wrote document.json');

Calling an extraction endpoint with cURL and Python

Every provider uses different authentication, upload fields, and operation names. Use the exact endpoint and headers from the provider’s current documentation. Keep secrets out of source control.

curl -X POST 'https://YOUR_PROVIDER.example/v1/extract' \
  -H 'Authorization: Bearer YOUR_API_TOKEN' \
  -F 'file=@report.pdf' \
  -F 'output=json' \
  -o extraction.json
import requests

with open('report.pdf', 'rb') as pdf:
    response = requests.post(
        'https://YOUR_PROVIDER.example/v1/extract',
        headers={'Authorization': 'Bearer YOUR_API_TOKEN'},
        files={'file': ('report.pdf', pdf, 'application/pdf')},
        data={'output': 'json'},
        timeout=120,
    )
response.raise_for_status()
with open('extraction.json', 'wb') as output:
    output.write(response.content)

Replace the placeholder URL and fields with the selected provider’s documented request. The normalization and validation stages remain the same.

Tables, reading order, and coordinates

Tables need their own validation

Text detection can find words without understanding that they belong to the same row or column. Check whether the API returns cells, rows, columns, spans, or relationships. Validate merged cells, wrapped text, nested headers, totals, and tables that continue across pages. If table fidelity is essential, preserve the page image or a provider rendition for review.

Reading order is a data decision

Two-column pages, sidebars, footnotes, captions, and repeated headers can produce a sequence that is technically valid but wrong for your downstream use. Keep element order and coordinates, then add rules for your document family. Never silently concatenate all page text before checking columns.

Coordinates require a consistent origin

Providers may express bounding boxes as normalized fractions or page units, and coordinate origins can differ. Store the provider’s raw geometry and page dimensions. Convert only at the application boundary that draws highlights or citations.

Errors and troubleshooting

Symptom Likely cause Fix
Authentication or authorization error Expired token, wrong region, missing permission, or malformed signature Check credentials, region, scopes, and the provider’s required signing method.
Password-protected or encrypted PDF fails The service cannot open the document Obtain an authorized decrypted copy or use the provider’s documented password support.
Unsupported language or poor OCR Language is not supported or scan quality is low Check supported languages, improve scan resolution and contrast, and validate against the page image.
Timeout or job never completes Large, complex, illustration-heavy, or table-heavy input Use asynchronous processing, split the PDF into smaller files, and retry with bounded backoff.
Missing table cells Table analysis was not enabled or layout is ambiguous Select the provider’s table feature and review merged, rotated, or borderless tables manually.
Wrong column order Reading-order inference differs from your document Use coordinates and element order to implement document-specific column grouping.
File-size or page-limit error Provider limits were exceeded Check current limits and split or compress the input where permitted.
Blank or incomplete output Corrupt PDF, vector-art-heavy page, or unsupported form technology Open the source with a PDF validator, render representative pages, and route unsupported files for a different parser or review.

Adobe specifically lists unsupported languages, XFA forms, restricted permissions, password-protected or corrupted PDFs, oversized files, page-limit violations, complex inputs or tables, and processing timeouts as failure conditions. Its guide notes that splitting a file can address a timeout and cautions that documents dominated by illustrations, CAD drawings, or other vector art may not produce quality results.

Performance, reliability, and cost planning

  • Choose synchronous versus asynchronous deliberately. Synchronous calls simplify small files. Queue large or variable jobs and persist job IDs so workers can resume after a process restart.
  • Make retries safe. Use an idempotency key when the provider supports one, otherwise record a document hash and extraction status before retrying.
  • Control concurrency. Match worker count to provider quotas, storage bandwidth, and downstream parsing capacity. Back off on rate-limit responses.
  • Cache by content hash. Avoid paying twice for an unchanged PDF. Include extraction options and provider version in the cache key.
  • Measure useful output. Track pages processed, failed documents, OCR confidence where available, table review rate, and cost per accepted document. Do not treat character count as extraction quality.
  • Review current pricing and quotas. The reviewed sources establish Adobe’s listed offer of 500 free Document Transactions per month, marked on its overview as updated May 1, 2026. Recheck current terms before relying on it. The dossier does not establish a comparable price analysis for Textract.

Validation checklist before processing a corpus

  • Include native text, image-only scans, multi-column pages, tables with merged cells, forms, footnotes, and repeated headers.
  • Check language, encryption, permissions, page count, file size, and provider limits.
  • Compare extracted text and order with rendered source pages.
  • Verify table rows, columns, spans, and numeric values.
  • Store page numbers and geometry for citations or review.
  • Define a fallback path for unsupported or low-confidence documents.
  • Version your normalized schema and retain the raw provider response.

Or skip the browser setup

ScreenshotNeo is a website screenshot API rather than a PDF-to-JSON extraction service. It can still help when your workflow needs a rendered visual reference of a public document or web page for validation. One GET request returns a PNG, JPEG, WebP, or PDF. Before capture, it accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response reports the result in X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. See the ScreenshotNeo documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests; r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90); open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

There are 1,000 free screenshots each month with no card. Paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.

FAQ

Does JSON guarantee correct document structure?

No. JSON is only the container. Confirm what the provider means by a heading, table, relationship, confidence value, and reading order, then validate representative pages.

Should I use OCR for every PDF?

No. Native-text PDFs may work better with text and layout extraction. Use OCR for image-only scans and test language and scan quality first.

Can I use one normalized schema for Adobe and Textract?

Yes, but through a mapping layer. Keep each raw response because the providers expose different element names, relationships, and geometry conventions.

When should extraction be asynchronous?

Use asynchronous processing when files are large, page counts vary, or the provider documents job-based processing. Persist job state and make retries resumable.

How do I prove an extracted value came from the PDF?

Store the page number, element identifier, geometry, and source-file hash, then let reviewers open the corresponding page image or rendition.