ScreenshotNeo

BlogHTML to image & PDF

How to Extract Fields from PDFs with the Stirling PDF API

Learn how to identify interactive, text-based, and scanned PDFs, verify Stirling-PDF's version-specific API, and extract form data safely.

By the ScreenshotNeo team1 October 20268 min read

Short answer: first determine what kind of PDF you have. Stirling-PDF’s documented form operations are intended for existing interactive fields such as text boxes, checkboxes, radio buttons, and combo boxes. The exact operation path, upload field, parameters, and response format depend on your installed version, so open /swagger-ui/index.html on the Stirling-PDF server you will call and use that schema as the contract.

A selectable-text PDF and a scanned PDF require different workflows. Text conversion can give you words to parse, while OCR can turn page images into machine-readable text. Neither process automatically means that arbitrary business fields can be extracted with the same form endpoint.

1. Identify the PDF representation

Input What is stored Likely workflow Output caveat
Interactive form AcroForm or similar controls with names and values Use the form-field extraction operation exposed by your Stirling-PDF version Best match for “extract fields”; verify whether the version exports CSV or XLSX
Selectable text Characters positioned on pages Use the version’s text or PDF-to-CSV/XML operation, then parse the result in your application Visual labels are not necessarily named fields
Scanned pages Page images Run OCR first, then map the recognized text into your own field schema OCR output is text; it does not guarantee structured form values

Quick checks

  • Can you tab between controls and type into the document? It probably contains interactive fields.
  • Can you select and copy words, but not focus controls? Treat it as a text-extraction problem.
  • Is every page an image? Plan for OCR and validation.
  • Does the PDF contain a signature or flattened form? A visually filled form may no longer contain editable field objects.

2. Use the exact Swagger contract for your server

Stirling-PDF generates API documentation from endpoint annotations and publishes it through OpenAPI. The local Swagger UI is therefore the safest source for a deployed instance because paths and schemas can change between releases or configuration profiles.

  1. Open https://YOUR-STIRLING-HOST/swagger-ui/index.html in a browser. Use your actual scheme and host.
  2. Search for operations mentioning form fields, extract form data, or export form data.
  3. Read the operation’s request schema. Record the HTTP method, multipart field name for the PDF, optional parameters, accepted content types, and response content type.
  4. Use the Swagger “Try it out” request with a non-sensitive sample PDF.
  5. Save the generated request example or OpenAPI operation ID in your integration tests.

Do not copy an endpoint path from a different installation. The available project material describes form extraction and CSV/XLSX export, but it does not establish one stable path or payload contract for every release.

Inspect the OpenAPI document programmatically

Swagger UI normally loads an OpenAPI document for the same server. The URL can be deployment-specific, so inspect the page or its network requests to find the JSON document URL. The following examples fetch a URL you supply and list operations whose path or description mentions form fields.

curl -fsSL "$STIRLING_OPENAPI_URL" \
  -H "Accept: application/json" \
  -H "X-API-KEY: $STIRLING_API_KEY" \
  -o stirling-openapi.json

python - <<'PY'
import json, os
spec = json.load(open("stirling-openapi.json", encoding="utf-8"))
for path, methods in spec.get("paths", {}).items():
    for method, operation in methods.items():
        if method.lower() not in {"get", "post", "put", "patch", "delete"}:
            continue
        text = " ".join([
            path,
            operation.get("summary", ""),
            operation.get("description", ""),
            operation.get("operationId", "")
        ]).lower()
        if "form" in text or "field" in text:
            print(method.upper(), path, operation.get("operationId", ""))
PY

The X-API-KEY header is the authentication mechanism described by the project README, subject to your deployment’s security settings. If your instance uses a different security configuration, follow the security scheme shown in its OpenAPI document.

3. Extract existing interactive form values

Once you have identified the operation in local Swagger, send the PDF using the exact request schema shown there. Because the project does not publish one universal path and parameter set for all versions, keep the operation-specific request in your code as configuration rather than assuming a hard-coded endpoint.

Configuration-driven Python client

import os
from pathlib import Path
import requests

# Copy these values from your instance's Swagger operation.
endpoint = os.environ["STIRLING_FORM_ENDPOINT"]
pdf_path = Path(os.environ.get("PDF_PATH", "input.pdf"))
file_field = os.environ.get("STIRLING_FILE_FIELD", "file")
api_key = os.environ["STIRLING_API_KEY"]

with pdf_path.open("rb") as pdf:
    response = requests.post(
        endpoint,
        headers={"X-API-KEY": api_key, "Accept": "application/json"},
        files={file_field: (pdf_path.name, pdf, "application/pdf")},
        timeout=120,
    )
response.raise_for_status()
content_type = response.headers.get("content-type", "")
if "json" in content_type:
    print(response.json())
else:
    Path("form-output.bin").write_bytes(response.content)
    print(f"Wrote form-output.bin ({content_type})")

Set STIRLING_FORM_ENDPOINT and STIRLING_FILE_FIELD to the values displayed by your Swagger schema. If the operation requires additional form parts, add only the names and values shown there.

Configuration-driven Node.js client

import { readFile } from 'node:fs/promises';

const endpoint = process.env.STIRLING_FORM_ENDPOINT;
const fileField = process.env.STIRLING_FILE_FIELD || 'file';
const apiKey = process.env.STIRLING_API_KEY;
const pdfPath = process.env.PDF_PATH || 'input.pdf';

if (!endpoint || !apiKey) throw new Error('Set STIRLING_FORM_ENDPOINT and STIRLING_API_KEY');
const pdf = await readFile(pdfPath);
const form = new FormData();
form.append(fileField, new Blob([pdf], { type: 'application/pdf' }), pdfPath);

const res = await fetch(endpoint, {
  method: 'POST',
  headers: { 'X-API-KEY': apiKey, 'Accept': 'application/json' },
  body: form
});
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
const type = res.headers.get('content-type') || '';
if (type.includes('json')) console.log(await res.json());
else await Bun.write('form-output.bin', Buffer.from(await res.arrayBuffer()));

On Node versions without Bun.write, replace the final line with writeFile from node:fs/promises. Keep the multipart field name and any extra parameters aligned with Swagger.

CSV and XLSX exports

Project discussion material describes exporting form data to CSV and XLSX. Treat that as a version-dependent capability: confirm the response type and export option in your server’s Swagger schema. Do not confuse this with PDF-to-CSV or PDF-to-XML conversion, which may represent page content rather than named interactive controls.

4. Handle selectable text and scanned PDFs

Selectable text

Use the text or conversion operation documented by your version, save the response, and parse it with a format-aware library. Expect layout problems: columns may be interleaved, labels may be separated from values, and repeated headers may appear on every page. Define a mapping layer that normalizes whitespace, dates, numbers, and page breaks before writing records.

Scanned pages and OCR

Stirling-PDF lists OCR support and associates it with Tesseract. OCR is a preparation step: it recognizes characters in page images, but it does not by itself identify your application’s “invoice number” or “policy holder” field. Follow the OCR operation in your local Swagger, then validate the recognized text and apply your own field-mapping rules.

  1. Detect whether the source is image-only.
  2. Run the documented OCR operation with the language and output options supported by your installation.
  3. Store the OCR result alongside the original PDF for auditability.
  4. Extract values with deterministic patterns or a separate document parser.
  5. Require validation for totals, identifiers, dates, and other high-impact values.

5. Reliability, performance, and cost considerations

  • Version pinning: pin the Stirling-PDF image or release and keep an exported OpenAPI document with your integration tests.
  • Timeouts: OCR and large PDFs can take longer than simple form reads. Set a client timeout appropriate for your page count and retry only requests that are safe to repeat.
  • File limits: check reverse-proxy and Stirling-PDF upload limits before production. Reject unexpectedly large files early.
  • Idempotency: hash the input file and operation settings so retries do not create duplicate records.
  • Validation: verify response content type, required fields, page count, and file identifiers before accepting output.
  • Privacy: self-hosting keeps processing under your deployment’s control; still apply access controls, TLS, key rotation, and log redaction.
  • Cost: Stirling-PDF is self-hosted software, so budget for the server, storage, OCR CPU, and operational maintenance rather than assuming a per-request vendor price.

6. Troubleshooting

Symptom Likely cause Fix
404 from an endpoint copied online Path differs in your release Use the local Swagger UI and copy the operation from that instance.
415 Unsupported Media Type Wrong upload encoding or content type Follow the operation’s request schema; use multipart only when Swagger specifies it.
401 or 403 Missing or invalid API key, or security settings differ Send X-API-KEY as configured and inspect the documented security scheme.
Empty field list PDF is flattened, text-only, or scanned Inspect the PDF representation and switch to text conversion or OCR as appropriate.
Garbled OCR Low resolution, skew, complex layout, or wrong language Improve the source scan, configure supported languages, and validate results.
CSV contains page text instead of fields Used a conversion operation rather than form export Confirm that the selected operation reads interactive controls and check its response schema.
Request hangs Large file, OCR workload, proxy timeout, or resource pressure Check server logs and proxy limits, increase timeouts, and process files asynchronously at your application layer.

7. Or skip the browser setup

If your application needs screenshots of the source document, a rendered preview, or a related web page, ScreenshotNeo provides a single HTTP request instead of maintaining browser automation. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result.

It also provides an MCP server for Claude, Cursor, and other MCP clients, with take_screenshot, get_page_info, and capture_pdf tools. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for the capture options. Sign up free to get 1,000 screenshots each month with no card.

8. FAQ

Can Stirling-PDF extract fields from any PDF?

No. Interactive controls are the clearest supported case. Scans and ordinary text require OCR or text processing plus your own mapping and validation.

Why should I use local Swagger instead of a copied example?

The running instance documents its own operation paths, request fields, response types, and security configuration. Those details can vary by version.

Does OCR recover the original form field names?

No. OCR recognizes visible characters. It does not recreate AcroForm metadata or guarantee semantic field names.

Can I treat CSV export as proof that values were extracted from controls?

Only after confirming the operation in your version. PDF-to-CSV conversion and form-data export are different capabilities.

How should I test a production integration?

Keep representative interactive, flattened, text, and scanned fixtures; assert status, content type, required values, and failure behavior; and rerun them when upgrading Stirling-PDF.