ScreenshotNeo

BlogHow-to

How to Extract Data from PDFs with Amazon Bedrock

Choose the right Bedrock parser for text, scans, tables, and reusable PDF search, with working setup guidance, code, costs, and troubleshooting.

By the ScreenshotNeo team1 October 20268 min read

Short answer: use a Bedrock Knowledge Base with the default parser when your PDFs contain selectable text and you need repeatable search. Choose Bedrock Data Automation (BDA) or a foundation-model parser when figures, charts, tables, images, or page layout affect the answer. For a one-off document, a direct model request may be simpler. Scanned PDFs need OCR or visual processing; the AWS Textract tutorial for Bedrock covers single-page JPG or PNG inputs and explicitly does not cover the separate asynchronous workflow required for multi-page PDFs.

“PDF extraction” in Bedrock is not one API or one parser. The right design depends on document type, whether you need one answer or a reusable corpus, and how much visual structure must be preserved.

1. Choose the extraction path

Situation Recommended path Why
Selectable text, reusable corpus Knowledge Bases default parser Parses text, then chunks, embeds, and indexes it. AWS says parsing has no usage charge.
Charts, tables, figures, images, or layout matter BDA or foundation-model parser Both support multimodal extraction for Knowledge Base retrieval.
Need a custom extraction instruction Foundation-model parser Its extraction prompt can be customized.
One document or small application-controlled job Direct model request, if the selected model supports the document input A vector store and Knowledge Base may be unnecessary.
Scanned pages Textract OCR plus Bedrock, or a visual parser OCR or visual interpretation is required before reliable text extraction.

Read the AWS documentation for parser behavior and pricing, Knowledge Base ingestion, and retrieval APIs before selecting a production design.

2. Build a Knowledge Base for text PDFs

  1. Upload PDFs to a supported unstructured data source. AWS demonstrates Amazon S3 for this workflow.
  2. Create an IAM role that grants Bedrock access to the source, embedding model, and vector store. Restrict permissions to the resources used by the Knowledge Base.
  3. Select the default parser for text-only documents. If you select BDA or a foundation-model parser, that parser is used for every PDF in the data source, including text-only files.
  4. Choose chunking, an embedding model, and a vector store.
  5. Start ingestion or sync. Bedrock parses, chunks, embeds, and writes vectors to the store.
  6. Query with Retrieve when your application will control answer generation, or RetrieveAndGenerate when Bedrock should generate a grounded answer with source attribution.

Keep collections with different parsing needs separate when that reduces multimodal parsing charges. Verify regional availability and current pricing for the parser, embedding model, vector store, and page volume.

Python setup and retrieval with boto3

import boto3

region = "us-east-1"
knowledge_base_id = "YOUR_KNOWLEDGE_BASE_ID"

client = boto3.client("bedrock-agent-runtime", region_name=region)

response = client.retrieve(
    knowledgeBaseId=knowledge_base_id,
    retrievalQuery={"text": "What is the cancellation policy?"},
    retrievalConfiguration={
        "vectorSearchConfiguration": {
            "numberOfResults": 5
        }
    }
)

for result in response.get("retrievalResults", []):
    text = result.get("content", {}).get("text", "")
    score = result.get("score")
    location = result.get("location", {})
    print(f"score={score} location={location}\n{text}\n")

Use retrieveAndGenerate when you want Bedrock to combine retrieval and generation. Preserve the returned citations or source chunks in your application so users can inspect the supporting pages.

CLI query example

aws bedrock-agent-runtime retrieve \
  --region us-east-1 \
  --knowledge-base-id YOUR_KNOWLEDGE_BASE_ID \
  --retrieval-query '{"text":"List the renewal terms"}' \
  --retrieval-configuration '{"vectorSearchConfiguration":{"numberOfResults":5}}'

3. Extract tables and visual content

The default parser extracts text but does not extract visual content from charts, figures, tables, or images. Select BDA for managed multimodal processing, or a foundation-model parser when you need to adjust the extraction prompt. BDA is billed by pages or images processed. Foundation-model parsing is billed by input and output tokens.

The parser choice applies to every PDF in that data source. A mixed source containing hundreds of text-only PDFs and a few charts can therefore incur advanced-parser charges for all of them. Separate sources when appropriate, and check current regional prices before estimating cost.

Foundation-model parser prompt considerations

  • Describe the fields and units you need.
  • Tell the parser to preserve table headers, row relationships, footnotes, and page references.
  • Specify how to represent missing, illegible, or ambiguous values.
  • Keep the prompt focused; validate extracted values against the source page.

Multimodal extraction improves retrieval of visual evidence, but it does not make extracted values ground truth. Review critical financial, legal, medical, compliance, and operational fields against the original page.

4. One-off PDF extraction

For a single document or a small application-controlled workload, a direct model request can avoid Knowledge Base and vector-store setup. Bedrock’s Converse API provides a common message interface for supported models. The API reference does not establish that every model accepts PDF bytes or identical document formats, so confirm the selected model’s document-input support, limits, region, and permissions first. If direct document input is unavailable, extract text or render page images before calling the model.

Generic Converse request shape

import boto3

client = boto3.client("bedrock-runtime", region_name="us-east-1")

response = client.converse(
    modelId="YOUR_MODEL_ID",
    messages=[
        {
            "role": "user",
            "content": [
                {"text": "Extract invoice_number, invoice_date, total, and currency. Return strict JSON."}
                # Add a document or image content block only when your chosen model supports it.
            ],
        }
    ],
    inferenceConfig={"temperature": 0}
)

print(response["output"]["message"]["content"])

Grant the runtime permission required by the chosen operation, including bedrock:InvokeModel where applicable. Treat model output as untrusted data: validate the JSON schema, types, required fields, and totals.

5. Scanned PDFs and OCR

A scanned PDF may contain only page images. OCR or visual interpretation must happen before the application can reliably work with its text. AWS’s Bedrock/Textract tutorial uses DetectDocumentText with a single-page JPG or PNG image. It explicitly excludes multi-page PDFs, which require Textract’s separate asynchronous document-processing workflow.

For production multi-page scans, confirm the current asynchronous Textract operation, input constraints, output format, region, and pricing before implementing it. Then pass the OCR text or structured blocks to Bedrock for classification, extraction, or question answering. Preserve page and block coordinates when downstream users need to verify a value visually.

6. Retrieve the original or parsed document

When an interface needs to show or download the source behind a Knowledge Base result, use GetDocumentContent. The response includes a pre-signed URL and MIME type. The URL expires after five minutes. The caller needs both bedrock:Retrieve and bedrock:GetDocumentContent. Pass user identity context when ACL-based access control is enabled.

7. Syncing, updates, and deletion

Run a data-source sync after adding, editing, or deleting PDFs so the Knowledge Base reflects the source. Some sources also support direct ingestion and deletion operations. Design your pipeline to record document identifiers, sync status, parser choice, embedding model, and ingestion timestamps. Re-index after changing chunking, embeddings, or parser configuration.

8. Cost, performance, and reliability

  • Parser cost: AWS says the default parser has no usage charge. BDA is charged by pages or images; foundation-model parsing by input and output tokens. Advanced parser charges apply to every PDF in the selected data source.
  • Embedding and storage: ingestion also creates embeddings and writes vectors, so include those services in estimates.
  • Query latency: Retrieve returns chunks for your own model call; RetrieveAndGenerate adds generation. Limit the number of retrieved results to the evidence your prompt needs.
  • Large documents: use deliberate chunking and metadata filters. Oversized chunks can reduce retrieval precision; tiny chunks can lose context.
  • Reliability: retry transient AWS errors with bounded exponential backoff, make ingestion jobs observable, and record failed documents for replay.
  • Validation: compare critical values with the source page, especially for low-quality scans, tables, handwriting, and compliance-sensitive material.
  • Security: keep source buckets private, use least-privilege IAM, avoid logging sensitive document contents, and scope pre-signed URL exposure to the five-minute access window.

AWS’s tutorial reports an estimate of less than USD 0.15 when completed within two hours and the notebook is deleted afterward. That is a bounded tutorial setup estimate, not a production workload forecast. Calculate production cost from page count, parser, model tokens, embedding, storage, queries, region, and retention.

9. Troubleshooting

Symptom Likely cause Fix
Tables or charts are missing Default text parser selected Use BDA or a foundation-model parser and re-sync the source.
Costs are higher than expected Advanced parser applied to all PDFs Split text-only and multimodal collections; verify current regional pricing.
No new document appears in results Source changed but sync did not run or finish Run a sync, inspect its status, and retry failed documents.
Scanned PDF returns little or no text Pages are images without OCR Use the appropriate asynchronous Textract flow for multi-page PDFs or a visual parser.
Direct Converse request rejects the document Model or region does not support that document input Check model-specific limits; extract text or page images first.
Retrieved answer lacks evidence Too few or poorly sized chunks, or generation not grounded Adjust chunking and result count, use Retrieve for controlled prompting, and display citations.
Source download fails later GetDocumentContent URL expired Request a new URL; pre-signed links expire after five minutes.
Access denied Missing IAM action or ACL identity context Grant bedrock:Retrieve and bedrock:GetDocumentContent; pass identity context when required.

10. Or skip the browser setup

If your workflow needs page images of PDFs or web documents before sending them to Bedrock, ScreenshotNeo provides a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf.

See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

There are 1,000 screenshots per month free with no card. Paid plans start at $5 for 3,000 screenshots, and every feature is included on every plan. Start with a free ScreenshotNeo account.

11. FAQ

Does Bedrock automatically understand every PDF?

No. Select the parser and model path for the document type. Text-only, scanned, and visually rich PDFs require different handling.

Should I always use a Knowledge Base?

No. Use one for a reusable searchable corpus. A direct model request can be simpler for a one-off document when the model supports its input format.

Which parser is cheapest?

The default parser has no AWS usage charge for parsing, but it is intended for text extraction. BDA and foundation-model parsing add charges and should be selected when their visual capabilities are needed.

Can I trust extracted totals without review?

No. Validate important values against the source page, especially when scans, tables, handwriting, or low-quality images are involved.

The pre-signed URL expires after five minutes, so request it when the user needs access.