How to Scrape Data from PDFs
Extract searchable text, tables, or OCR text from PDFs with Python. Learn how to choose a method, export results, and check them against the source.
To scrape data from a PDF, first identify what kind of content it contains. If you can select and copy its text, extract the text layer with a PDF library. For tables, use a table-aware extractor and inspect the cells. If the page is a scan or image, use OCR. In every case, treat the extracted data as a draft and check it against the rendered pages.
This guide uses Python for the extraction workflows. It covers searchable text, tables, and scanned pages, with options for CSV output and a browser based alternative for capturing web pages as PDFs.
1. Identify what is inside the PDF
Open the file and try selecting a sentence. If you can select individual words, the document likely has a machine-readable text layer. If the page behaves like one large image, or copying produces no useful text, it needs OCR. Tables require their own check: text can extract correctly while the rows and columns are scrambled.
| PDF content | First method to try | Typical output |
|---|---|---|
| Selectable text | PyMuPDF text extraction | Text per page |
| Tables with visible rules | PyMuPDF table detection or Camelot lattice | Rows and columns |
| Borderless tables | Text based table detection, such as PyMuPDF’s text strategy or Camelot stream | Rows and columns, subject to layout checks |
| Scanned pages or text stored as images | OCR with Tesseract through PyMuPDF | Recognized text |
No one method reliably handles every PDF layout. Page structure, table borders, scan quality, rotation, and small print affect what can be recovered. For table extraction, PyMuPDF documents that line based detection can miss borderless tables and tables shown through background colors. PyMuPDF’s table extraction FAQ describes trying text based detection when drawn borders are absent.
2. Set up Python and install the libraries
Use a virtual environment so the PDF tools are isolated from other projects. For the text, table, and OCR examples below, install PyMuPDF. Install Camelot only if you want its alternate table extraction workflow.
python -m venv .venv
# macOS or Linux
source .venv/bin/activate
# Windows PowerShell
# .venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install pymupdf
For Camelot, install its package separately and follow the current installation guidance for your operating system and selected parser options: Camelot quickstart. OCR also requires Tesseract installed as a separate application; installing the Python package alone does not install the OCR engine. See PyMuPDF’s OCR guide.
3. Extract searchable text page by page
Save this as extract_text.py beside input.pdf. It writes one UTF-8 text file with page markers, so you can trace a result to its original page.
import pymupdf
input_path = "input.pdf"
output_path = "extracted.txt"
with pymupdf.open(input_path) as document:
with open(output_path, "w", encoding="utf-8") as output:
for page_number, page in enumerate(document, start=1):
text = page.get_text("text")
output.write(f"\\n\\n--- Page {page_number} ---\\n")
output.write(text)
print(f"Wrote {output_path}")
Run it with:
python extract_text.py
Page.get_text("text") is a good starting point for ordinary reading order. Other output modes are useful when you need positions or structure: PyMuPDF supports text extraction in formats including blocks, words, HTML, and dictionaries. For documents with multiple columns or unusual reading order, inspect the output and consider extracting words or blocks with coordinates so you can reconstruct the intended order.
Keep page references and handle empty pages
Page markers are valuable when reviewing, indexing, or importing records. An empty string can mean the page has no text layer, but it can also mean that text is encoded or positioned unusually. Render and inspect the page before deciding it needs OCR.
4. Extract tables to structured data
PyMuPDF includes table finding on pages. The following script finds detected tables and writes each one to its own CSV file, with the source page represented in the filename.
import csv
import pymupdf
input_path = "input.pdf"
with pymupdf.open(input_path) as document:
for page_number, page in enumerate(document, start=1):
result = page.find_tables()
for table_number, table in enumerate(result.tables, start=1):
rows = table.extract()
output_path = f"page-{page_number}-table-{table_number}.csv"
with open(output_path, "w", newline="", encoding="utf-8") as output:
writer = csv.writer(output)
writer.writerows(rows)
print(f"Wrote {output_path}: {len(rows)} rows")
Run with python extract_tables.py after saving the snippet as that filename. Install PyMuPDF as described above.
By default, table detection relies on vector graphics such as lines and rectangles. For a borderless table, try a text based strategy:
result = page.find_tables(strategy="text")
Detection settings include strategies and tolerances for snapping nearby lines, joining segments, and identifying intersections. Adjust them only after looking at the page and the detected result. A setting that improves one page can introduce false rows or columns elsewhere in the file. The Page API documentation lists the available find_tables() arguments.
Use Camelot when you want a table focused workflow
Camelot exports tables as CSV, JSON, Excel, HTML, Markdown, and SQLite. Its documented parser options include lattice and stream: lattice is suited to tables defined by lines, while stream uses whitespace and text positioning to infer structure. Choose based on the visual layout, then review each exported table.
import camelot
tables = camelot.read_pdf("input.pdf", pages="1-end", flavor="lattice")
print(f"Detected {tables.n} tables")
tables.export("table.csv", f="csv")
For a borderless table, try flavor="stream". The export method may create a separate file per detected table, with page and table numbers in the filenames. See the Camelot quickstart and parser overview for supported formats and parser details.
Do not choose a library based on a single claimed accuracy number. A table extractor’s output depends on the document’s construction and layout; evaluate it on representative pages and compare it with the source.
5. OCR scanned pages
PyMuPDF can create an OCR text page using Tesseract. The example below checks each page for ordinary text, OCRs pages with no extractable text, and writes the result. Tesseract must be installed and available to PyMuPDF first.
import pymupdf
input_path = "scanned.pdf"
output_path = "ocr_text.txt"
with pymupdf.open(input_path) as document:
with open(output_path, "w", encoding="utf-8") as output:
for page_number, page in enumerate(document, start=1):
text = page.get_text("text").strip()
if not text:
text_page = page.get_textpage_ocr(language="eng", dpi=200, full=True)
text = page.get_text("text", textpage=text_page)
output.write(f"\\n\\n--- Page {page_number} ---\\n{text}")
print(f"Wrote {output_path}")
Adjust language to the installed Tesseract language data; multiple languages can be specified with a plus sign, such as eng+spa. dpi affects image resolution and recognition time. full=True requests OCR for the full page; the default can focus on image areas and retain regular text. For mixed pages, test the behavior against the specific document.
OCR is much slower than ordinary text extraction. PyMuPDF’s documentation estimates it at about one thousand times slower and recommends doing OCR once per page, storing the returned text page, then reusing it for extraction or search. It also notes that OCR text does not preserve original font emphasis and that Tesseract does not recognize vector drawings as text. See the OCR guide.
6. Validate the extracted result
Before sending extracted data into a database, spreadsheet, or downstream model, compare it with the rendered source. Check more than whether the script completed: parsers can return plausible looking but incorrectly aligned content.
- For text, compare paragraph order, columns, headers, footers, and page numbers.
- For tables, verify column count, header placement, empty cells, merged cells, totals, and rows that continue across pages.
- For OCR, inspect names, dates, decimal points, minus signs, small type, rotated text, and similar looking characters.
- Record the PDF page for each extracted record where traceability matters.
- Use representative pages from each distinct layout, not just the first page.
For difficult tables, crop to the relevant area or inspect the parser’s detected cells and boundaries. pdfplumber is another Python option when character coordinates and visual table debugging are useful; it exposes text and table methods, plus configurable line and text strategies. Its project documentation says it works best on machine generated rather than scanned PDFs. See the pdfplumber documentation.
7. Choose a workflow by output and constraints
| Need | Starting point | Tradeoff to plan for |
|---|---|---|
| Plain text from digital PDFs | PyMuPDF get_text() |
Reading order may need review on columns and complex layouts |
| Ruled tables | PyMuPDF find_tables() or Camelot lattice |
Decorative or broken rules can confuse cell boundaries |
| Borderless tables | PyMuPDF text strategy, pdfplumber text strategy, or Camelot stream | Whitespace and alignment heuristics can split or merge cells |
| Scanned documents | PyMuPDF OCR with Tesseract | Requires a separate OCR installation and substantially more processing time |
| Detailed positions and debugging | pdfplumber | More layout-specific tuning may be needed |
8. Performance, reliability, and cost
Local extraction with the libraries in this guide has no per request API fee. The costs are compute, installation, maintenance, and review time. Ordinary text extraction is generally the lightest path. Table detection adds layout work, and OCR is substantially slower; avoid running OCR on pages that already have a usable text layer.
For repeatable jobs, store the source file identifier, extraction method, library versions, page number, and review status alongside the output. Cache OCR text per page rather than rerunning it on every query. For large files, process and persist pages incrementally so a failure does not require discarding all prior output.
PDFs can be malformed, encrypted, or restricted from text extraction. Use an authorized password when required and respect document permissions. If a library rejects a file or page, preserve the original, isolate the failing page, and test a rendered or repaired copy rather than silently dropping content. No parser can promise correct output for every arbitrary layout, so high consequence data needs human verification.
9. Troubleshooting
| Symptom | Likely cause | What to try |
|---|---|---|
| Extracted text is empty | The page is scanned, image based, or uses unusual text encoding | Render the page and inspect it; use OCR if it is image based |
| Text appears in the wrong order | Multiple columns, positioned text, or layout complexity | Extract blocks or words with coordinates and reconstruct per region; compare with rendered page |
| No table is detected | Borderless cells, color-only boundaries, or unusual vector construction | Try a text strategy, crop the page, or use a different table parser; inspect detections visually |
| Extra or missing table columns | Broken rules, merged cells, inconsistent spacing, or decorative lines | Tune tolerances or table area for that layout, then compare every row with the source |
| OCR import or execution fails | Tesseract is missing, not on the executable path, or language data is absent | Install Tesseract and the required language data, then rerun on one page |
| OCR output has incorrect characters | Low resolution, skew, rotation, small print, or poor scan quality | Inspect the rendered page, adjust OCR resolution, rotate or improve the source scan, and manually verify critical fields |
| Encrypted PDF cannot be opened or extracted | Password required or extraction permissions restrict access | Use the authorized password and confirm you have permission to extract the content |
| CSV has broken rows or odd characters | Cells contain line breaks, encoding issues, or inconsistent table structure | Open with a UTF-8 aware tool, preserve quoting, and inspect the original cells before normalizing values |
10. FAQ
Can I scrape a PDF directly into Excel?
Yes. Extract tables to CSV and open that file in Excel, or use a table tool that exports Excel workbooks. Check headers, types, and merged cells before relying on the spreadsheet.
Can I extract only selected pages?
Yes. Iterate over the page numbers you need rather than every page. Camelot also accepts a page selection such as "1,3,5-8" in its pages argument.
Will OCR recover charts and diagrams?
OCR recognizes text in images; it does not turn a chart’s geometry or a diagram’s relationships into structured data. Extracting those requires a separate, layout-specific process.
Or skip the browser setup
If the PDF you need to process is a web page you can access by URL, ScreenshotNeo can capture that page as an image or PDF with one GET request. It does not extract text or tables from an existing PDF; use the Python workflows above for that.
Example cURL request:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
For Python, Node.js, available parameters, and response details, see the ScreenshotNeo API documentation.
- Cookie banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing status.
- An MCP server lets AI agents use
take_screenshot,get_page_info, andcapture_pdf. - The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for ScreenshotNeo and get 1,000 free screenshots a month with no card.


