ScreenshotNeo

BlogHow-to

How to Convert PDF to Excel with PowerShell

Convert PDF tables to XLSX with a reliable PowerShell workflow: extract, normalize, validate, and export with the right tool for each PDF.

By the ScreenshotNeo team1 October 20268 min read

Short answer: PowerShell does not include a universal PDF-table converter. Use it to orchestrate two explicit stages: extract tables with a PDF-aware tool, then write validated rows to .xlsx with a workbook tool such as ImportExcel. For scanned PDFs, add OCR before table extraction.

This separation matters. ImportExcel creates Excel workbooks from structured data; it does not establish PDF table extraction capability. Camelot provides PDF table extraction strategies, while Excel Power Query and Adobe Acrobat provide GUI-based alternatives. Treat every extraction as data that must be checked against the source PDF.

1. Choose the right workflow

Approach Strength Limit Best fit
PowerShell + PDF extractor + ImportExcel Automatable, repeatable, and can create XLSX without Excel installed Extraction and workbook writing are separate steps Batch jobs and PowerShell pipelines
Excel Power Query PDF import Detected tables appear in Navigator for inspection and transformation Requires the documented connector components and .NET Framework 4.5 or higher Occasional imports and interactive review
Adobe Acrobat export Direct PDF-to-XLSX flow with worksheet, numeric-separator, and text-recognition settings Commercial software; verify current access and terms GUI conversion and OCR-oriented workflows

Microsoft documents the Excel route under Power Query data-source imports. Adobe documents XLSX conversion and recognition settings in its Acrobat conversion help and PDF-to-Excel guide. Camelot’s extraction modes are described in its Quickstart documentation.

2. Classify the PDF before writing code

  • Text PDF: You can select and copy table text. A PDF-aware extractor can usually read its character positions.
  • Scanned PDF: Each page is an image. Run OCR first, then inspect the extracted text and table boundaries.
  • Hybrid PDF: Some pages contain text and others are scans. Process pages separately and merge only after validation.

OCR can make scanned text searchable, but it does not guarantee that every row, column, decimal, or merged cell is reconstructed correctly. Adobe describes text-recognition behavior and export settings in its documentation; use those settings deliberately and compare the result with the source.

3. Install the PowerShell workbook stage

ImportExcel is a PowerShell module for creating and reading Excel files without requiring Microsoft Excel. Install it from the PowerShell Gallery:

Install-Module ImportExcel -Scope CurrentUser
Import-Module ImportExcel
Get-Command Export-Excel

Export-Excel expects objects, CSV rows, or other structured input. It is the output stage of this workflow, not the PDF parser.

4. Extract a table with Camelot, then export it from PowerShell

Camelot is a Python library and command-line tool, so PowerShell invokes it as an external process or runs a Python script. Its documented strategies include lattice for ruled tables, stream for whitespace-aligned tables, and network, hybrid, and automatic approaches. Match the method to the page layout.

4.1 Create a small Python extractor

Install Python and Camelot according to the current Camelot documentation. Save this as extract_tables.py:

import argparse
import csv
from pathlib import Path

import camelot

parser = argparse.ArgumentParser()
parser.add_argument("pdf")
parser.add_argument("csv_path")
parser.add_argument("--pages", default="1-end")
parser.add_argument("--flavor", choices=["lattice", "stream"], default="lattice")
args = parser.parse_args()

# Use lattice when visible ruling lines define cells; use stream for aligned text.
tables = camelot.read_pdf(
    args.pdf,
    pages=args.pages,
    flavor=args.flavor,
)

if tables.n == 0:
    raise SystemExit("No tables detected")

output = Path(args.csv_path)
with output.open("w", newline="", encoding="utf-8-sig") as handle:
    writer = csv.writer(handle)
    for table_index, table in enumerate(tables):
        data = table.df.fillna("").values.tolist()
        if table_index:
            writer.writerow([])
        writer.writerows(data)

print(f"Wrote {tables.n} table(s) to {output}")

4.2 Orchestrate extraction and XLSX creation in PowerShell

This script checks the input, calls the extractor, imports the resulting CSV, and writes an XLSX file. It keeps each stage inspectable so you can validate before publishing the workbook.

param(
    [Parameter(Mandatory = $true)]
    [string] $PdfPath,

    [string] $Python = "python",
    [ValidateSet("lattice", "stream")]
    [string] $Flavor = "lattice",
    [string] $Pages = "1-end",
    [string] $OutputPath = ""
)

$ErrorActionPreference = "Stop"

$pdf = (Resolve-Path -LiteralPath $PdfPath).Path
if ([string]::IsNullOrWhiteSpace($OutputPath)) {
    $OutputPath = [IO.Path]::ChangeExtension($pdf, ".xlsx")
}
$work = Join-Path ([IO.Path]::GetDirectoryName($pdf)) ([IO.Path]::GetFileNameWithoutExtension($pdf) + "-extracted.csv")
$extractor = Join-Path $PSScriptRoot "extract_tables.py"

if (-not (Test-Path -LiteralPath $extractor)) {
    throw "Missing extractor: $extractor"
}

& $Python $extractor $pdf $work --pages $Pages --flavor $Flavor
if ($LASTEXITCODE -ne 0) {
    throw "PDF extraction failed with exit code $LASTEXITCODE"
}

Import-Module ImportExcel
$rows = Import-Csv -LiteralPath $work -Header Column1,Column2,Column3,Column4,Column5,Column6,Column7,Column8
$rows = $rows | Where-Object {
    ($_.PSObject.Properties.Value -join "").Trim().Length -gt 0
}

$rows | Export-Excel -Path $OutputPath -WorksheetName "Extracted" -AutoSize -FreezeTopRow -BoldTopRow

Write-Host "Created $OutputPath"
Write-Host "Intermediate CSV: $work"

Run it with:

pwsh ./Convert-PdfToExcel.ps1 -PdfPath ./invoice.pdf -Flavor lattice
pwsh ./Convert-PdfToExcel.ps1 -PdfPath ./statement.pdf -Flavor stream -Pages 1-5 -OutputPath ./statement.xlsx

The generic CSV headers in this example are intentional: extraction often produces inconsistent column counts across pages. For production data, inspect the CSV, define the expected columns, convert types explicitly, and reject rows that do not meet your schema.

5. Normalize and validate the extracted rows

Before exporting a workbook that other systems will consume, check:

  • Column boundaries and headers, including repeated page headers.
  • Merged cells and labels that span multiple rows.
  • Rows split across page breaks.
  • Negative numbers, currency symbols, decimal separators, and thousands separators.
  • Dates and locale assumptions.
  • Blank cells that mean zero versus blank or unavailable.
  • Subtotal and total rows that should not be treated as ordinary records.
  • Duplicate rows created by overlapping extraction regions.

A simple PowerShell validation pass can flag rows with unexpected widths:

$expectedColumns = 6
$badRows = Import-Csv ./invoice-extracted.csv -Header C1,C2,C3,C4,C5,C6 |
    Where-Object { ($_.PSObject.Properties.Value | Where-Object { $_ -ne $null }).Count -ne $expectedColumns }

if ($badRows) {
    $badRows | Format-Table
    throw "Validation failed: one or more rows have an unexpected shape."
}

For financial or operational data, compare representative pages and totals with the PDF manually. Official documentation does not provide a universal conversion-accuracy percentage, and PDF layout variation makes a single success rate misleading.

6. Handle scanned PDFs and OCR

  1. Run OCR with a tool that supports the document’s language and layout.
  2. Export or save a searchable PDF.
  3. Extract tables from the OCR result.
  4. Review characters commonly confused by OCR, such as 0/O, 1/I, decimal points, and minus signs.
  5. Reconcile row counts and totals with the original scan.

OCR improves text access; it does not recreate table structure perfectly. If a scan contains skew, shadows, handwriting, or low contrast, expect more manual correction.

7. Excel and Acrobat alternatives

Excel Power Query

In Excel, choose Data > Get Data > From File > From PDF. Select detected tables in Navigator, then load them or transform them. Microsoft notes that the PDF connector requires .NET Framework 4.5 or higher and may show a message that additional components must be installed.

Adobe Acrobat

Acrobat can export a PDF to Microsoft Excel XLSX. Its documented settings include worksheets per table, page, or document; numeric separators; and text recognition options. Use this route when a GUI review or OCR setting is more useful than a repeatable PowerShell pipeline.

8. Troubleshooting

Symptom Likely cause Fix
No tables detected The PDF is scanned, protected, or the chosen flavor does not match its layout Check text selection, run OCR if needed, then try stream or lattice on a small page range
Columns are shifted Whitespace or ruling lines do not describe the real cells Switch extraction strategy, adjust the PDF-aware tool’s table settings, and inspect page-level output
Repeated headers appear as records Each page header was extracted as a normal row Remove known header rows after identifying them by exact normalized values
Rows are split at page breaks A record continues on the next page Merge continuation rows using a stable key, or extract affected pages with a layout-specific setting
Numbers import as text Currency symbols, separators, or OCR characters remain Normalize strings first, then cast with the intended locale and validate totals
Export-Excel is not recognized ImportExcel is not installed or imported Run Install-Module ImportExcel -Scope CurrentUser and Import-Module ImportExcel
Excel shows a connector component error Power Query PDF prerequisites are missing Install the components listed by Microsoft and verify the .NET Framework requirement
Python exits with an error Python, Camelot, or a PDF backend is unavailable Run the extractor directly, read the reported dependency error, and install the dependency using Camelot’s current instructions

9. Performance, reliability, and cost

  • Performance: Process only the required page range, run a representative page first, and avoid reprocessing unchanged PDFs. Large scans spend most of their time in OCR and image analysis.
  • Reliability: Keep the intermediate CSV, log the PDF hash and extraction options, and fail the job when row counts or totals violate your checks.
  • Repeatability: Pin your Python and module versions in the environment used for scheduled jobs, and store the exact script alongside the output.
  • Cost: PowerShell and ImportExcel do not require Excel for workbook creation. OCR, Acrobat, or hosted extraction services may have separate licensing or usage costs; verify current terms before deployment.

10. Or skip the browser setup

If the PDF starts as a web page or you need a clean visual capture of a source page before processing, ScreenshotNeo provides a single HTTP request for a PNG, JPEG, WebP, or PDF. Its consent step accepts cookie banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing result.

See the ScreenshotNeo API documentation for the complete option list.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also has an MCP server so Claude, Cursor, and other MCP clients can take screenshots, inspect pages, and capture PDFs. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Can PowerShell convert a PDF to Excel by itself?

No. Use PowerShell as the orchestrator, a PDF-aware extractor for tables, and an XLSX writer such as ImportExcel.

Should I use lattice or stream extraction?

Use lattice when visible ruling lines define cells. Use stream when columns are aligned by whitespace without rules. Test both on representative pages when the layout is unclear.

Will OCR preserve the original table exactly?

No. OCR can expose text in a scan, but merged cells, symbols, and column boundaries still require inspection and validation.

Do I need Microsoft Excel installed for ImportExcel?

No. ImportExcel is designed to create workbooks from PowerShell without Excel installed.

What should I do with multi-table PDFs?

Extract and validate each table separately, preserve page or table identifiers, then combine normalized records into the final workbook.