ScreenshotNeo

BlogHTML to image & PDF

How to Convert PDF to Text with PowerShell

Use PowerShell to run Apache PDFBox, extract selectable PDF text, save it, and troubleshoot versions, passwords, page ranges, and scanned files.

By the ScreenshotNeo team1 October 20266 min read

PowerShell does not parse PDF files by itself. Use PowerShell to launch a PDF extraction tool such as Apache PDFBox, then read or process the generated text file. PDFBox 3.x uses the export:text command; PDFBox 2.x uses the older ExtractText command. Check the version you downloaded before choosing a command.

What you need

Convert a PDF with PDFBox 3.x

Place the PDF and the PDFBox application JAR in a working directory. Replace the placeholder JAR name with the exact filename you downloaded.

java -jar .\pdfbox-app-3.y.z.jar export:text -i=.\input.pdf -o=.\output.txt

PDFBox documents -i/--input and -o/--output for this command. The documented default output encoding is UTF-8.

Run the same command from PowerShell

$pdfBox = Join-Path $PWD 'pdfbox-app-3.y.z.jar'
$inputPdf = Join-Path $PWD 'input.pdf'
$outputTxt = Join-Path $PWD 'output.txt'

java -jar $pdfBox export:text -i=$inputPdf -o=$outputTxt

if ($LASTEXITCODE -ne 0) {
    throw "PDFBox failed with exit code $LASTEXITCODE"
}

if (-not (Test-Path -LiteralPath $outputTxt)) {
    throw "Expected output file was not created: $outputTxt"
}

Get-Content -LiteralPath $outputTxt -Raw

Get-Content -Raw returns the complete file as one string. Without -Raw, PowerShell returns an array of lines. Microsoft documents both behaviors in its Get-Content reference.

Use PDFBox 2.x syntax

Do not combine the 2.x command with the 3.x command. PDFBox 2.x documents this form:

java -jar .\pdfbox-app-2.y.z.jar ExtractText .\input.pdf .\output.txt

Use the exact version-specific help output if your release exposes different options.

A reusable PowerShell script

This script supports PDFBox 3.x, validates paths, and lets you pass optional arguments such as page ranges or a password.

param(
    [Parameter(Mandatory = $true)]
    [string]$PdfPath,

    [Parameter(Mandatory = $true)]
    [string]$PdfBoxJar,

    [string]$OutputPath,

    [string]$Password,

    [int]$StartPage,

    [int]$EndPage,

    [switch]$Sort
)

$ErrorActionPreference = 'Stop'

if (-not (Test-Path -LiteralPath $PdfPath -PathType Leaf)) {
    throw "Input PDF not found: $PdfPath"
}
if (-not (Test-Path -LiteralPath $PdfBoxJar -PathType Leaf)) {
    throw "PDFBox JAR not found: $PdfBoxJar"
}
if (-not $OutputPath) {
    $OutputPath = [System.IO.Path]::ChangeExtension((Resolve-Path -LiteralPath $PdfPath), '.txt')
}

$args = @(
    '-jar', $PdfBoxJar,
    'export:text',
    "-i=$((Resolve-Path -LiteralPath $PdfPath).Path)",
    "-o=$OutputPath"
)

if ($Password) { $args += "-password=$Password" }
if ($StartPage) { $args += "-startPage=$StartPage" }
if ($EndPage) { $args += "-endPage=$EndPage" }
if ($Sort) { $args += '--sort' }

& java @args
if ($LASTEXITCODE -ne 0) {
    throw "PDFBox failed with exit code $LASTEXITCODE"
}

if (-not (Test-Path -LiteralPath $OutputPath -PathType Leaf)) {
    throw "PDFBox completed without creating $OutputPath"
}

Write-Output "Created $OutputPath"
Get-Content -LiteralPath $OutputPath -Raw

PDFBox 3.x documents page selection, password handling, sorting, and other text-export options. Option names can change between releases, so run the installed version’s help and compare it with the official 3.x documentation.

Launch PDFBox with Start-Process

PowerShell’s call operator (&) is usually simplest when you need the exit code. Start-Process is useful when you need a separate process object or redirected streams:

$arguments = @(
    '-jar', '.\pdfbox-app-3.y.z.jar',
    'export:text',
    '-i=.\input.pdf',
    '-o=.\output.txt'
)

$process = Start-Process -FilePath 'java' -ArgumentList $arguments -Wait -PassThru
if ($process.ExitCode -ne 0) {
    throw "Java/PDFBox failed with exit code $($process.ExitCode)"
}

Get-Content -LiteralPath '.\output.txt' -Raw

Microsoft documents Start-Process for launching executables and warns that untrusted data should not be used as its FilePath. Keep the executable path fixed or validate it before invoking it.

Useful extraction options

Need Approach
Only selected pages Use the PDFBox 3.x start and end page options documented for export:text.
Password-protected PDF Supply the documented password option. A wrong password or restricted file can stop extraction.
Reading order Try the documented sorting option, then inspect results because multi-column layouts can still require cleanup.
UTF-8 output UTF-8 is the documented PDFBox 3.x default; verify the installed release if you override encoding.
Markdown output PDFBox documentation notes Markdown output availability since 3.0.4. Confirm support in your installed version.

Process many PDFs

$pdfBox = Join-Path $PWD 'pdfbox-app-3.y.z.jar'
New-Item -ItemType Directory -Force -Path '.\text' | Out-Null

Get-ChildItem -LiteralPath '.\pdfs' -Filter '*.pdf' -File | ForEach-Object {
    $destination = Join-Path '.\text' ($_.BaseName + '.txt')
    & java -jar $pdfBox export:text "-i=$($_.FullName)" "-o=$destination"
    if ($LASTEXITCODE -ne 0) {
        Write-Error "Failed: $($_.FullName)"
    }
}

For large batches, write outputs to a separate directory, preserve the source filename, and log each exit code. Avoid running an unbounded number of Java processes at once; start with sequential processing and add controlled parallelism only after measuring memory use.

When the PDF is a scan

A scanned PDF may contain only page images. PDFBox text extraction cannot turn those pixels into words unless a text layer is present. Test a sample by opening the PDF and trying to select or copy text. If selection is impossible, you need an OCR workflow; the researched PDFBox command-line material does not establish an OCR method, so do not assume the commands above will recognize the scan.

Troubleshooting

Symptom Likely cause Fix
java is not recognized Java is missing or not on PATH. Install Java, reopen PowerShell, and verify with java -version.
Unable to access jarfile The JAR filename or working directory is wrong. Use an absolute path or list files with Get-ChildItem; copy the exact filename.
Unknown command or option You used PDFBox 2.x syntax with 3.x, or the reverse. Check the JAR release and use export:text for 3.x or ExtractText for 2.x.
Output is empty The PDF may be image-only, encrypted, or contain unusual font/layout data. Check for selectable text, provide the correct password, and inspect a different page.
Text order is confusing Columns, positioned text, headers, and footers do not map cleanly to reading order. Try the documented sorting option and post-process the text for your document layout.
File not found in a script Relative paths resolve from the current directory, which may differ in scheduled jobs. Resolve paths explicitly and use -LiteralPath for filenames containing wildcard characters.
PowerShell appears hung Java is still processing a large or complex PDF. Wait for the process, check CPU and memory, and add your own timeout or job supervision around the process.

Performance, reliability, and cost

  • Performance: Extraction time depends on PDF size, page count, fonts, images, and layout complexity. Selecting a page range reduces unnecessary work.
  • Reliability: Check the process exit code and confirm the output file exists before consuming it. Keep source PDFs and generated text separately so failed runs can be retried.
  • Encoding: Keep UTF-8 unless a downstream system requires another encoding, and validate non-ASCII text with representative documents.
  • Security: Treat PDFs and filenames as untrusted input. Do not pass untrusted values as the executable path, and avoid exposing passwords in command history or logs.
  • Cost: PDFBox is software you run locally; this workflow has no per-page API charge. Your practical costs are compute, storage, and any separate OCR service you choose for scans.

Or skip the browser setup

If your workflow starts with a web page that you need to capture as an image or PDF before downstream processing, ScreenshotNeo provides a single request instead of maintaining browser automation. It is a screenshot API and MCP server; it does not replace OCR or PDF text extraction.

Use the ScreenshotNeo API documentation for the complete option list and authentication details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners, newsletter popups, and chat widgets are removed before the shot.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify the page verdict and billing status.
  • An MCP server lets AI agents such as Claude or Cursor take screenshots.
  • The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account to try it with 1,000 screenshots a month and no card.

FAQ

Can Get-Content convert a PDF directly?

No. Get-Content reads text files. A PDF parser such as PDFBox must create the text first.

Which PDFBox command should I use?

Use export:text with PDFBox 3.x and ExtractText with PDFBox 2.x.

Will this extract text from a scanned invoice?

Only if the scan already has a text layer. Image-only pages require OCR, which is outside the documented extraction route here.

Can I extract only a few pages?

Yes. PDFBox 3.x documents page-range controls; check the help for your exact release and pass the corresponding options.

Why does extracted text lose columns?

PDFs store positioned text rather than a guaranteed reading order. Sorting can help, but complex layouts may still need document-specific cleanup.