ScreenshotNeo

BlogHow-to

How to Convert PDF to Word with PowerShell

Automate PDF-to-DOCX conversion with Microsoft Word COM automation, batch processing, quality checks, and fixes for layout and OCR problems.

By the ScreenshotNeo team1 October 20267 min read

Yes, you can convert a PDF to an editable Word document with PowerShell by automating Microsoft Word’s built-in PDF import and DOCX save operation. This requires Windows and desktop Microsoft Word. Word converts a copy of the PDF, so the original file remains unchanged. The method works best when the PDF is mostly text; scanned, image-heavy, and structurally complex files need OCR or manual review.

What you need

  • Windows PowerShell or PowerShell 7 running on Windows.
  • Desktop Microsoft Word installed and activated on the machine that runs the script.
  • Read access to the source PDF and write access to the destination folder.

Word’s documented behavior is to open the PDF, convert its contents into a format Word can display, and save that content as a Word document. Microsoft says this works best for PDFs that are mostly text. See Microsoft’s PDF support guidance.

Convert one PDF to DOCX

Save this as Convert-PdfToWord.ps1:

param(
    [Parameter(Mandatory = $true)]
    [string]$InputPdf,

    [Parameter(Mandatory = $true)]
    [string]$OutputDocx
)

Set-StrictMode -Version Latest
$ErrorActionPreference = 'Stop'

$inputPath = [System.IO.Path]::GetFullPath((Resolve-Path -LiteralPath $InputPdf).Path)
$outputPath = [System.IO.Path]::GetFullPath($OutputDocx)
$outputDirectory = Split-Path -Parent $outputPath

if (-not (Test-Path -LiteralPath $outputDirectory)) {
    New-Item -ItemType Directory -Path $outputDirectory -Force | Out-Null
}

$word = $null
$document = $null

try {
    $word = New-Object -ComObject Word.Application
    $word.Visible = $false
    $word.DisplayAlerts = 0

    # Open the PDF read-only. Word converts a working copy internally.
    $document = $word.Documents.Open($inputPath, $false, $true)

    # 16 is wdFormatDocumentDefault, the DOCX format.
    $document.SaveAs2($outputPath, 16)
    Write-Host "Created $outputPath"
}
finally {
    if ($document -ne $null) {
        $document.Close($false)
        [void][System.Runtime.InteropServices.Marshal]::ReleaseComObject($document)
    }
    if ($word -ne $null) {
        $word.Quit()
        [void][System.Runtime.InteropServices.Marshal]::ReleaseComObject($word)
    }
    [GC]::Collect()
    [GC]::WaitForPendingFinalizers()
}

Run it from PowerShell:

Set-ExecutionPolicy -Scope Process Bypass
.\Convert-PdfToWord.ps1 -InputPdf .\invoice.pdf -OutputDocx .\output\invoice.docx

The execution-policy setting applies only to the current PowerShell process. Use your organization’s policy if scripts are centrally controlled.

Batch-convert every PDF in a folder

This script creates a matching .docx beside each PDF, or in a separate output directory:

param(
    [Parameter(Mandatory = $true)] [string]$InputDirectory,
    [Parameter(Mandatory = $true)] [string]$OutputDirectory
)

Set-StrictMode -Version Latest
$ErrorActionPreference = 'Stop'

$inputRoot = (Resolve-Path -LiteralPath $InputDirectory).Path
if (-not (Test-Path -LiteralPath $OutputDirectory)) {
    New-Item -ItemType Directory -Path $OutputDirectory -Force | Out-Null
}
$outputRoot = [System.IO.Path]::GetFullPath($OutputDirectory)

$word = New-Object -ComObject Word.Application
$word.Visible = $false
$word.DisplayAlerts = 0

try {
    Get-ChildItem -LiteralPath $inputRoot -Filter *.pdf -File | ForEach-Object {
        $document = $null
        try {
            $destination = Join-Path $outputRoot ($_.BaseName + '.docx')
            $document = $word.Documents.Open($_.FullName, $false, $true)
            $document.SaveAs2($destination, 16)
            Write-Host "Converted $($_.Name) -> $(Split-Path $destination -Leaf)"
        }
        catch {
            Write-Warning "Failed to convert $($_.FullName): $($_.Exception.Message)"
        }
        finally {
            if ($document -ne $null) {
                $document.Close($false)
                [void][System.Runtime.InteropServices.Marshal]::ReleaseComObject($document)
            }
        }
    }
}
finally {
    $word.Quit()
    [void][System.Runtime.InteropServices.Marshal]::ReleaseComObject($word)
    [GC]::Collect()
    [GC]::WaitForPendingFinalizers()
}

Run it with:

.\Convert-Folder.ps1 -InputDirectory .\pdfs -OutputDirectory .\docx

For nested folders, replace Get-ChildItem with Get-ChildItem -Recurse and recreate the relative directory structure before saving each result.

What the conversion preserves well

Text-heavy business, legal, and scientific PDFs generally produce the most useful editable documents. Headings, paragraphs, ordinary lists, and many inline images often import acceptably.

Word warns that the result may not match the PDF page for page. Line wrapping, page breaks, columns, and spacing can change because a PDF stores positioned output while DOCX stores editable document structure.

What commonly changes or fails

  • Scanned PDFs: Pages may become images with no editable text. Run OCR first or use a converter that includes OCR.
  • Charts and graphics: A graphics-heavy page may be represented as one image, so its text cannot be edited.
  • Tables: Cell spacing, merged cells, borders, and unusual layouts can shift.
  • Columns and text boxes: Reading order and alignment may change.
  • Footnotes and endnotes: Multi-page notes can move or lose their original placement.
  • Tracked changes, comments, bookmarks, and tags: These PDF structures may not map cleanly to Word features.
  • Fonts and effects: Missing fonts, special effects, transparency, and unusual glyphs can alter appearance.
  • Active content: Audio, video, forms, scripts, and other interactive PDF elements are not equivalent to editable DOCX content.

Keep the original PDF as the authoritative visual copy and inspect the DOCX before distributing it.

Quality-control checklist

  1. Open the DOCX in Word and confirm that every page is present.
  2. Search for several phrases from the PDF to verify text extraction.
  3. Check headings, columns, tables, page breaks, footnotes, images, and hyperlinks.
  4. Compare totals, dates, names, and other high-risk values against the PDF.
  5. Review accessibility: heading structure, reading order, table headers, and alternative text.
  6. Check fonts and printing at the required paper size.
  7. Retain the source PDF alongside the converted document.

PowerShell and Word options that matter

Option Purpose
$word.Visible = $false Runs Word without displaying its window.
$word.DisplayAlerts = 0 Prevents modal prompts from stopping unattended runs.
Documents.Open(path, $false, $true) Opens the PDF without conversion prompts and read-only.
SaveAs2(path, 16) Saves using Word’s DOCX format.
Get-ChildItem -Filter *.pdf -File Enumerates only PDF files in a batch.

Always close each document and release COM objects. Otherwise, hidden WINWORD.EXE processes can accumulate and lock files.

Troubleshooting

New-Object cannot create the Word COM object

Cause: Desktop Word is not installed, COM registration is damaged, or the process architecture and Office installation are incompatible. Fix: Confirm that Word opens interactively on the same machine, repair Office, and run the script under the same user account that has Word installed.

The script says the PDF cannot be found

Cause: The path is relative to a different working directory or contains special characters. Fix: Pass a quoted absolute path, for example -InputPdf 'C:\Jobs\April report.pdf', and verify it with Test-Path -LiteralPath.

Access is denied when saving

Cause: The destination is read-only, the DOCX is open, or another process has a lock. Fix: Choose a writable directory, close the existing document, and retry with a new filename.

Word remains running after the script exits

Cause: A document or another COM child object still holds a reference. Fix: Close the document in finally, release every object you create, then force garbage collection as shown above. Terminate leftover Word processes only after checking that no user’s Word session is being used.

The DOCX is mostly pictures

Cause: The PDF is scanned or graphics-heavy. Fix: Run OCR before conversion or use a PDF converter with OCR. Word’s import is not a substitute for OCR.

Tables or page breaks look wrong

Cause: PDF coordinates do not encode the same layout model as DOCX. Fix: Rebuild complex tables, adjust section and column settings, and use the quality-control checklist. Preserve the PDF for exact visual reproduction.

Performance, reliability, and cost considerations

  • Start one Word COM instance for a batch and close each document promptly.
  • Process files sequentially unless you have isolated Windows workers; multiple Word instances increase resource use and can create file locks.
  • Write outputs to a local writable directory, then copy completed files to network storage.
  • Log the source path, destination path, timestamp, and exception message for each file.
  • Use a retry only for transient file-lock or startup failures. Do not repeatedly retry a malformed PDF without recording the error.
  • Desktop Word requires licensing and an interactive Windows installation. It is a poor fit for Linux containers or serverless workers.

Or skip the browser setup

ScreenshotNeo is not a PDF-to-DOCX converter. If your workflow also needs a clean visual capture of a web page or PDF URL, its API can return a screenshot or PDF without installing browser automation. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots, and 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000. See the ScreenshotNeo documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/document.pdf -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/document.pdf"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/document.pdf' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Learn more at ScreenshotNeo, then sign up free for 1,000 screenshots each month with no card.

Alternatives for unattended workflows

When desktop Word is unavailable, a connector or hosted conversion service may fit better. Microsoft Learn documents the PDF4me Convert connector, which accepts file content and a file name, supports quality and language parameters, and returns converted Word file content. Adobe Acrobat and other third-party PDF converters are additional options listed by Microsoft.

FAQ

Does the script modify the PDF?

No. It opens the PDF read-only and writes a separate DOCX.

Can this convert password-protected PDFs?

Only if Word can open them with the required password or access rights. Supply credentials through an approved process; do not embed passwords in scripts or logs.

Can I run this on Linux?

Not with Word COM automation. Use a Windows worker with desktop Word or a service/API designed for unattended conversion.

Will the DOCX look exactly like the PDF?

No guarantee exists. The conversion targets editable Word content, so pagination and layout can differ.

Is OCR included?

No. Scanned documents generally require an OCR step or an OCR-capable converter before you can edit the text.