ScreenshotNeo

BlogHow-to

How to Convert HTML to DOCX

Convert HTML to DOCX with Pandoc, reference styles, automation examples, fidelity limits, troubleshooting, and a clean capture option.

By the ScreenshotNeo team1 October 20266 min read

Short answer: use Pandoc to translate an HTML file into a Word document:

pandoc -f html input.html -o output.docx

Pandoc also infers formats from file extensions:

pandoc input.html -o output.docx

The conversion preserves document structure, but it is not a pixel-perfect browser print. Review the resulting DOCX, especially tables, page breaks, fonts, and CSS-driven layout. For repeatable Word styling, provide a reference DOCX.

1. Convert a local HTML file

Save the page as input.html, open a terminal in that directory, and run:

pandoc -f html input.html -o output.docx

Use explicit paths in automation:

pandoc -f html ./html/input.html -o ./docx/output.docx

Open output.docx in Word or another DOCX reader and inspect headings, lists, links, images, tables, and page layout.

2. Use a reference DOCX for Word styles

A reference document supplies styles and document properties such as margins, page size, headers, and footers:

pandoc -f html input.html -o output.docx --reference-doc=styles.docx

This is useful when generated documents must match an existing Word template. Start from Pandoc’s default reference document and edit the styles and page properties you need. See the Pandoc User’s Guide for reference-document details.

3. Prepare HTML that converts predictably

Use semantic elements

Prefer h1–h6, p, ul, ol, li, blockquote, pre, code, table, thead, tbody, tr, th, and td. Semantic structure maps more reliably to Word paragraphs, headings, lists, and tables than deeply nested layout containers.

Keep important content in the HTML

Browser-only effects are not a DOCX content model. Content inserted after page load by JavaScript, text painted into a canvas, and layout created only by CSS may need to be rendered or rewritten as ordinary HTML before conversion. Use real text and ordinary image files for anything that must appear in Word.

Make assets available

Use local, readable image paths or URLs available to the conversion process. In CI, copy assets into the build workspace and use stable relative paths. A missing image can leave an empty area in the DOCX.

Keep tables simple

Simple rectangular tables convert best. Complex nested tables, merged cells, and layout tables can lose formatting because Pandoc translates through an intermediate representation that cannot express every source-format detail. Flatten complex tables or move explanatory text outside them.

4. A repeatable conversion script

This script creates the output directory and applies a reference template:

#!/usr/bin/env bash
set -euo pipefail

input="${1:-input.html}"
output="${2:-build/output.docx}"
reference="${3:-styles.docx}"

mkdir -p "$(dirname "$output")"
pandoc -f html "$input" -o "$output" --reference-doc="$reference"
echo "Wrote $output"

Run it with:

chmod +x convert.sh
./convert.sh report.html build/report.docx company-styles.docx

For a one-off conversion without a template, omit --reference-doc.

5. Python automation

Call Pandoc as a subprocess when your Python application owns the input and output paths:

from pathlib import Path
import subprocess

source = Path("input.html")
target = Path("output.docx")
reference = Path("styles.docx")

target.parent.mkdir(parents=True, exist_ok=True)
subprocess.run([
    "pandoc", "-f", "html", str(source),
    "-o", str(target),
    "--reference-doc", str(reference),
], check=True)
print(f"Wrote {target}")

Pass a list of arguments instead of one shell string so paths containing spaces remain safe. The python-docx library creates and updates DOCX files, but it is not itself a turnkey HTML converter. An HTML-to-DOCX-specific project such as html2docx is a separate option when you need an in-process Python workflow; evaluate its supported HTML and CSS subset before adopting it.

6. Browser-based conversion

A Pandoc browser application lists HTML and DOCX support and includes a reference-document upload control. It can be convenient for a one-off conversion, but check its current data-handling terms before uploading confidential or regulated content; the research for this guide did not verify those terms. For private or repeatable jobs, local Pandoc keeps the conversion in your environment.

7. What converts well, and what needs review

HTML feature Typical result Review action
Headings and paragraphs Word headings and paragraphs Check heading levels and spacing.
Ordered and unordered lists Word lists Check nesting and numbering restarts.
Links Clickable hyperlinks Open several links after conversion.
Images Figures when assets are available Check resolution, sizing, and missing files.
Simple tables Word tables Check column widths and wrapping.
Advanced CSS layout Partial or changed layout Rewrite as semantic structure or accept a manual pass.
JavaScript-generated content May be absent Materialize the content into HTML first.

Pandoc aims to preserve structure through an intermediate representation. Because that representation is less expressive than many input and output formats, exact visual equivalence is not guaranteed. Inspect tables, page breaks, margins, headers, footers, and custom styling before distributing the DOCX.

8. Troubleshooting

“pandoc: command not found”

Pandoc is not installed or is not on the process PATH. Install Pandoc from its official distribution, reopen the terminal or restart the service so PATH is refreshed, then run pandoc --version.

The output is blank or missing recent page content

The HTML may depend on JavaScript, a login session, or an API call that never ran in the conversion process. Save a fully rendered HTML snapshot first, or replace dynamic sections with ordinary HTML before invoking Pandoc.

Images are missing

Check every src path, run from the expected working directory, and make sure the process can read local files or reach remote URLs. Prefer copied local assets in CI.

Styles look wrong

Browser CSS does not map one-to-one to Word styles. Use semantic HTML and a reference DOCX. Set margins, page size, and Word styles in that reference document instead of relying on complex CSS.

Tables overflow or lose merged cells

Simplify the table, avoid layout tables, and review merged cells manually. Split a very wide table or move long prose into paragraphs.

Fonts differ between machines

The DOCX may reference fonts unavailable to the reader. Choose common fonts in the reference document or install required fonts wherever the document is rendered.

The command succeeds but the DOCX is still wrong

A zero exit status means Pandoc produced a file, not that every visual detail matches the browser. Add a review step that opens the DOCX and checks representative pages, tables, and figures.

9. Performance, reliability, and cost

  • Performance: keep assets local and avoid unnecessarily large images. Process files in batches with a script while preserving each command’s exit status.
  • Reliability: pin the Pandoc version in build images, keep the input HTML and reference DOCX together, and archive conversion logs and source files.
  • Cost: local Pandoc is software you run yourself, so practical costs are compute, storage, and maintenance. Verify any online service’s privacy and pricing terms before sending sensitive HTML.
  • Quality control: compare representative output after template or Pandoc upgrades. No relevant published accuracy benchmark was found in the reviewed sources.

10. Or skip the browser setup

If your HTML lives at a public URL and you need a clean visual capture for review, archiving, or a later document workflow, ScreenshotNeo returns an image or PDF with one GET request. It does not replace Pandoc’s HTML-to-DOCX translation; it removes the browser-capture setup step.

See the ScreenshotNeo API docs for all options:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Before capture, ScreenshotNeo accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

There are 1,000 screenshots per month free with no card. Paid plans start at $5 for 3,000 shots, and every feature is included on every plan. Create a free ScreenshotNeo account.

11. FAQ

Can Pandoc preserve the exact browser appearance?

No. It preserves structure through an intermediate representation, so advanced CSS and complex tables can change. Use a reference DOCX and inspect the output.

Do I need a reference DOCX?

No for a basic conversion. Use one when you need consistent Word styles, margins, page size, headers, or footers.

Is python-docx an HTML converter?

It is a library for creating and updating DOCX files. Use Pandoc or an HTML-to-DOCX-specific library for conversion.

Should confidential HTML go through an online converter?

Prefer local Pandoc unless you have verified the service’s current data-handling terms and your organization allows the upload.

Can ScreenshotNeo produce a DOCX?

ScreenshotNeo produces PNG, JPEG, WebP, or PDF captures. Use Pandoc when the required output is an editable DOCX.