ScreenshotNeo

BlogHTML to image & PDF

How to Convert HTML Files to PDF While Preserving Local Images

Convert an HTML file to PDF without losing local images. Learn how paths, file access, Chrome, WeasyPrint and wkhtmltopdf affect the result.

By the ScreenshotNeo team4 October 20268 min read

To convert an HTML file to PDF and preserve its local images, use a renderer that can read both the HTML and every image it references. Make relative paths resolve from the right directory, allow local-file access where required, and inspect the finished PDF. Saving as PDF does not repair a broken path or guarantee that delayed, script-generated, or inaccessible images loaded.

For a quick manual conversion, open the file in a browser and choose Print → Save to PDF. For repeatable conversion, use Chrome Headless, WeasyPrint, or wkhtmltopdf. The choice depends on whether the page needs browser JavaScript, how it references local assets, and what rendering engine your output requires.

1. Check the HTML and image paths first

A renderer can include an image only if it can find and load it. Suppose your files are arranged like this:

project/
├── page.html
└── images/
    └── photo.jpg

This markup refers to the image relative to the HTML document’s base URL:

<img src="images/photo.jpg" alt="A sample photo">

Before converting:

  1. Open page.html in the browser or environment you plan to use.
  2. Confirm that each image appears. Check CSS backgrounds as well as <img> elements.
  3. Identify the base for relative paths. A <base href="..."> element can change how relative links resolve.
  4. Check capitalization and spelling. Some operating systems distinguish Photo.jpg from photo.jpg.
  5. Check that the conversion process has permission to read the HTML, images, stylesheets, and fonts.

Paths such as C:\images\photo.jpg or /home/me/images/photo.jpg may work only on the machine where they were authored. Relative paths are usually easier to move with a project, provided the converter resolves them from the intended directory.

2. Print to PDF from a browser

For a one-off conversion, open the HTML in a browser that can access its local assets, then choose Print and select Save to PDF. This is often the simplest choice when the page depends on browser JavaScript or modern CSS. In the print dialog, review paper size, orientation, margins, scale, and background graphics; these settings affect layout and whether background images appear.

For scripting, Chrome Headless documents --print-to-pdf and a timeout option. The following command uses an absolute file URL; replace the path with the full path to your HTML file:

chrome --headless --disable-gpu \
  --no-pdf-header-footer \
  --print-to-pdf="output.pdf" \
  "file:///absolute/path/to/project/page.html"

On systems where the executable is named google-chrome or chromium, substitute that command. A timeout can give a page more time before capture:

chrome --headless --no-pdf-header-footer --timeout=5000 \
  --print-to-pdf="output.pdf" \
  "file:///absolute/path/to/project/page.html"

Chrome’s documentation shows URL targets. Confirm that your installed Chrome build handles the local file:// URL and its sibling assets as expected. A timeout only limits waiting; it cannot fix incorrect paths, blocked file access, or an image that never loads. See the Chrome Headless CLI documentation.

3. Convert with WeasyPrint in Python

WeasyPrint accepts a local filename and can resolve relative URLs using a base URL. Install it in an environment following its official installation instructions, then save this as convert.py:

from pathlib import Path
from weasyprint import HTML

source = Path("project/page.html").resolve()
output = Path("output.pdf").resolve()

HTML(
    filename=str(source),
    base_url=str(source.parent),
).write_pdf(str(output))

print(f"Wrote {output}")

Run it with:

python convert.py

When the HTML is loaded by filename, WeasyPrint can infer a base from that filename. Setting base_url explicitly makes the intended directory clear, especially if you later switch to passing HTML as a string. For <img src="images/photo.jpg">, the base should be the directory containing the images folder. A document’s own <base> element can also affect relative URL resolution. Consult the WeasyPrint API reference for URL resolution, supported image formats, and version-specific behavior.

To pass HTML directly as a string, specify the base explicitly:

from pathlib import Path
from weasyprint import HTML

html_text = Path("project/page.html").read_text(encoding="utf-8")
base = Path("project").resolve()
HTML(string=html_text, base_url=str(base)).write_pdf("output.pdf")

Without a base URL, relative resources in string input may not resolve. WeasyPrint is a dedicated HTML/CSS-to-PDF renderer; if your page depends on browser-only JavaScript to create images or markup, use a browser-based workflow or prepare those resources before conversion.

4. Convert with wkhtmltopdf

wkhtmltopdf’s usage documentation lists image loading as enabled by default, while local-file access is disabled by default unless enabled or the necessary files are allowed. For a project whose assets are in /absolute/path/project, a command can look like this:

wkhtmltopdf \
  --enable-local-file-access \
  /absolute/path/project/page.html \
  /absolute/path/output.pdf

Where supported by your installed version, restrict access to the project directory instead of allowing broader local-file access:

wkhtmltopdf \
  --allow /absolute/path/project \
  /absolute/path/project/page.html \
  /absolute/path/output.pdf

Use --no-images only when you intentionally want images omitted. For diagnosis, consult the installed version’s help for its load-error and media-error handling options. wkhtmltopdf is based on Qt WebKit; rendering may differ from a current browser. Check the project’s usage documentation and official homepage for the version you run.

5. Choose a conversion method

Method Good fit Check before relying on it
Browser print dialog One-off conversion; page uses browser features Print settings, local asset access, and whether the page has finished rendering
Chrome Headless Automated conversion with browser rendering Local file:// behavior, timing, and Chrome version
WeasyPrint Python workflows and explicit control over relative URL resolution CSS and script compatibility; set the base URL when needed
wkhtmltopdf Existing workflows built around this renderer Local-file access settings and differences from modern browser rendering

There is no universal guarantee that all HTML, CSS, and image paths will render identically across these tools. Choose based on the features your document uses, then check representative output from the exact renderer and version you will deploy.

6. Verify the PDF

After conversion, open the PDF and inspect every page that contains important imagery. Confirm that:

  • Images are present, legible, and not unexpectedly cropped or stretched.
  • CSS background images appear if they are required.
  • Large or late-loading images are not blank.
  • Page breaks, margins, and scaling have not cut off content.
  • The PDF opens in more than one viewer if it is being distributed broadly.

For repeat jobs, make this inspection part of a review process. A successful command exit or a generated PDF file does not prove every resource was embedded correctly.

7. Common problems and fixes

Symptom Likely cause What to try
An image is missing in the PDF The path is wrong, relative to an unexpected base, or inaccessible Open the HTML first; check src, CSS url(...), <base>, spelling, and file permissions. Set WeasyPrint’s base_url or allow the required files in wkhtmltopdf.
The HTML displays images, but conversion does not The converter runs from a different working directory or has a different file-access policy Use absolute paths for diagnosis, verify the process user can read the assets, and check the converter’s local-resource settings.
Only CSS background images are missing Print styling or browser print settings omit backgrounds Enable background graphics in the print dialog; check print media CSS and the renderer’s support for the relevant styles.
Images generated by JavaScript are blank The capture occurred before the script finished, or the chosen renderer does not run the required script Use a browser workflow, wait for the content to appear, or save the generated HTML and image resources before converting.
Chrome produces a PDF before images appear The page needs more time, or the resource is delayed or blocked Try a longer documented timeout and check the resource path and network or file access. Waiting cannot repair an invalid URL.
wkhtmltopdf reports a load or access error Local-file access is disabled, the path is outside an allowed directory, or a resource failed to load Use --allow for the required asset directory or enable local-file access when appropriate; review version-specific error options.
WeasyPrint cannot resolve a relative image HTML was provided as a string without a base, or the base points to the wrong directory Pass base_url for the directory against which relative paths should resolve.
Images look blurry or oversized The source resolution, CSS dimensions, or PDF scaling is unsuitable Check the source file and its rendered dimensions; adjust print scale or image sizing, then inspect the output again.

8. Performance, reliability, and cost

For an individual file, the main practical cost is setup and review time. For automated batches, conversion time and memory depend on the document, image sizes, renderer, and environment; the cited documentation does not establish a universal speed or memory comparison. Large images and complex page layouts can make output heavier and conversion slower.

  • Keep assets together: package the HTML and its referenced files, and preserve their relative directory structure.
  • Use stable inputs: avoid changing asset files while a conversion is running.
  • Control access: grant the renderer access only to the files it needs where practical.
  • Pin and review versions: renderer options and output can change; verify the locally installed version before relying on flags.
  • Check failures explicitly: log conversion errors and review generated PDFs, especially for recurring or high-volume jobs.
  • Do not assume remote assets are local: an online image requires network access and may be unavailable in a restricted environment.

There is no named image-preservation rate or cross-tool benchmark in the cited documentation. Treat the workflow as reliable only after validating your own HTML, assets, renderer, and version.

9. Or skip the browser setup

If the page is available at a public URL, ScreenshotNeo can return a screenshot or PDF from one API request. This is for capturing a website URL; it does not upload or convert a local HTML file and its private local image paths. See the ScreenshotNeo API docs for parameters and formats.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -d format=pdf \
  -o page.pdf

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, no card required.

Frequently asked questions

Will converting HTML to PDF embed the local image files?

It can, if the renderer can resolve and load them. Verify the resulting PDF rather than relying on the file extension or a successful command alone.

Should image paths be absolute or relative?

Either can work when accessible. Relative paths travel more easily with a project, but depend on the correct document base; absolute paths are useful for diagnosis but tie the HTML to a machine’s directory layout.

Can I convert a local HTML file through ScreenshotNeo?

The ScreenshotNeo call shown here captures a website at a URL. A local HTML file and local image folder are not a public URL; use a local renderer for that workflow.

Which method should I use if I need browser JavaScript?

Start with a browser-based workflow such as Chrome, then confirm scripts and images have finished loading before the PDF is created.