ScreenshotNeo

BlogHTML to image & PDF

Why Are Images Missing From My HTML to PDF Conversion?

Find why images disappear in HTML-to-PDF output and fix URL resolution, access, timing, JavaScript, and print-CSS issues in wkhtmltopdf, WeasyPrint, and Puppeteer.

By the ScreenshotNeo team4 October 20268 min read

Images usually disappear from an HTML-to-PDF conversion because the renderer cannot resolve their URLs, cannot access the files or remote server, or creates the PDF before JavaScript-driven images are ready. Print CSS can also hide an image or omit a CSS background. First identify the renderer and version, then test the exact image URL or path from the same environment and process that creates the PDF.

This guide covers the diagnostic steps and fixes for wkhtmltopdf, WeasyPrint, and Puppeteer. Their defaults and controls differ, so use the instructions for your actual engine and check its version-specific documentation.

1. Identify the renderer and reproduce the failure

Record the renderer and version, how you provide the HTML (URL, file, or string), the command or API call, and the environment where conversion runs. A path that works on a developer’s laptop may not exist inside a container or on a server.

  1. Save or inspect the exact HTML passed to the converter.
  2. List every missing image’s src or CSS url(...) value. Check for empty attributes, typos, malformed URLs, and relative paths.
  3. Determine whether the image is static HTML, inserted by JavaScript, lazy-loaded, or used as a CSS background.
  4. Try one known-good absolute HTTPS image URL and one local image in a minimal document. This helps separate URL, access, timing, and styling problems.
  5. Check the renderer’s logs or diagnostic output for failed resource requests and their URLs.

2. Check how image URLs resolve

A relative URL such as images/logo.png needs a base location. The browser may infer that location from the page you viewed, while a converter given an in-memory HTML string may have no base URL at all. Inspect the final URL the renderer should request, not just the string in the source.

Use an explicit base URL with WeasyPrint

For HTML passed as a string, provide a base URL when the markup contains relative references. The following runnable example uses a local directory as the base; change it to the directory or URL that actually contains your assets.

from weasyprint import HTML

HTML(
    string='''<html><body><img src="images/logo.png"></body></html>''',
    base_url='file:///app/site/'
).write_pdf('output.pdf')

For remote assets, an HTTPS base URL can resolve relative references against a site. WeasyPrint documents how base URLs are determined and the base_url API option in its first steps guide and API reference.

Use absolute paths or URLs where appropriate

  • For remote images, use a complete URL such as https://example.com/assets/photo.png.
  • For local images, verify the path exists inside the converter’s runtime and is readable by its user.
  • Do not assume a browser-relative URL, a working directory, or a developer-machine path is shared by the conversion process.

3. Verify the conversion process can fetch the image

Test access from the same host, container, account, and network context that performs conversion. A remote asset can fail because of DNS, TLS, firewall rules, an HTTP error, a redirect, authentication requirements, or a timeout. A local asset can fail because it is absent, unreadable, or outside the renderer’s allowed file access.

Check the exact response and final destination for remote URLs. If the image requires cookies or authentication, confirm the renderer supports the needed mechanism and that credentials are passed safely. WeasyPrint’s default URL fetching does not support cookies and authentication; its documentation describes custom URL fetchers for integrations that need different behavior. See its API reference.

For local files, grant access only to the required asset directory when the renderer provides access controls. Do not enable broad local-file access for untrusted HTML. WeasyPrint also advises restricting resource access for untrusted HTML and CSS in its security guidance.

4. Wait for JavaScript and lazy-loaded images

If JavaScript adds an image or a page uses lazy loading, the image may not exist or may not have finished loading when PDF generation starts. A generic delay can help diagnose a timing issue, but an application-specific ready condition is more reliable. Network idle alone does not prove that a particular image loaded successfully.

Puppeteer: wait for page activity, then check images

const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch();
  try {
    const page = await browser.newPage();
    await page.goto('https://example.com/report', {
      waitUntil: 'networkidle0',
      timeout: 60000,
    });

    // Optional diagnostic: inspect whether each image loaded.
    const images = await page.evaluate(() =>
      Array.from(document.images, image => ({
        src: image.currentSrc || image.src,
        loaded: image.complete && image.naturalWidth > 0,
      }))
    );
    console.log(images);

    await page.pdf({ path: 'output.pdf', printBackground: true });
  } finally {
    await browser.close();
  }
})();

For a page you control, prefer waiting for a specific completion signal before calling page.pdf(), then inspect any failed image URLs. Puppeteer’s waitForNetworkIdle reference describes the network-idle wait; the PDF options document output settings.

wkhtmltopdf: use its JavaScript wait controls

wkhtmltopdf documents JavaScript enablement, a JavaScript delay, and a window-status wait option. For example, a short delay can help determine whether asynchronous page setup is the cause:

wkhtmltopdf --javascript-delay 2000 https://example.com/report output.pdf

For production, prefer the documented window-status mechanism when your page can set a known status after the images are ready. Check the options available in your installed build. The wkhtmltopdf usage documentation describes these controls.

5. Check renderer settings and print styles

Some renderers print using print media rules, which can differ from the screen view. Inspect @media print rules, visibility, dimensions, clipping, and whether the element is an <img> or a CSS background. Puppeteer’s Page.pdf() uses print CSS media by default.

In Puppeteer, PDF backgrounds are configurable. Set printBackground: true when the missing visual is a CSS background or when backgrounds should be included. This option does not repair a missing or unreachable image URL.

For wkhtmltopdf, image loading is enabled by default in the documented usage, while --no-images disables it. Check that this option has not been set by your command or wrapper. Its documentation also describes local-file-access controls; use the narrowest permissions your assets need. Defaults and option availability can depend on the installed build. See the usage documentation.

6. Renderer-specific checks

Renderer Likely cause What to check
wkhtmltopdf Images disabled, local-file access blocked, or capture happens before JavaScript finishes. Review --images/--no-images, local-file-access controls, JavaScript settings, delay, and window-status options in the installed version’s usage docs.
WeasyPrint Relative URL has no base, fetch fails, or a remote resource requires authentication or cookies. Set base_url for string input, test the exact resource from the conversion environment, and review the URL-fetching and security details in the first steps guide and API reference.
Puppeteer Page not ready, image request failed, print CSS hides it, or the visual is a background omitted from PDF. Wait for an application-specific ready condition, inspect image load status and requests, review print styles, and set PDF background output when needed. See Page, network idle, and PDF options.

7. Troubleshooting common failures

Symptom Likely cause Fix
Relative images disappear when HTML is supplied as a string No valid base URL is available. Set the renderer’s base URL or use a correct absolute URL/path. In WeasyPrint, pass base_url.
Image works in a browser but not in the PDF job The job runs in a different container, account, filesystem, or network. Test the exact path or URL inside the conversion environment and fix its permissions or network access.
Only protected images are missing The request needs authentication, cookies, or headers the renderer does not send. Provide supported credentials through a safe integration, or make the asset reachable to the conversion job. Check renderer-specific fetcher support.
Images inserted by scripts are missing PDF generation starts before page code or image loading completes. Wait for an explicit application-ready condition. Use delay or network-idle controls only as a diagnostic or when appropriate.
CSS background artwork is missing in Puppeteer PDFs Background printing is disabled. Enable printBackground in page.pdf() and verify the CSS background URL itself resolves.
Image is hidden or clipped only in the PDF Print media styles differ from screen styles, or layout clips the image. Inspect @media print, visibility, sizing, overflow, and page layout.
Local files fail after enabling security restrictions The renderer cannot read the asset path. Grant access to the specific required directory using the renderer’s supported controls; avoid unrestricted file access for untrusted input.
Some remote images fail intermittently Remote server, network, redirects, or timeouts are inconsistent. Log the requested URL and response, check redirects and reachability from the job, and set an appropriate timeout where supported.

8. Reliability, performance, and cost considerations

  • Reliability: Make dependencies explicit. Use stable asset URLs or packaged local assets, set a base URL, and wait for a page-specific readiness signal. Log failed resource URLs so a blank spot in a PDF can be traced to a request.
  • Security: HTML and CSS can reference local files or remote resources. Restrict access to trusted paths and networks, especially when converting user-supplied markup. Avoid broad file permissions as a quick fix.
  • Performance: Waiting for every network request to stop can add latency or stall on pages with long-lived requests. A targeted ready condition can be faster and more meaningful. Large images and slow remote hosts also increase conversion time.
  • Cost: The research sources do not establish universal renderer pricing or a cost benchmark. Account for the compute time, network transfer, and any hosted rendering service you use; measure against your own documents and workload.

9. A compact diagnostic checklist

  • [ ] Renderer and version recorded.
  • [ ] Exact HTML and image URL/path inspected.
  • [ ] Relative paths resolved against the intended base.
  • [ ] Image fetched successfully from the conversion runtime.
  • [ ] Local file permissions and renderer access policy checked.
  • [ ] Authentication, redirects, TLS, and timeouts investigated for remote assets.
  • [ ] JavaScript and lazy-loading completion verified.
  • [ ] Print CSS and background output checked.
  • [ ] Renderer logs reviewed for failed resource requests.

Or skip the browser setup

If your goal is a clean screenshot of a web page, ScreenshotNeo is a website screenshot API and MCP server for developers. A single GET request returns an image or PDF. See the API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Responses identify the page verdict and billing status in headers.
  • An MCP server lets AI agents using Claude, Cursor, or another MCP client take screenshots, get page information, and capture PDFs.
  • The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan.

Sign up free for 1,000 screenshots a month, with no card required.

FAQ

Why does a relative image path work in my browser but not in the PDF?

The browser and converter may be using different base locations or running in different environments. Give the renderer a valid base URL and verify the resolved path from the conversion process.

Does waiting for network idle guarantee every image is in the PDF?

No. It indicates a period without network activity, not that every intended image loaded successfully. Inspect image load status and wait for an application-specific ready condition.

Can I safely allow access to local files to fix missing images?

Allow access only to the required asset paths, particularly when input HTML may be untrusted. Broad local-file access can expose files the document should not read.