How to Convert HTML with Images to PDF Using an API
Learn how to convert HTML pages with images into reliable PDFs, handle remote assets, wait for dynamic content, and choose the right API options.
Direct answer: choose an HTML-to-PDF API that accepts your input format (raw HTML, a public URL, or an uploaded file/archive), authenticate, submit rendering options, and save the returned PDF bytes. Images must be reachable by the renderer or included through the provider’s documented upload, archive, or data-URI mechanism. For JavaScript-generated images, configure a wait strategy. Set page size, margins, print CSS, backgrounds, and viewport deliberately, then inspect representative PDFs for missing images and layout changes.
Provider behavior differs. Some APIs accept only one of URL, file, or HTML; others support ZIP archives or reusable assets. Treat the examples below as integration patterns and verify the exact request schema in your selected provider’s current documentation.
1. Choose the input mode
| Input | Use it when | Image considerations |
|---|---|---|
| Raw HTML | Your application creates a document from a template. | Use absolute URLs, inline data where supported, or upload the assets. |
| Public URL | The page is already deployed and reachable from the provider’s network. | Confirm images do not require your browser’s cookies, VPN, or local filesystem. |
| File or archive | The document uses private or local assets and the provider supports uploads. | Preserve relative paths and include every referenced image and stylesheet. |
HTMLPDF documents mutually exclusive URL, file, and HTML inputs, along with image loading, viewport, and print-media controls. Adobe PDF Services documents conversion from static or dynamic HTML, ZIP, and URL inputs. These are provider-specific capabilities, not a universal API contract.
2. Make every image resolvable
A browser rendering your page successfully does not prove that a remote conversion service can fetch its images. For each <img src>, CSS background, and font:
- Prefer an absolute HTTPS URL when the asset can be public.
- Check that the URL returns an image, not an HTML login page or redirect.
- Confirm the certificate is valid and the host is reachable from the provider’s network.
- Do not rely on your browser’s cookies, local paths, VPN, or service-worker cache.
- Use the provider’s documented upload or archive feature for private assets.
- Use data URIs only when the provider supports them and the resulting request stays within its size limits.
Remember that an <img> and a CSS background are separate cases. An API may load images while omitting background graphics unless a print-background option is enabled.
Minimal HTML with an external image
<!doctype html>
<html>
<head>
<meta charset="utf-8">
<style>
@page { size: A4; margin: 18mm; }
body { font-family: sans-serif; }
img { max-width: 100%; height: auto; }
</style>
</head>
<body>
<h1>Quarterly report</h1>
<img src="https://example.com/assets/chart.png" alt="Quarterly chart">
</body>
</html>
3. Render dynamic HTML before creating the PDF
If JavaScript inserts an image or fetches data after the initial response, the converter needs a documented JavaScript engine and a suitable wait rule. Common choices are:
- Selector wait: continue when a known chart or image element exists.
- Network idle: continue after network activity quiets down; useful for client-rendered pages, but analytics or long polling can prevent idle.
- Fixed delay: simple and predictable, but slower and still vulnerable to unusually slow requests.
PDFSpark documents JavaScript rendering and a network-idle example. HTMLPDF documents JavaScript and a configurable delay. Confirm the option names and timeout limits for your provider.
Self-hosted reference with Playwright (Node.js)
import { chromium } from 'playwright';
const browser = await chromium.launch();
const page = await browser.newPage({
viewport: { width: 1280, height: 900 },
deviceScaleFactor: 1
});
await page.goto('https://example.com/report', { waitUntil: 'networkidle' });
await page.waitForSelector('img.chart');
await page.emulateMedia({ media: 'print' });
await page.pdf({
path: 'report.pdf',
format: 'A4',
printBackground: true,
margin: { top: '18mm', right: '18mm', bottom: '18mm', left: '18mm' }
});
await browser.close();
This illustrates the workflow when you control the browser. A hosted API may expose equivalent settings under different names.
4. Set page and print options intentionally
| Option | Why it matters | Typical checks |
|---|---|---|
| Paper size or custom dimensions | Controls pagination and available line width. | A4 versus Letter, or a receipt-sized custom page. |
| Margins | Prevents content and headers from colliding with page edges. | Measure long headings, tables, and footer space. |
| Orientation | Landscape can prevent wide tables from wrapping. | Check charts and table columns. |
| Print media | Chooses print-specific CSS instead of screen CSS. | Verify visibility rules and print-only layout. |
| Backgrounds | Controls CSS backgrounds and colored sections. | Compare branding blocks and charts with the browser. |
| Viewport | Affects responsive breakpoints before pagination. | Set it explicitly for predictable wrapping. |
| Wait time | Allows fonts, images, and charts to finish loading. | Use a selector or network rule where possible. |
| Headers and footers | Adds document metadata and page numbers. | Keep enough top and bottom margin; avoid overlap. |
PDF.co documents print media, background, page, margin, header, and footer controls, including page-number variables. Adobe’s example includes page layout and a wait setting. Defaults differ, so make the settings that affect your layout explicit.
5. Send the request and validate the response
Authentication and request encoding are provider-specific. A robust client should:
- Build either HTML, URL, or an upload according to the provider contract.
- Send credentials through the documented header or credential field.
- Set a request timeout long enough for browser rendering.
- Check the HTTP status and content type before writing bytes as a PDF.
- Record the provider’s error body when the response is not a PDF.
- Save or stream the PDF without converting binary data to text.
# Generic request shape (pseudocode)
input = html_or_url
options = {
"page_format": "A4",
"margins": "18mm",
"print_background": true,
"wait_for": "img.chart"
}
response = POST provider_endpoint with credentials, input, options
if response is a successful PDF:
save(response.body, "result.pdf")
else:
log(response.status, response.body)
6. Inspect output with a repeatable checklist
- Every image appears, at the expected resolution and aspect ratio.
- Lazy-loaded images and charts are present.
- Images do not push headings or captions onto unexpected pages.
- Tables do not clip at the right edge.
- Print backgrounds, borders, and colors are intentional.
- Fonts are available or have an acceptable fallback.
- Headers and footers do not overlap body content.
- Links, outlines, and page ranges behave as expected.
- Protected images are handled through the provider’s supported authentication or upload mechanism.
7. Or skip the browser setup
ScreenshotNeo can capture a reachable page as a PDF with one GET request. It accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its PDF options include paper size, margins, landscape mode, and page ranges.
See the ScreenshotNeo API documentation for the current parameter list.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com/report \
-d format=pdf \
-d paper_size=A4 \
-d print_background=true \
-o report.pdf
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={
"access_key": "YOUR_API_KEY",
"url": "https://example.com/report",
"format": "pdf",
"paper_size": "A4",
"print_background": "true",
},
timeout=90,
)
r.raise_for_status()
content_type = r.headers.get("content-type", "")
if "pdf" not in content_type:
raise RuntimeError(f"Expected PDF, got {content_type}")
open("report.pdf", "wb").write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com/report',
format: 'pdf',
paper_size: 'A4',
print_background: 'true'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned ${res.status}`);
const type = res.headers.get('content-type') || '';
if (!type.includes('pdf')) throw new Error(`Expected PDF, got ${type}`);
const fs = await import('node:fs/promises');
await fs.writeFile('report.pdf', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo also supports full-page capture with lazy images loaded, custom CSS and JavaScript, click actions, selector waits, delays, network-idle waits, request blocking, custom headers and cookies, timezone and geolocation, caching with a chosen TTL, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, and an MCP server with take_screenshot, get_page_info, and capture_pdf tools for AI agents.
Free accounts include 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account.
8. Troubleshooting missing images and layout errors
| Symptom | Likely cause | Fix |
|---|---|---|
| Image area is blank | Relative URL, blocked host, authentication, or failed redirect. | Use an absolute URL, upload the asset, or provide documented credentials. |
| Only some images appear | Lazy loading or a slow origin. | Wait for a selector, network idle, or a documented delay; inspect image responses. |
| CSS background is missing | Print backgrounds are disabled. | Enable the provider’s background option and test print CSS. |
| Charts are empty | JavaScript has not finished. | Enable JavaScript and wait for a chart selector or stable network state. |
| PDF is actually JSON or HTML | Authentication or validation failed. | Check status and content type; log the error body before saving. |
| Text wraps differently | Viewport, print media, fonts, or paper size differ. | Set viewport, media, fonts, page size, and margins explicitly. |
| Content overlaps footer | Margins are too small for header/footer templates. | Increase margins and reserve space for the templates. |
| Request times out | Slow resources, long polling, or an overly strict timeout. | Block unnecessary requests, use selector waits, and tune timeout limits. |
9. Performance, reliability, and cost
- Performance: smaller images, fewer third-party scripts, explicit waits, and request blocking reduce render time. Network-idle waits can be slow on pages with analytics or long polling.
- Reliability: retry transient transport failures with backoff, but avoid blindly retrying invalid HTML or authentication errors. Store a request identifier and the input URL with each output.
- Idempotency: cache stable documents when the provider supports a TTL. For changing pages, include a version or timestamp in your cache key.
- Cost: compare billing units, asset upload limits, async-job fees, and cache behavior in the provider’s current pricing. The research does not establish comparable prices or benchmarks across vendors.
- Quality: maintain fixtures containing remote images, lazy images, CSS backgrounds, web fonts, tables, and JavaScript charts. Review PDFs after template or dependency changes.
10. FAQ
Can an API convert a web page URL to PDF?
Yes, when the provider supports URL input and its renderer can reach the page and its resources. Private pages usually require a documented upload or authentication mechanism.
Should I inline every image as a data URI?
Only when the selected API documents data-URI support and the request remains within its size limits. Uploading an archive or using reachable URLs is often easier to maintain.
Why does the browser show an image that the PDF does not?
The converter may lack your cookies, network access, JavaScript wait time, or permission to fetch the asset. Inspect the image URL from the converter’s environment and configure the relevant option.
How do I preserve CSS backgrounds?
Enable print-background behavior and verify that your print stylesheet does not hide the element. Background loading and ordinary image loading can be separate settings.
Is a screenshot API suitable for multi-page documents?
Use a provider that supports PDF output, full-page capture, paper settings, margins, and page ranges. Validate page breaks and image resolution with representative documents.


