How to Load External JavaScript When Converting HTML to PDF in Node.js
Render HTML in Chromium, load the external script, wait for a page-specific ready signal, then generate the PDF with Puppeteer or Playwright.
Use a real Chromium page. Navigate to the HTML with Puppeteer or Playwright, ensure the external JavaScript file loads, wait for the application to finish rendering, and only then call the PDF API. Loading the script and finishing application rendering are separate events, so a fixed delay or networkidle alone is often insufficient.
1. Complete Puppeteer example
This example loads a document that already references an external script, waits for a page-specific readiness flag, and writes a PDF.
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch();
const page = await browser.newPage();
page.on('console', message => {
console.log('[browser]', message.type(), message.text());
});
page.on('pageerror', error => {
console.error('[page error]', error);
});
page.on('requestfailed', request => {
console.error('[request failed]', request.url(), request.failure()?.errorText);
});
await page.goto('https://example.com/report.html', {
waitUntil: 'networkidle2'
});
// Use this only when report.html does not already include the dependency.
// await page.addScriptTag({ url: 'https://cdn.example.com/report.js' });
await page.waitForFunction(() => window.reportReady === true);
// PDF uses print media by default. Use screen media when the CSS was designed
// for the browser viewport instead of print styles.
// await page.emulateMediaType('screen');
await page.pdf({
path: 'report.pdf',
format: 'A4',
printBackground: true
});
await browser.close();
See the Puppeteer PDF generation guide, addScriptTag API, and page.pdf API for the documented methods.
When the HTML contains the script already
<script src='https://cdn.example.com/report.js'></script>
Let the document load its own dependency. Calling addScriptTag as well can execute the library twice and produce duplicate event handlers or duplicated output.
When you must inject the script
await page.goto('file:///absolute/path/report.html', {
waitUntil: 'domcontentloaded'
});
await page.addScriptTag({
url: 'https://cdn.example.com/report.js'
});
await page.waitForFunction(() => window.reportReady === true);
The browser process must be able to reach the URL. A script URL that works in your laptop browser can still fail in a container because of DNS, proxy, TLS, authentication, CSP, or mixed-content rules.
2. Make readiness deterministic
networkidle2 means that only a small number of network connections remain. It does not prove that a chart, table, or client-side application has finished rendering. Add a condition owned by the page.
<script>
async function renderReport() {
const data = await fetch('/api/report').then(response => response.json());
document.querySelector('#total').textContent = data.total;
document.querySelector('#report').classList.add('ready');
window.reportReady = true;
}
renderReport().catch(error => {
window.reportError = error.message;
});
</script>
await page.waitForFunction(() => window.reportReady === true, {
timeout: 30000
});
const error = await page.evaluate(() => window.reportError);
if (error) throw new Error(`Report failed: ${error}`);
A rendered selector is another good contract:
await page.waitForSelector('#report.ready', { visible: true });
Use a timeout as a safety limit, not as the definition of readiness. If the page has no controllable flag, wait for the selector that represents the final content and then verify its text or element count.
3. Playwright equivalent
import { chromium } from 'playwright';
const browser = await chromium.launch();
const page = await browser.newPage();
page.on('console', message => console.log('[browser]', message.text()));
page.on('pageerror', error => console.error('[page error]', error));
page.on('requestfailed', request => {
console.error('[request failed]', request.url(), request.failure()?.errorText);
});
await page.goto('https://example.com/report.html', {
waitUntil: 'networkidle'
});
// Only inject when the document has no script reference.
// await page.addScriptTag({ url: 'https://cdn.example.com/report.js' });
await page.waitForFunction(() => window.reportReady === true);
await page.pdf({
path: 'report.pdf',
format: 'A4',
printBackground: true
});
await browser.close();
Playwright documents load, domcontentloaded, networkidle, and commit navigation states and its PDF API. Treat networkidle as a navigation aid, then use an application-specific assertion.
4. PDF fidelity options that affect JavaScript output
Print versus screen CSS
Puppeteer generates PDFs with print media by default. If your layout only works with screen rules, call:
await page.emulateMediaType('screen');
Print styles can hide elements, change colors, or alter layout. Keep a dedicated print stylesheet when the PDF is a separate document.
Fonts and pagination
Puppeteer waits for fonts by default during page.pdf(). You can make the dependency explicit:
await page.evaluate(() => document.fonts.ready);
await page.pdf({ path: 'report.pdf', printBackground: true });
Late font swaps can change line wrapping and page breaks. Use stable font files, wait for document.fonts.ready, and avoid measuring text before fonts are loaded.
Colors and backgrounds
await page.addStyleTag({
content: '* { -webkit-print-color-adjust: exact !important; print-color-adjust: exact !important; }'
});
await page.pdf({ path: 'report.pdf', printBackground: true });
printBackground includes CSS backgrounds. Color adjustment can improve fidelity but may increase ink usage when users print the PDF.
Viewport and page size
await page.setViewport({ width: 1440, height: 900, deviceScaleFactor: 1 });
await page.pdf({
path: 'report.pdf',
width: '210mm',
height: '297mm',
margin: { top: '16mm', right: '14mm', bottom: '16mm', left: '14mm' },
preferCSSPageSize: true,
printBackground: true
});
Use either a named format such as A4 or explicit dimensions. CSS @page rules and preferCSSPageSize can control the final paper size.
5. External script edge cases
- CSP: A page policy can reject injected or CDN scripts. Check the browser console and response headers; serve an allowed origin or adjust the policy.
- Authentication: Set cookies or headers before navigation when the script or data endpoint requires a session.
- Mixed content: An HTTPS page may block an HTTP script. Use HTTPS for every dependency.
- Cross-origin frames: A script running in an iframe changes that frame, not the top page. Wait in the correct frame and print the page that owns the content.
- Module scripts: Preserve
type='module'and wait for the module’s application-ready signal. Import failures appear as page errors. - Workers: A web worker may finish after the DOM looks complete. Expose a flag only after worker results have been applied to the document.
- Animations: Disable transitions and wait for the final state to prevent half-rendered charts or moving elements.
- Lazy content: Scroll required sections into view or trigger the page’s loading routine before waiting for readiness.
6. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| JavaScript content is missing | The script failed, was injected into another frame, or PDF capture ran too early. | Attach console, page-error, request-failed, and response listeners; verify the URL and wait for a DOM marker or ready flag. |
addScriptTag rejects |
Bad URL, blocked network access, CSP, TLS, or authentication. | Open the URL from the browser context, inspect the failed request, and provide required cookies or headers. |
| PDF is blank | Navigation failed, the page redirected to a login screen, or the app never rendered. | Check page.url(), response status, screenshot the page during diagnosis, and fail when the ready condition times out. |
| Works visually but not in PDF | Print media rules hide or restyle the content. | Inspect print CSS, call emulateMediaType('screen') when appropriate, and enable printBackground. |
| Wrong page breaks or text wraps | Fonts loaded late or viewport and paper settings differ. | Await document.fonts.ready, pin fonts, set viewport and margins, and define @page. |
| Timeout at network idle | Analytics, sockets, or polling keep connections open. | Use a bounded navigation wait, then wait for a selector or application flag instead of indefinite network-idle waiting. |
| Browser process hangs | An exception skipped cleanup or a page remains open. | Wrap work in try/finally and always close the browser after the PDF buffer or file is produced. |
7. Reliability, performance, and cost
Reliability checklist
- Pin Puppeteer or Playwright and the browser revision in deployment.
- Use a per-job timeout and cancel work that exceeds it.
- Record URL, navigation status, script failures, readiness duration, and PDF duration.
- Reuse a browser process carefully, but create isolated pages and clear cookies between jobs.
- Retry transient network failures with a limit; do not retry deterministic CSP or JavaScript errors blindly.
- Close pages and browsers in
finallyblocks.
Performance
Browser startup is expensive, so a worker can keep one browser alive and create a fresh page per job. Reuse trusted static assets through caching, but avoid sharing page state between customers. Waiting for a specific readiness signal usually finishes sooner and more reliably than adding a long sleep. Disable unnecessary images, trackers, and animations only when they are not part of the document.
Cost
Self-hosting means paying for compute, browser memory, bandwidth, and operational maintenance. The PDF API itself has no per-page price in Puppeteer; your infrastructure is the cost. External CDNs and data APIs can add their own usage charges, and retries multiply that traffic.
8. Or skip the browser setup
ScreenshotNeo provides a hosted capture API and can produce screenshots or PDFs without maintaining Chromium workers. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. Its MCP server lets Claude, Cursor, and other MCP clients use take_screenshot, get_page_info, and capture_pdf.
For the complete parameter list and PDF options, see the ScreenshotNeo documentation.
curl -G 'https://api.screenshotneo.com/v1/shot' \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
import requests
r = requests.get(
'https://api.screenshotneo.com/v1/shot',
params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'},
timeout=90,
)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const body = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', body));
Every feature is available on every plan: 1,000 shots per month are free with no card, and paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
9. FAQ
Can a PDF library execute an external script by itself?
Only if it embeds a browser engine that executes the HTML. Otherwise use Puppeteer, Playwright, or a hosted browser capture service.
Is a five-second delay enough?
It is not a reliable rule. Network speed and application work vary; wait for a deterministic page condition with a timeout.
Should I inject a script with addScriptTag?
Inject it only when the source HTML does not already reference that script. Duplicate loading can run initialization twice.
Why does the browser view look correct while the PDF does not?
PDF generation uses print media by default, and fonts, backgrounds, viewport, and page margins can change layout. Inspect print CSS and wait for fonts.
Can I use Playwright instead of Puppeteer?
Yes. Both use Chromium page rendering, script loading, readiness waits, and a PDF API. Choose the tool that fits your existing browser versions and operational setup.


