ScreenshotNeo

BlogHTML to image & PDF

How to Load JavaScript from a URL When Generating a PDF in Python

Use Playwright for Python to load or inject a remote JavaScript file, wait for the page to finish rendering, and save the result as a PDF.

By the ScreenshotNeo team29 September 202610 min read

How to Load JavaScript from a URL When Generating a PDF in Python

If a page needs JavaScript to build its content before you print it, use a JavaScript-capable browser such as Playwright for Python. Navigate to the page, add the remote script with page.add_script_tag(url=...) if the page does not already include it, wait for the application’s own ready signal, then call page.pdf(). WeasyPrint can fetch remote resources, but it does not execute JavaScript during rendering. Playwright’s Page API documents script injection and PDF output; WeasyPrint’s API documentation describes its resource fetching behavior.

The key distinction is that downloading a JavaScript file is not the same as running it in a browser page. Pick the renderer based on whether your document depends on script execution, then wait for asynchronous page work before producing the PDF.

1. Choose the renderer that matches the page

Requirement Use Reason
Static HTML and CSS rendered through a Python PDF API WeasyPrint It converts HTML and CSS to PDF and can fetch network resources. It does not run page JavaScript.
Content or layout created by JavaScript Playwright with Chromium It runs the page in a browser, can inject a script from a URL, and can print the rendered page.
An archival PDF/A output Check the requested variant and renderer constraints PDF/A variants have restrictions, including restrictions on JavaScript as active content in the PDF. That is different from running a page script before creating a static PDF.

WeasyPrint may still be the right tool if your HTML has no JavaScript dependency. But if a page calls an API, builds a chart, hydrates a frontend, or reveals content only after a script runs, fetching the script as a resource will not make that content appear in a WeasyPrint document. See the WeasyPrint PDF/A guidance for output-format limitations.

2. Install Playwright for Python

Install the Python package and its Chromium browser. Playwright’s browser installation is a separate step from installing the package, so run both commands in the environment that will generate the PDF:

python -m pip install playwright
python -m playwright install chromium

For Linux deployments, consult Playwright’s installation documentation for system dependencies and deployment-specific setup. Pin package versions in your project’s dependency management so upgrades are deliberate. The example below uses Playwright’s synchronous Python API.

3. Load the page, inject the script, and print

Save this as make_pdf.py. Replace the example page URL and script URL. Replace window.reportReady with a readiness signal your page actually sets after its data and rendering are complete.

A browser must execute the remote script and finish its asynchronous page work before printing the rendered result.
A browser must execute the remote script and finish its asynchronous page work before printing the rendered result.
from playwright.sync_api import sync_playwright

PAGE_URL = "https://example.test/report"
SCRIPT_URL = "https://example.test/app.js"
OUTPUT_PATH = "report.pdf"

with sync_playwright() as playwright:
    browser = playwright.chromium.launch()
    page = browser.new_page()

    response = page.goto(
        PAGE_URL,
        wait_until="domcontentloaded",
        timeout=60_000,
    )
    if response is not None and not response.ok:
        raise RuntimeError(
            f"Page navigation failed: HTTP {response.status} for {PAGE_URL}"
        )

    # Omit this if the page already includes and runs the required script.
    page.add_script_tag(url=SCRIPT_URL, timeout=30_000)

    # Use the page's real readiness condition. Script onload alone does not
    # mean that later API calls, rendering, or animations have finished.
    page.wait_for_function(
        "window.reportReady === true",
        timeout=60_000,
    )

    # page.pdf() uses print CSS media by default. Keep this line if the
    # PDF should match the page's print layout.
    page.pdf(
        path=OUTPUT_PATH,
        format="A4",
        print_background=True,
        prefer_css_page_size=True,
    )
    browser.close()

The code checks the navigation response when one is available, injects a script by URL, waits for an app-specific condition, and writes an A4 PDF. The signal shown is illustrative: it is not built into Playwright. Your app might set a global flag, render a known selector, update a status element, or expose another dependable condition. Playwright documents that add_script_tag resolves when the script’s load event fires or its contents have been injected; that event does not promise that the script’s follow-up asynchronous work is complete. See the Page API reference.

If the page already loads the script

Do not inject it a second time if its own HTML already contains the correct script tag. Navigate, wait for the application-specific ready state, and print:

from playwright.sync_api import sync_playwright

with sync_playwright() as playwright:
    browser = playwright.chromium.launch()
    page = browser.new_page()
    page.goto("https://example.test/report", wait_until="domcontentloaded")
    page.wait_for_function("window.reportReady === true", timeout=60_000)
    page.pdf(path="report.pdf", format="A4", print_background=True)
    browser.close()

When the script does not expose a global flag

Wait for something observable that means the content is ready. For example, if the app adds a report container only after data arrives:

page.locator("#report-chart").wait_for(state="visible", timeout=60_000)

A visible element may still be incomplete if the app fills it incrementally. Prefer a signal that means the data and layout are done, or wait for a specific final value. If you control the page, a dedicated readiness flag set after rendering finishes is usually clearer than guessing with a fixed delay.

4. Decide whether the PDF should use print or screen styling

page.pdf() uses print CSS media by default. This is usually appropriate for a document intended for printing: the page can use @media print rules, page breaks, and print-specific margins. If the PDF should instead preserve the screen layout, switch media before printing:

Playwright prints with print media by default; emulate screen media when that layout is intended.
Playwright prints with print media by default; emulate screen media when that layout is intended.
page.emulate_media(media="screen")
page.pdf(path="report.pdf", print_background=True)

Playwright also documents PDF sizing and rendering options. The example uses format="A4", print_background=True, and prefer_css_page_size=True. Choose the options to match the document’s CSS and delivery requirements; avoid setting conflicting CSS page dimensions and PDF dimensions without checking which should take precedence. For exact printed colors, Playwright’s documentation describes the -webkit-print-color-adjust CSS property. Review the Page API options for the supported parameters.

5. Handle remote scripts and asynchronous content

Remote scripts introduce two separate waits: the browser must load and execute the script, and the application may then need to fetch data or render content. add_script_tag covers the script tag’s load or injection, not every task started by that code. Likewise, page.goto() reaching a navigation milestone does not automatically mean a single-page application has finished its own work.

  • Use the smallest adequate navigation milestone. domcontentloaded lets you begin waiting on an app signal without requiring every image and subresource to finish first.
  • Wait for the real completion condition. Use a readiness flag or locator tied to final content.
  • Set finite timeouts. A page that never becomes ready should fail with a useful error rather than hang indefinitely.
  • Check for script errors. A failed script request, a content security policy, or a runtime exception can prevent the readiness condition from ever becoming true.
  • Do not rely on an arbitrary sleep as the main synchronization method. A fixed delay can be too short on a slow run and waste time on a fast one.

If you cannot change the page, use the strongest observable completion condition available and document its limits. Network-idle-style signals can help with pages that settle, but analytics, polling, streaming, and long-lived connections can keep a page active. A selector or app-owned ready flag is often more specific.

6. WeasyPrint: what fetching a URL does and does not do

WeasyPrint’s Python API can load HTML and fetch resources such as stylesheets, fonts, and images. That resource handling does not turn the renderer into a JavaScript browser. If a remote script is supposed to populate a chart before PDF creation, WeasyPrint will not execute that script and the chart will not be generated by the JavaScript.

For static HTML, a basic API call looks like this:

from weasyprint import HTML

HTML(url="https://example.test/static-report").write_pdf("report.pdf")

Use this when the document is already represented in HTML and CSS, or when any dynamic content has been generated before WeasyPrint receives the document. If JavaScript must run, use a browser engine such as Playwright to render first. If you need a PDF/A variant, validate the output requirements against the chosen PDF/A specification and renderer documentation; PDF/A is not a setting that makes browser-executed page JavaScript part of the resulting PDF.

7. Security, reliability, and runtime considerations

A script loaded from a URL executes in the page context with the page’s capabilities. Use trusted script origins, controlled page input, and a constrained runtime environment. Avoid accepting an arbitrary user-supplied URL and then allowing the renderer to fetch every resource it references without controls.

WeasyPrint’s security guidance warns that untrusted HTML and CSS can cause security problems, resource exhaustion, long-running renders, and local-file access through file:// URLs. It recommends limiting time and memory, sanitizing inputs, controlling filesystem and network access, and using a custom fetcher to filter resource requests. Apply similar resource and network boundaries to browser automation. Playwright’s BrowserType API exposes a Chromium sandbox launch option and documents its default as false; review and configure the browser isolation behavior appropriate for your deployment. See WeasyPrint security guidance and Playwright BrowserType options.

Browser PDFs also consume browser memory and CPU. Reuse a browser process for batches of documents when appropriate, while giving each job an isolated page or context and closing resources when finished. Keep per-job timeouts and concurrency bounded. The exact cost depends on the page, its assets, browser environment, and workload; there is no universal render-time figure that applies to every site.

8. Troubleshooting

Symptom Likely cause Fix
PDF contains no chart or dynamic content The script did not run, or its asynchronous work had not completed. Check the script URL and browser console or network activity. Wait for the application’s actual ready signal before printing.
add_script_tag times out The remote script is unreachable, blocked by policy, or too slow. Open the script URL from the rendering environment, inspect the page’s content security policy and network errors, and increase the timeout only if a slower load is expected.
wait_for_function times out The condition is wrong, never set, or set before it reflects complete rendering. Check the app’s actual readiness behavior. Wait for a final selector or correct the flag and make sure page errors are not preventing it from being set.
PDF looks different from the browser window PDF output uses print media by default, or print CSS changes colors and layout. Use page.emulate_media(media="screen") before printing when screen styling is intended. Otherwise adjust the print stylesheet.
Backgrounds or colors are missing Print output may omit backgrounds or adjust colors for printing. Try print_background=True and review the print color rules documented by Playwright.
WeasyPrint output omits script-generated content WeasyPrint does not execute JavaScript. Render the page in Playwright first, or generate the content server-side before passing the document to WeasyPrint.
PDF job stalls or exhausts memory Slow remote resources, expensive content, unbounded concurrency, or untrusted input. Set time and memory limits, restrict reachable resources, cap concurrent jobs, and inspect requests that stall.
Browser launch fails on a deployment host Chromium or required system dependencies are missing, or deployment isolation differs from development. Install the Playwright browser and required host dependencies in the deployment environment, then review the documented launch and sandbox settings.

9. Or skip the browser setup

If the goal is a screenshot of a web page rather than a multi-page document assembled from browser content, ScreenshotNeo provides a website screenshot API. A single GET request returns an image or PDF, and the same parameter names used by other screenshot APIs also work. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server includes screenshot, page-info, and PDF capture tools for AI agents. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. Sign up for 1,000 free screenshots a month with no card.

10. FAQ

Can I use add_script_tag with an ES module?

Playwright’s method has a type argument; its Page API documents using type="module" for an ES6 module. Confirm the script’s module and origin requirements, then wait for your app’s readiness signal.

Does a generated PDF run the page’s JavaScript when opened?

The workflow here runs JavaScript in the browser before PDF creation and prints the rendered result. The generated pages represent that output; they are not a substitute for an interactive browser application. PDF/A formats also carry restrictions on JavaScript as active PDF content.

Can I use Python’s requests library to load the script?

You can download a script file with an HTTP client, but downloading it does not execute it or reproduce its browser environment. Use Playwright when the page needs a browser to run the code and render its result.

Should I use a fixed sleep after injecting the script?

Only as a fallback when no meaningful completion signal is available. A page-specific flag or final content selector makes the wait more reliable and gives a clearer timeout failure.

Primary references