ScreenshotNeo

BlogHTML to image & PDF

How to Make a PDF Archive of a Website That Requires JavaScript

Save JavaScript-rendered pages as PDFs by waiting for the content you need, then printing from a browser or automating the capture with Playwright.

By the ScreenshotNeo team4 October 20269 min read

To save a JavaScript-driven website as a PDF, first wait until the page has rendered the content you need, then use the browser’s Print command and choose Save as PDF. For repeatable captures, use Chromium with Playwright and wait for a site-specific readiness condition before calling page.pdf(). A generic delay can help, but it cannot tell you whether a particular app has finished loading its data.

A PDF is a fixed-layout record of a rendered page. It does not preserve the page’s scripts, network resources, or behavior as a replayable website. If you need an archive that packages web resources for later preservation, investigate WARC instead.

1. Save a page manually with browser Print

  1. Open the exact page in a current browser.
  2. Complete legitimate sign-in and consent steps required to access the content.
  3. Wait until the content you want to preserve is visible. Check charts, results lists, article text, and content that loads below the fold.
  4. Open the browser’s Print command and choose Save as PDF as the destination.
  5. Review the resulting PDF, including its first and last pages, page breaks, images, tables, and text selection.

Web apps often fetch content after the initial document loads. A visible shell, spinner, or empty chart is not evidence that the page is ready. Use a specific visible element or data state as your readiness signal, then inspect the saved file.

2. Print to PDF with Chrome Headless

For a simple command-line capture, Chrome Headless supports --print-to-pdf. Its command-line reference says this flag saves the target page as output.pdf in the current working directory. The --timeout option sets a maximum wait before capture, and --no-pdf-header-footer omits the print header and footer.

google-chrome --headless --disable-gpu \
  --timeout=10000 \
  --no-pdf-header-footer \
  --print-to-pdf=archive.pdf \
  "https://example.com/app/report"

Replace the URL with the page you are authorized to access. The timeout is a cap on waiting, not proof that an asynchronous application has finished rendering. For a page that requires a particular interaction or app-specific readiness check, use browser automation such as Playwright.

3. Automate PDF creation with Playwright

Playwright’s page.pdf() exports the rendered page using print CSS by default. The example below waits for a site-specific selector, then writes a PDF. Change [data-report-ready] to a selector that appears only when the content you need is ready.

Install

npm install playwright
npx playwright install chromium

Runnable Node.js example

const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch({ headless: true });
  const page = await browser.newPage();

  try {
    await page.goto('https://example.com/app/report', {
      waitUntil: 'domcontentloaded',
      timeout: 30000,
    });

    // Replace this with a condition specific to the target site.
    await page.locator('[data-report-ready]').waitFor({
      state: 'visible',
      timeout: 30000,
    });

    await page.pdf({
      path: 'archive.pdf',
      format: 'A4',
      printBackground: true,
      displayHeaderFooter: false,
      margin: {
        top: '12mm',
        right: '12mm',
        bottom: '12mm',
        left: '12mm',
      },
      tagged: true,
    });
  } finally {
    await browser.close();
  }
})();

The readiness selector is deliberately site-specific. If the page has no useful selector, wait for a known text value or application state. A delay can be a fallback, but it is less reliable than checking that the expected content has appeared.

Python alternative

Playwright also provides a Python API. Install it and its Chromium browser, then use a locator that represents the page’s ready state.

pip install playwright
playwright install chromium
import asyncio
from playwright.async_api import async_playwright

async def main():
    async with async_playwright() as p:
        browser = await p.chromium.launch()
        page = await browser.new_page()
        try:
            await page.goto(
                "https://example.com/app/report",
                wait_until="domcontentloaded",
                timeout=30_000,
            )
            # Replace with a selector that indicates the needed data is ready.
            await page.locator("[data-report-ready]").wait_for(
                state="visible",
                timeout=30_000,
            )
            await page.pdf(
                path="archive.pdf",
                format="A4",
                print_background=True,
                display_header_footer=False,
                margin={
                    "top": "12mm",
                    "right": "12mm",
                    "bottom": "12mm",
                    "left": "12mm",
                },
                tagged=True,
            )
        finally:
            await browser.close()

asyncio.run(main())

Make readiness checks reflect the page

  • Wait for a selector: use a result container, chart, or status element that appears after the data loads.
  • Wait for known text: useful when the page displays a stable report title or completion message.
  • Wait for network idle cautiously: analytics, polling, or persistent connections can keep a page active, while a quiet network does not always mean the right content loaded.
  • Use a timeout as a limit: a timeout prevents an automation job from waiting forever; it does not make an incomplete page complete.

4. Choose PDF settings for the result you need

Setting What it changes When to adjust it
Print CSS or screen media Playwright uses print media by default. Print styles can hide navigation or change layout. Use the default for a document-like print layout. Call page.emulate_media(media="screen") before PDF creation if you need screen styling.
Paper size Sets page dimensions, for example A4 or Letter. Choose the size expected by readers or the destination printer.
Margins Reserve space around printed content. Adjust if text is clipped, cramped, or too close to the page edge.
Background printing Playwright’s print_background option defaults to false. Enable it when background colors or images carry information; review contrast and ink-heavy areas.
Headers and footers Controls PDF header/footer output. Disable them for a cleaner document or configure them when page context is useful.
Scale Changes how page content fits the PDF. Adjust only after inspecting clipping and page breaks; no single scale works for every layout.
Tagged output Requests a tagged PDF structure in supported Playwright versions. Consider it when accessibility matters, then validate the file against the accessibility requirements for your use case.

To request screen styling in the Node.js example, add this before page.pdf():

await page.emulateMedia({ media: 'screen' });

To request it in Python, use:

await page.emulate_media(media="screen")

5. Inspect the PDF before treating it as an archive

  • Confirm that the expected JavaScript-generated content appears, including values, charts, and loaded images.
  • Check page boundaries for clipped paragraphs, split tables, and awkward breaks.
  • Check that colors and backgrounds are present when they convey meaning.
  • Verify links and text selection if readers need to navigate or copy from the PDF.
  • Record the source URL, capture date and time, browser and version, access context, and relevant print settings when the record matters.

Those capture details help explain how a particular PDF was produced. They do not turn the PDF into a complete copy of the website.

6. PDF snapshot or replayable web archive?

Choose PDF when people need a stable, readable rendering of a page. A PDF does not package the scripts and related resources needed to replay the site’s behavior. WARC is a web-archive format for captured resources and associated information; how archived resources are rendered can depend on the implementation. Dynamic features may also need a separate representation.

Need Better fit
Read or share a fixed rendering PDF
Repeatably generate page PDFs Browser automation such as Playwright
Preserve a collection of captured web resources for archival use Investigate WARC and the requirements of your preservation workflow

7. Troubleshooting

Symptom Likely cause Fix
PDF is blank or the app content is missing The capture happened before JavaScript fetched or rendered the data. Wait for a page-specific selector, text, or state. Reopen the page and verify the content is visible before printing.
Some results are missing even though the page opened The initial navigation completed, but later data or lazy content had not loaded. Wait for the relevant data state and scroll or otherwise trigger lazy content where the site requires it. Inspect the full PDF.
PDF looks different from the browser Playwright uses print media by default, and the site may have print-specific CSS. Use screen media emulation if screen styling is needed. Check paper size, margins, and scaling.
Backgrounds or colors are missing Background printing is disabled by default in Playwright PDF output. Set printBackground: true in Node.js or print_background=True in Python.
Long content is clipped or breaks badly Content does not fit the chosen page size, margins, or print styles. Inspect the affected page boundaries and adjust size, margins, scale, or print CSS. Recheck the output after each change.
Wait for selector times out The selector is wrong, hidden, blocked by a failed request, or never appears for that account or URL. Inspect the live page and choose a condition that actually represents readiness. Handle access and application errors before exporting.
PDF is inaccessible or reading order is poor Visual layout and document structure may not provide the semantics assistive technology needs. Try tagged output where supported and validate the PDF with tools and requirements appropriate to the use case. Tagging alone does not guarantee conformance.
Need the saved site to work offline A PDF stores a fixed rendering rather than the site’s scripts and resources. Use an archival workflow designed to capture related web resources, such as one based on WARC, and assess dynamic content separately.

Or skip the browser setup

For a screenshot of a rendered page, ScreenshotNeo provides a single-call website screenshot API. It returns an image or PDF, and its options include full-page capture, lazy-image loading, PDF paper size and margins, and wait conditions. See the ScreenshotNeo API documentation for supported parameters.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed, and response headers report the page verdict and billing status. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Learn about ScreenshotNeo or sign up for 1,000 free screenshots a month, with no card.

Performance, reliability, and cost

For one page, browser Print is usually the shortest setup. Automation adds browser installation and execution time, but makes repeated captures reproducible. The main reliability factor is readiness: waiting for a specific rendered state is more dependable than choosing one delay for every site. PDF generation itself can also fail when pages are large or the browser process exits, so close the browser in a cleanup path and verify that the expected output file was created.

Chrome Headless and Playwright are software workflows; plan for the compute and storage used by your own environment. This research provides no universal runtime, success rate, or cost benchmark, so measure with your own pages and capture settings. A PDF is typically a single output file, while a replayable archive must account for captured resources and preservation requirements.

FAQ

Does a JavaScript website need to be converted to static HTML first?

No. A browser can execute the page’s JavaScript, then print the rendered result. You do need to wait for the particular content you want to preserve.

Will Save as PDF preserve animations or interactive controls?

A PDF records a fixed rendering. Do not expect it to preserve the original page’s live behavior.

Can I use a fixed delay for every page?

You can, but it is fragile: pages load data at different times, and some may never become ready. Prefer a condition tied to the expected content, with a timeout as a safety limit.

Does tagged PDF output guarantee accessibility?

No. Tagged output is an option, not a guarantee of conformance. Check the resulting document against the accessibility needs of your use case.

References