ScreenshotNeo

BlogHTML to image & PDF

How to Export Dynamically Paginated Data to PDF

Export every intended record to PDF without missing rows or broken pages. Prepare the data, wait for rendering, apply print styles, and validate the result.

By the ScreenshotNeo team30 September 202613 min read

How to Export Dynamically Paginated Data to PDF

To export dynamically paginated data to PDF reliably, first decide which records the document should contain, then make all of those records available in a print-oriented view. Wait for your application data, charts, fonts, and images to finish rendering; apply print styles; generate the PDF through a browser’s print-to-PDF path; and inspect the resulting pages. Printing the current screen alone does not fetch records on other server-paginated pages.

This guide shows a browser-based approach with runnable Node.js and Python examples, explains when to use explicit pagination or a dedicated PDF engine, and covers readiness, page layout, edge cases, troubleshooting, performance, and cost. CSS paged media defines page size, margins, and orientation; fragmentation rules influence where content breaks. The browser’s output still needs validation against the data and layout you intend to ship. MDN’s CSS paged media guide documents these controls.

1. Decide what “all data” means

Before coding, define the export’s data contract. A table with server-side pagination might show 25 rows at a time while the report query matches 40,000. Calling print on that table usually prints the rendered view; it does not automatically retrieve every page. Decide whether a PDF represents:

  • The current visible page.
  • Rows the user selected.
  • Every record matching the current filters and sort order.
  • A consistent snapshot of the report at the time export began.

Make the choice visible in the product, especially when a large export may take time. The export endpoint should receive the same filters and permissions as the report, and it should enforce authorization server-side. For large datasets, prefer a server export job or a dedicated export data endpoint over asking a browser tab to fetch and retain every row.

Keep data semantics separate from layout. The application determines which records belong; the PDF renderer determines how those records flow across pages. Treat these as two explicit stages so an attractive PDF cannot conceal missing records.

2. Build a complete export view

A robust flow creates a print-specific route or component from the complete export dataset. It can reuse report labels and formatting, but should not rely on a virtualized table, infinite scroll, or the interactive page’s current pagination state. Virtualized interfaces render only rows near the viewport, so content outside that window may not exist in the DOM to print.

Reliable exports separate data completeness, rendering readiness, and page layout into distinct steps.
Reliable exports separate data completeness, rendering readiness, and page layout into distinct steps.

For modest datasets, load all intended records, render them, and then print. For larger datasets, fetch records in batches on the server, generate a report model, or create an asynchronous export job and let the user download the finished file. Avoid assuming that increasing a table’s page size solves every issue: response limits, memory, browser layout, and PDF size can become constraints.

Give the export view a clear completion signal. For example, set a root attribute only after the report data is loaded and charts have drawn:

// In the export page after the data and charts are ready:
document.documentElement.dataset.exportReady = "true";

Do not set it just because a network request finished if the UI still has state updates or chart work pending. Wait for the condition that means the document is actually printable.

3. Add print styles and page rules

Use @media print to hide controls that do not belong on paper and adapt screen layout for pages. Use @page for the page box, such as paper size, orientation, and margins. CSS fragmentation properties can discourage breaks within a row or heading, and can request a break before a report section. Browser support and exact pagination can vary, so inspect output in the renderer and version you deploy.

Printing the visible table is not the same as exporting every record in the report query.
Printing the visible table is not the same as exporting every record in the report query.
@page {
  size: A4 landscape;
  margin: 14mm 12mm 16mm;
}

@media print {
  .toolbar, .pagination-controls, .screen-only { display: none !important; }
  body { color: #111; background: #fff; font-size: 9pt; }
  .report-section { break-before: page; }
  h1, h2 { break-after: avoid; }
  table { width: 100%; border-collapse: collapse; }
  thead { display: table-header-group; }
  tr { break-inside: avoid; }
  th, td { border: 1px solid #bbb; padding: 4px; }
  a { color: inherit; text-decoration: none; }
}

The CSS declarations are requests to the paged layout engine, not guarantees that every oversized element fits. A table row taller than the printable area may still need to split or overflow. Wide tables may require landscape pages, smaller typography, selected columns, or a separate detail section. Prefer a readable PDF to shrinking everything until it fits.

Decide whether browser-generated headers and footers should be enabled. They can add page metadata, but may duplicate a title or conflict with a designed footer. Likewise, background graphics are often controlled by the PDF renderer separately from print CSS; set that option deliberately if colored backgrounds carry meaning.

4. Generate a PDF with a headless browser

Here is a runnable Node.js example using Playwright. Install Playwright and its browser according to its official documentation, then save this as export.mjs. It opens an export route, waits for the application-specific readiness marker, waits for fonts and images, and writes the PDF. Replace the URL and readiness signal with those used by your application.

import { chromium } from 'playwright';

const browser = await chromium.launch({ headless: true });
try {
  const page = await browser.newPage();
  await page.goto('http://localhost:3000/reports/export?reportId=42', {
    waitUntil: 'domcontentloaded',
    timeout: 60_000,
  });

  // Application contract: set only after all report data and charts are ready.
  await page.waitForFunction(
    () => document.documentElement.dataset.exportReady === 'true',
    { timeout: 60_000 },
  );
  await page.evaluate(async () => {
    await document.fonts.ready;
    const images = [...document.images];
    await Promise.all(images.map(img => img.complete ? Promise.resolve() :
      new Promise((resolve, reject) => {
        img.onload = resolve;
        img.onerror = () => reject(new Error(`Image failed: ${img.src}`));
      })));
  });

  await page.pdf({
    path: 'report.pdf',
    format: 'A4',
    landscape: true,
    printBackground: true,
    preferCSSPageSize: true,
    displayHeaderFooter: false,
    margin: { top: '14mm', right: '12mm', bottom: '16mm', left: '12mm' },
  });
} finally {
  await browser.close();
}

In production, authenticate the export route with a short-lived session or signed authorization mechanism. Do not place a durable secret in a URL that may appear in logs or browser history. Restrict access to the report and apply the same row-level permissions as the interactive view.

The example waits for image load events. If an image can legitimately fail or is decorative, decide whether to fail the export or omit it; silently waiting forever is not a recovery strategy. You may also need explicit waits for chart libraries, client-side data fetches, or custom web components.

5. Configure page size, breaks, and repeated content

Choose one source of truth for page dimensions. If the CSS @page rule defines size, enable the renderer’s preference for CSS page size where supported. Otherwise specify paper format and orientation in the PDF call. Mixing conflicting settings makes page dimensions harder to reason about.

Requirement Practical choice
Standard report A4 or Letter; portrait unless tables need width.
Wide data table Landscape, concise columns, repeated table headings.
Section starts on a new page Use a print-only class with break-before: page.
Keep heading with its content Use break-after: avoid and test long sections.
Avoid splitting a row Try break-inside: avoid on rows, while handling unusually tall rows.
Branded color blocks Enable background printing and verify contrast without color.

Browsers generally repeat table header groups when table markup and print layout permit it. Verify this with long tables and page boundaries. For charts, make sure their rendered dimensions are suitable for print; a screen-sized canvas can be blurry when enlarged, and a chart that loads after the readiness signal can be absent.

6. Wait for the right things

Navigation completion is not the same as application completion. A page may have completed its document load while a fetch, chart animation, font, lazy image, or component is still pending. Browserless’ PDF documentation exposes separate image and font readiness controls and notes that those options default to false in the described API. The same principle applies in a self-managed browser: explicitly wait for application data and required assets. See the Browserless PDF API documentation for its documented rendering options.

Prefer condition-based waits to arbitrary delays. A fixed sleep may be too short on a slow run and waste time on a fast one. When a dependency has no readiness event, use a bounded delay only as a fallback and keep a timeout with a useful error message. If network idle is used, account for pages with analytics, polling, or long-lived requests that may prevent idle from arriving.

Lazy-loaded images often appear only after scrolling into view. A print-specific export can disable lazy behavior or explicitly load image URLs. Do not assume a full-page PDF renderer will trigger every application’s lazy-loading logic.

7. When to add explicit pagination

Browser print layout is usually the simplest choice when HTML and CSS should remain the source of truth. If users need to preview exact pages in the application or the default fragmentation repeatedly produces poor results, consider an explicit pagination layer such as Paged.js. It describes in-browser fragmentation and a paginated preview, with headless browser automation available for output generation. Read its documentation and test the browser features your template depends on; paged-media behavior differs across implementations.

A direct HTML-to-PDF engine can suit a controlled template, but its CSS support may differ substantially from a browser. For example, tc-lib-pdf documents page flow and repeated table headers while stating that flexbox and grid are not implemented. Check the supported subset for the exact engine version before adapting a browser-built template. Its project documentation is the place to confirm the current details.

Choose based on fidelity to your source HTML, dataset size, control over page composition, need for a preview, deployment constraints, accessibility needs, and the renderer’s actual CSS support. A hosted rendering API may reduce browser operations, but assess its documented controls, privacy, deployment, and commercial terms against your requirements.

8. Python example with Playwright

The same workflow is available in Python. Install the Playwright Python package and its browser, then run this script. It uses the same export readiness contract and CSS page settings.

import asyncio
from pathlib import Path
from playwright.async_api import async_playwright

async def main():
    async with async_playwright() as p:
        browser = await p.chromium.launch()
        try:
            page = await browser.new_page()
            await page.goto(
                'http://localhost:3000/reports/export?reportId=42',
                wait_until='domcontentloaded', timeout=60_000
            )
            await page.wait_for_function(
                "document.documentElement.dataset.exportReady === 'true'",
                timeout=60_000
            )
            await page.evaluate("document.fonts.ready")
            await page.pdf(
                path='report.pdf', format='A4', landscape=True,
                print_background=True, prefer_css_page_size=True,
                display_header_footer=False,
                margin={'top': '14mm', 'right': '12mm',
                        'bottom': '16mm', 'left': '12mm'}
            )
        finally:
            await browser.close()

asyncio.run(main())

If images are essential, add an explicit image readiness check as in the Node.js version. The file write itself completes before the script exits; for an export service, also return a clear status, log the report identifier and failure stage, and clean up temporary files according to your retention policy.

9. Troubleshoot common failures

Symptom Likely cause Fix
Only the visible rows appear The view is client- or server-paginated, or virtualized. Build the export from the full intended query or selected dataset; do not print the interactive page state.
Last rows or charts are missing PDF generation began before fetches, chart drawing, or component rendering completed. Wait for an app-owned readiness marker set after all required work is done.
Fonts differ or text wraps onto extra pages Font files were not ready, unavailable, or blocked. Wait for document.fonts.ready, check font requests and fallbacks, then recheck page breaks.
Images are blank Lazy loading, failed requests, or premature capture. Load images in the export view, await required images, and decide how failed assets should affect the export.
Rows split awkwardly Fragmentation rules cannot keep the content together, or a row exceeds printable height. Try break-inside: avoid; shorten or restructure oversized cell content and inspect boundary cases.
Table columns are cut off Screen-width layout was carried into a portrait page. Use landscape, reduce nonessential columns, wrap long values, or create a separate detail page.
Blank pages appear Conflicting explicit breaks, oversized elements, or screen layout constraints. Inspect print styles and remove unnecessary fixed heights, min-heights, and break rules.
Export times out Readiness condition never becomes true, a request hangs, or network-idle is blocked by ongoing traffic. Log each export phase, bound every wait, report the failed condition, and avoid relying only on network idle.
PDF is missing colors Background printing is disabled. Enable background graphics in the renderer and preserve sufficient text contrast.
Layout differs from browser preview Different browser versions, media modes, or CSS support. Generate and preview with the same renderer family and version; simplify unsupported print layout.

10. Performance, reliability, accessibility, and cost

PDF layout cost grows with rendered content and complexity. Large DOM trees, many high-resolution images, complex charts, and huge tables all use memory and CPU. Batch data on the server where possible, avoid rendering invisible interactive controls, and consider asynchronous jobs for reports that take too long for a synchronous web request. Set limits for record count, output size, and execution time based on your product’s needs; the cited sources do not establish universal safe thresholds.

For reliability, make exports repeatable: record a report query or snapshot timestamp, use deterministic sorting with a unique tie-breaker, and include an export identifier in logs. Retry transient data or browser failures thoughtfully, but avoid restarting expensive work without limits. Ensure cleanup runs when a browser crashes or a job is cancelled. For sensitive reports, minimize retention and restrict access to both the source data and generated file.

Cost depends on where rendering runs. A self-hosted browser consumes your compute and operational time; a hosted renderer introduces provider pricing and data-handling considerations. A dedicated PDF engine can have a different licensing and maintenance profile. The available research provides no universal performance benchmark or service price, so measure representative reports in your deployment and compare complete operational costs rather than just render time.

Preserve semantic headings, table headers, labels, and meaningful reading order. Tagged PDF output is not automatically a compliance guarantee: Browserless documents that tagged structure depends on accessible source markup and that its tagged output is not certified PDF/UA. If accessibility or formal conformance is required, validate the actual output with an appropriate tool and conduct a compliance review. See its PDF API notes.

11. cURL example for a rendering service

If you use a hosted browser API, follow its own authentication, request schema, wait options, and PDF settings. The research dossier documents Browserless as one such option, but does not include an endpoint or credential format suitable for a runnable cURL command, so do not copy a guessed URL. Consult the provider’s current PDF API reference and pass only the data and access scope required for the export.

For services that accept a page URL, protect that URL from unauthorized access, avoid exposing secrets in query strings, and ensure the service can reach required assets. A successful HTTP response does not prove the document contains all expected records; validate page count and content as part of the export workflow.

12. Or skip the browser setup

For a screenshot of a report page or a PDF capture of a rendered page, ScreenshotNeo is a website screenshot API and MCP server. It takes one GET request for a URL and can return PNG, JPEG, WebP, or PDF. A screenshot of a URL captures what that page renders; it does not replace your export data path or guarantee that unseen server-paginated records are included. Build a complete export view first, then capture that view.

See the ScreenshotNeo API documentation for its parameters and PDF options. Example cURL request:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Cookie banners are accepted and removed before the shot, along with known newsletter popups and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. An MCP server lets AI agents in Claude, Cursor, or any MCP client use screenshot tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Sign up free for 1,000 screenshots a month, no card required.

13. Validation checklist

  1. Confirm the export query includes exactly the intended records and respects permissions.
  2. Test a short report, a report spanning many pages, and data at page boundaries.
  3. Test long text, wide values, oversized rows, charts, images, and missing assets.
  4. Run with slow and failed font or image loads; confirm the exporter waits or reports a useful error.
  5. Check page size, orientation, margins, repeated table headings, footer behavior, and background colors.
  6. Compare the generated document with the same browser and renderer version used in production.
  7. Validate accessibility and formal PDF requirements with suitable tools when those requirements apply.
  8. Track job duration, failure stage, record count, and output size without logging sensitive report contents.

Frequently asked questions

Does printing a paginated table include every page?

Only if all intended rows are present in the printable document. A server-paginated table usually contains just the current page, and a virtualized table may contain only visible rows.

Should every row be kept on one PDF page?

Try avoiding breaks inside normal rows, but very tall rows may not fit. Design long cell content so it can flow, or split it into a summary row and detail section.

Is a tagged PDF automatically accessible?

No. Structure depends on the source markup and renderer, and tagged output alone does not certify PDF/UA conformance. Validate the generated file against the applicable requirements.

When should I use Paged.js?

Use it when an explicit paginated preview or more deliberate page composition is useful and the extra pagination layer fits your rendering pipeline.

Can I export an arbitrarily large dataset in the browser?

There is no universal safe size. Large exports can strain browser memory and layout. For substantial reports, move data assembly and PDF generation into a bounded server-side job.