ScreenshotNeo

BlogHTML to image & PDF

Why Is Lazy-Loaded Text Missing From My Full-Page PDF Capture?

Full-page capture does not guarantee lazy content has loaded. Find the trigger, wait for the text, and check whether print styles hide it.

By the ScreenshotNeo team4 October 20267 min read

Full-page PDF capture sets how much of a page to capture; it does not guarantee that viewport-triggered text has loaded. Scroll the relevant content into view, wait until the text appears in the DOM, then generate the PDF. If the text is present on screen but missing in the PDF, check print CSS and the capture media mode.

Lazy loading can mean different things: images or iframes may load near the viewport, while site JavaScript may fetch or insert text after an intersection, timer, or user action. Find the page’s actual trigger before choosing a wait strategy. Google’s guidance is that lazy-loaded content should load when it becomes visible in the viewport (Fix Lazy-Loaded Website Content).

1. Identify whether text failed to load or failed to print

Before changing timeouts, establish where the failure occurs:

  1. Open the page in a browser and inspect the missing region.
  2. Check whether its text exists in the DOM before scrolling. If it appears only after scrolling or interaction, the page has not loaded it yet.
  3. Scroll the region into view and wait for the actual text or a page-specific ready condition.
  4. Inspect the page in print preview or print media. If the text is visible on screen but absent in print, look for @media print rules or print-specific layout changes.
  5. Check the PDF itself. Text can be visually present but absent from its selectable or extractable text layer; that is a different issue from missing rendered content.

A full-page screenshot or PDF describes the output extent, not the application’s readiness. Capture APIs do not know that a particular paragraph is still waiting for an intersection callback, data request, or user action.

2. Capture with Puppeteer

This runnable Node.js example scrolls through the document in increments, waits for the target text, and creates a PDF. Replace the URL and expected text with values from the page. Install Puppeteer with npm install puppeteer; run the script with Node.js.

const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch({ headless: true });
  try {
    const page = await browser.newPage({ viewport: { width: 1280, height: 900 } });
    await page.goto('https://example.com/article', {
      waitUntil: 'domcontentloaded',
      timeout: 60000,
    });

    // Advance through the document so viewport-sensitive loaders can run.
    await page.evaluate(async () => {
      const step = Math.max(300, Math.floor(window.innerHeight * 0.8));
      for (let y = 0; y < document.documentElement.scrollHeight; y += step) {
        window.scrollTo(0, y);
        await new Promise(resolve => setTimeout(resolve, 150));
      }
      window.scrollTo(0, 0);
    });

    // Prefer a meaningful condition over a delay alone.
    await page.waitForFunction(
      () => document.body.innerText.includes('Expected paragraph text'),
      { timeout: 30000 },
    );

    // page.pdf() uses print media by default. Use this only if screen styling
    // is desired in the PDF instead of print styling.
    // await page.emulateMediaType('screen');

    await page.pdf({
      path: 'page.pdf',
      format: 'A4',
      printBackground: true,
      waitForFonts: true,
      timeout: 60000,
    });
  } finally {
    await browser.close();
  }
})().catch(error => {
  console.error(error);
  process.exitCode = 1;
});

The scroll loop is a general trigger, not a guarantee. A site may observe a particular container, require a click, or fetch content only after another application state changes. For a known target, scrolling that element into view can be more precise:

await page.locator('#target-section').scrollIntoViewIfNeeded();
await page.waitForFunction(() => {
  const el = document.querySelector('#target-section');
  return el && el.innerText.includes('Expected paragraph text');
}, { timeout: 30000 });

Puppeteer’s PDF API uses print CSS media by default. Call page.emulateMediaType('screen') before page.pdf() if the desired output should use screen styles. Its waitForFonts option waits for document.fonts.ready; it does not trigger lazy loading or fetch page data.

3. Capture with Playwright

Playwright can generate a PDF in Chromium. As with Puppeteer, first trigger the page’s loading behavior and verify the text, then produce the document. Install with npm install playwright.

const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch({ headless: true });
  try {
    const page = await browser.newPage({ viewport: { width: 1280, height: 900 } });
    await page.goto('https://example.com/article', {
      waitUntil: 'domcontentloaded',
      timeout: 60000,
    });

    await page.evaluate(async () => {
      const step = Math.max(300, Math.floor(window.innerHeight * 0.8));
      for (let y = 0; y < document.documentElement.scrollHeight; y += step) {
        window.scrollTo(0, y);
        await new Promise(resolve => setTimeout(resolve, 150));
      }
      window.scrollTo(0, 0);
    });

    await page.getByText('Expected paragraph text', { exact: false }).waitFor({
      state: 'visible',
      timeout: 30000,
    });

    await page.pdf({ path: 'page.pdf', format: 'A4', printBackground: true });
  } finally {
    await browser.close();
  }
})().catch(error => {
  console.error(error);
  process.exitCode = 1;
});

Playwright documents PDF generation and full-page screenshots as separate capture operations. Choose the output that matches the need; neither capture extent nor API choice proves that site-specific lazy content has loaded.

4. Capture with Chrome headless

Chrome’s command-line options can help when the page needs more time, but a timeout is only a maximum wait. It does not scroll the page or prove that a particular request completed.

chrome --headless --no-sandbox \
  --timeout=15000 \
  --print-to-pdf=page.pdf \
  'https://example.com/article'

Chrome documents --timeout for headless DOM dumping, screenshots, and PDF printing. It also supports --virtual-time-budget to fast-forward timer-driven code:

chrome --headless --no-sandbox \
  --timeout=20000 \
  --virtual-time-budget=10000 \
  --print-to-pdf=page.pdf \
  'https://example.com/article'

Virtual time may help with setTimeout or setInterval behavior, but it is not evidence that a network request, viewport trigger, or user interaction succeeded. See the Chrome headless documentation.

5. Check print CSS and PDF settings

If the missing text exists in the DOM before capture, investigate rendering rather than adding more scrolling:

  • Search stylesheets for @media print, display: none, visibility rules, clipping, and print-only layout changes.
  • Compare screen and print media. Puppeteer’s default page.pdf() output uses print media; screen emulation changes which CSS rules apply.
  • Check whether the text is inside a fixed-height or overflow-clipped container that behaves differently on paper.
  • Confirm whether the text is visually absent or only missing from copy-and-paste or PDF text extraction.
  • Wait for the content itself before waiting for fonts. Font readiness affects typography, not whether lazy text has loaded.

6. Nested scroll containers and other edge cases

A document-level scroll may not move an inner panel. If the missing content is in a chat transcript, table, feed, or other independently scrolling area, find its scroll container and advance that element. Check whether the text appears in the DOM after doing so.

Some pages load content after clicking “Show more,” expanding an accordion, accepting a prompt, or changing a filter. Reproduce the user action in the capture flow. Other pages append results as you scroll; a single jump to the bottom can skip intermediate visibility events, so use incremental scrolling and confirm the expected content.

For very long pages, a scroll loop that waits at every step can take time. Narrow it to the regions containing required content when possible, while preserving the page’s expected trigger behavior.

7. Troubleshooting common failures

Symptom Likely cause Fix
Text appears only after manual scrolling Viewport-triggered loading has not run before capture. Scroll in increments or bring the target into view, then wait for the text in the DOM.
Longer timeout changes nothing The trigger requires visibility, interaction, or a separate request condition. Trigger the site behavior explicitly; do not use time alone as proof of readiness.
Text is in the DOM but absent from the PDF Print styles, clipping, or print layout changes hide or move it. Inspect print CSS and compare print media with screen media.
Top-level scrolling does not load panel content The content uses an independently scrolling element. Scroll the inner container and confirm the text appears before printing.
PDF looks right, but copied text is missing The rendered page and PDF text layer differ. Separate visual rendering from text extraction when diagnosing; inspect the PDF output with the intended reader or extraction workflow.
PDF times out waiting for fonts Fonts have not reached document.fonts.ready within the capture window. Check font requests and use an appropriate font wait timeout; remember this does not load lazy text.

8. Performance, reliability, and cost

Incremental scrolling and explicit content checks are more reliable than an arbitrary large delay because they test the page’s state. They add time, especially on long documents, so target the necessary regions and use a condition tied to the expected text. A fixed delay can still be a practical fallback for a known page, but network and rendering time can vary.

For repeatable captures, record the URL, browser and library version, viewport, readiness condition, PDF media mode, and timeout. This makes it easier to distinguish a site change from a capture configuration change. No single timeout or virtual-time budget guarantees that every page’s custom loading logic has completed.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. Its PDF capture options include waiting for a selector, delay, or network idle, plus full-page capture. One GET request can return a PDF; see the ScreenshotNeo API docs.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com/article \
  -d format=pdf \
  -o page.pdf

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, and failed loads are never billed, and response headers say which page verdict occurred and whether it was billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.

Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

10. FAQ

Does a full-page PDF capture scroll the page?

Do not assume it does. Full-page output extent is separate from whether a site’s viewport-sensitive loading triggers ran. Explicitly scroll or otherwise trigger the page, then verify the target text.

Will waiting for network idle fix missing text?

Not in every case. The page may not request the content until it becomes visible or a user action occurs. Trigger that condition first, then wait for an application-specific ready signal.

Should I use a PDF or a full-page screenshot?

Use a PDF when you need a paginated document or print output. Use a full-page screenshot when a single image of the rendered page is the intended artifact. In either case, confirm lazy content is present before capture.