ScreenshotNeo

BlogHow-to

How to capture a long scrolling webpage with APITemplate.io without cutting off content

Use APITemplate.io to render a long webpage to PDF, prompt lazy-loaded content to appear, and check for missing sections before sharing the result.

By the ScreenshotNeo team4 October 20268 min read

Direct answer: APITemplate.io documents a URL-to-PDF endpoint for publicly accessible webpages. For pages that load content as you scroll, render the page in a browser, scroll through it in viewport-height steps, wait for new content, and then generate the PDF. This can prompt lazy-loaded content to appear, but it does not guarantee a complete capture. Inspect the beginning, middle, and end of the resulting PDF.

The documented URL endpoint returns a paginated PDF, not one continuous, extra-tall screenshot image. If you need a screenshot image instead, use a screenshot capture workflow and confirm that it supports the page’s loading behavior.

1. Generate a PDF from a public URL

APITemplate.io’s URL-to-PDF API accepts a URL and an API key. Its documentation describes headless Chromium rendering, so JavaScript and modern CSS can render in the browser before the PDF is created. Follow its current API documentation for the exact request schema and available settings: 4 Ways to Generate a PDF.

curl -X POST "https://api.apitemplate.io/v2/create-pdf-from-url" \
  -H "X-API-KEY: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/long-article",
    "settings": {
      "paper_size": "A4",
      "orientation": "portrait",
      "margin_top": "12mm",
      "margin_bottom": "12mm",
      "margin_left": "12mm",
      "margin_right": "12mm",
      "print_background": true
    }
  }' \
  --output webpage.pdf

Replace the URL and API key. The documentation lists layout settings such as paper size, orientation, margins, print backgrounds, and header/footer options. Check the current endpoint schema for accepted values and response behavior before using this request in production.

2. Prompt lazy-loaded content to render

A page may load images or sections only when they approach the viewport. APITemplate.io’s Puppeteer article recommends scrolling through the page and waiting before generating the PDF. Its example estimates the page height from the body, scrolls in viewport-sized increments, pauses, and returns to the top. The article explicitly cautions that its body-height calculation does not always work. A page that appends items during scrolling can outgrow an earlier measurement.

The following Puppeteer pattern illustrates the workflow. It is a browser-side approach you can adapt to your own rendering pipeline; it is not a documented parameter that can be passed to the URL-to-PDF endpoint. A reliable implementation should recheck height after scrolling and stop only after a deliberate completion condition, such as a known end marker or several passes with no new height. Some infinite-scroll pages never reach a stable end.

import puppeteer from "puppeteer";

const url = "https://example.com/long-article";
const browser = await puppeteer.launch({ headless: true });

try {
  const page = await browser.newPage({
    viewport: { width: 1365, height: 900 },
  });
  await page.goto(url, { waitUntil: "domcontentloaded", timeout: 60000 });

  // Give initial scripts and layout a chance to settle.
  await page.waitForTimeout(1000);

  let previousHeight = 0;
  let stablePasses = 0;
  const maxPasses = 80;

  for (let pass = 0; pass < maxPasses; pass += 1) {
    const height = await page.evaluate(() =>
      Math.max(document.body.scrollHeight, document.documentElement.scrollHeight)
    );

    if (height <= previousHeight) {
      stablePasses += 1;
    } else {
      stablePasses = 0;
    }

    if (stablePasses >= 3) break;
    previousHeight = height;

    await page.evaluate(() =>
      window.scrollBy(0, Math.max(300, window.innerHeight - 100))
    );
    await page.waitForTimeout(500);
  }

  await page.evaluate(() => window.scrollTo(0, 0));
  await page.waitForTimeout(300);
  await page.pdf({
    path: "webpage.pdf",
    format: "A4",
    printBackground: true,
    margin: { top: "12mm", right: "12mm", bottom: "12mm", left: "12mm" },
  });
} finally {
  await browser.close();
}

This bounded loop avoids an unending scroll, but it still cannot prove completeness. A slow request may arrive after a pause; content might be gated on a button click; a site may virtualize earlier items out of the DOM; or the page may add content without changing document height. Set timeouts and stopping conditions for the target site, and inspect the PDF.

APITemplate.io’s source for this scrolling approach: 7 Tips for Generating PDFs with Puppeteer.

3. Set PDF layout separately from page loading

Paper settings decide how rendered content is paginated. They do not fetch content that has not loaded.

Setting What it affects Practical guidance
Paper size Page dimensions and line wrapping Choose a common size such as A4 or Letter based on the audience and intended printer.
Orientation Available width and height per page Portrait suits articles; landscape can help with wide tables, though it may reduce readability for text.
Margins Printable area and whitespace Use enough margin to keep text and page furniture away from edges.
Print backgrounds Whether background colors and images appear Enable when the design depends on backgrounds; check the file size and contrast afterward.
Headers and footers Repeated page details Use only when useful; verify they do not overlap page content.
CSS page size Whether CSS @page dimensions control output APITemplate.io documents this behavior with rendering engine 153 and preferCSSPageSize; consult the current docs for the exact configuration.

See APITemplate.io’s documentation on custom page size and margins. For authored HTML, CSS page-break rules can influence pagination. For example, break-inside: avoid can keep a suitable short card or table row together. These rules cannot restore lazy-loaded content that was never fetched. The HTML template editor and advanced PDF topics cover related layout controls.

4. Review the output for cut-off content

  1. Check the first page. Confirm the title, opening content, and any expected hero image are present.
  2. Sample the middle. Look for missing sections, repeated blocks, image placeholders, or abrupt jumps.
  3. Check the end. Make sure the final expected section appears and the capture did not stop mid-list.
  4. Review page boundaries. Watch for headings stranded at a page bottom, clipped tables, overlapping headers, or unexpected blank pages.
  5. Compare against the live page. For infinite scroll, use an expected item count, final marker, or other site-specific signal when possible.

APITemplate.io distinguishes previewing a document from generating a PDF with settings applied. Treat the downloaded output as the artifact to review, not the preview alone.

5. Common problems and fixes

Symptom Likely cause What to try
PDF ends before the visible page does The renderer captured before scroll-triggered content loaded, or the scroll loop stopped early. Scroll in smaller steps, wait after each step, remeasure height after scrolling, and use a site-specific end condition. Inspect the final pages.
Images are blank or missing Images are lazy-loaded, remote requests are slow, or the page has not finished rendering them. Scroll images into view, wait for image loads where practical, and confirm the source URL is publicly accessible to the renderer.
Only the initial items appear on an infinite-scroll page The page appends items only after a scroll event or user interaction. Trigger the page’s actual loading behavior. If it requires clicking “Load more,” automate that action; scrolling alone may not be sufficient.
Content is clipped at page edges Paper size, orientation, margins, or wide fixed-width content do not fit. Adjust layout settings or the page’s print CSS; use landscape only when it improves readability.
Elements split awkwardly across pages Print pagination rules are absent or unsuitable. For controlled HTML, apply appropriate page-break rules such as break-inside: avoid to short elements. Recheck tables and long blocks.
Blank PDF or navigation error The page may be inaccessible to the renderer, require a login, block automated access, or fail before rendering. Confirm it is publicly reachable and loads without a session. The reviewed endpoint documentation specifies publicly accessible pages; it does not establish a login-gated capture guarantee.
Request is slow or times out Large pages, slow scripts, or third-party resources delay rendering. Set a suitable client timeout, wait only for the readiness condition you need, and avoid unbounded scrolling. Check current API limits and settings in the official docs.

6. Performance, reliability, and cost considerations

Long pages take longer to render because the browser may need to execute scripts, load images, and lay out many sections. Scrolling in small steps with pauses improves the chance that lazy content is requested, but adds time and still cannot establish that every site-specific condition has completed. Bound the number of scroll passes and total runtime. If the page is known to have a finite item count, use that as a completion check.

Reliability depends on the page as well as the PDF renderer: network latency, client-side errors, consent overlays, bot checks, authentication, and infinite-scroll behavior can all affect the result. Keep representative captures under review when the source page changes. The available research does not provide a completeness guarantee or numeric performance benchmark.

APITemplate.io’s getting-started guide lists US, EU, Australia, and Singapore endpoint regions. Region options and behavior can change, so check the current first-request guide if latency or data location matters. Confirm current pricing and usage limits directly with the provider; no pricing figures are established by the sources used here.

7. Or skip the browser setup

For an image capture, ScreenshotNeo provides a website screenshot API and MCP server. Its one-request API returns PNG, JPEG, WebP, or PDF, with options for full-page captures and lazy images. This is an image/PDF capture workflow; it does not turn an endless page into a guaranteed complete document.

See the ScreenshotNeo API documentation. Example request for a full-page WebP:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com", "full_page": "true"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com',
  full_page: 'true'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

These examples use the documented API base and the full-page option; check the docs for output-format parameters and other settings. ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture, and each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for 1,000 free screenshots a month, with no card required.

FAQ

Does APITemplate.io return one tall screenshot?

The documented URL endpoint generates a PDF. PDFs are paginated according to document settings; the reviewed sources do not establish that this endpoint returns one continuous screenshot image.

Will scrolling guarantee that every infinite-scroll item is included?

No. Scrolling can trigger lazy loading, but the source article warns that its height estimate does not always work. Use a page-specific completion signal and inspect the output.

Can I capture a page that requires a login?

The cited URL-to-PDF documentation describes publicly accessible pages. It does not establish a general guarantee for login-gated pages.

Does changing the paper size fix missing content?

No. Paper size and margins change pagination and layout. They do not cause unloaded sections to appear.

Official sources