ScreenshotNeo

BlogHTML to image & PDF

How to Convert a Long HTML Page to PDF Without Cutting Off Content in PDFCrowd

Fix clipped PDFCrowd output by checking content fitting, lazy-loaded content, page breaks, and usable page area. Choose paginated or single-page output deliberately.

By the ScreenshotNeo team4 October 20268 min read

To stop a long HTML page from being cut off in PDFCrowd, first decide whether you want a normal paginated PDF or one continuous tall page. For ordinary pages, inspect content_fit_mode and the available page area; for content that appears only after scrolling, adjust content_viewport_height; for a truly continuous page, set page_height to -1. These settings solve different problems, and none guarantees that every source page will fit correctly.

1. Identify what “cut off” means

Before changing settings, inspect the output and locate the first missing or clipped content. A wide table disappearing beyond the right edge points to a width-fitting problem. Content missing near the bottom may be lazy-loaded or may need a larger content viewport. Overlapping headers, footers, or body content points to insufficient usable page area. A section splitting awkwardly across pages is usually a page-break issue.

Also choose the intended output:

  • Paginated PDF: normal paper-sized pages, suitable for reading and printing. Content must be scaled or broken across pages.
  • Single continuous page: one unusually tall PDF page. This avoids ordinary page breaks but can be inconvenient to view, print, or open in some PDF readers.

2. Check PDFCrowd’s content fitting mode

PDFCrowd’s HTTP API documents these content_fit_mode values: auto, smart-scaling, no-scaling, viewport-width, content-width, single-page, and single-page-ratio. The right setting depends on the layout and desired output. In particular, PDFCrowd warns that no-scaling can leave content cut off when it extends beyond page boundaries.

Mode When to consider it
auto A sensible first setting when you want PDFCrowd to choose fitting behavior.
smart-scaling When automatic fitting should reduce oversized content while retaining a paginated layout.
no-scaling When preserving rendered size matters and you have confirmed the content fits; it can clip content outside page boundaries.
viewport-width When fitting to the browser viewport width is appropriate for the source layout.
content-width When fitting based on the content width is more appropriate than viewport width.
single-page, single-page-ratio When you want a one-page result; inspect the resulting dimensions and readability.

For content clipped at the horizontal edge, compare the width-oriented options. For content that is oversized in both dimensions, start with auto or smart-scaling and inspect the result. Do not assume one mode will fix every page: CSS layout, content dimensions, and page settings all affect the outcome.

3. Load content that appears lazily

Some pages render more content only after scrolling or after the page has had time to load. PDFCrowd provides content_viewport_height for this case. Its documentation says auto uses the print area height and is sufficient for most pages. If extensive lazy-loaded content is missing, try large or a custom numerical height, then check whether the missing section appears.

The HTTP reference includes 10000px as an example, while the command-line reference includes 5000px. Those are examples, not universal values. A larger viewport can change how much content the converter loads, so use the smallest value that reliably exposes the content you need and review the resulting PDF.

4. Use one continuous page only when that is the goal

For a single tall page, PDFCrowd documents page_height=-1, which expands the page vertically to fit the content. PDFCrowd gives 200 inches as the safe maximum because larger pages may fail to open in some viewers. A very tall page is not the usual choice for a document intended for ordinary reading or printing. If you need a conventional document, use paginated output and tune fitting and page breaks instead.

5. Control page breaks in HTML you own

If you control the source HTML, use print CSS to request breaks at sensible boundaries and to keep small elements, such as table rows, together where possible. PDFCrowd’s examples show a forced break before a section and a rule to avoid breaking table rows. Its example also warns: “Content taller than a page still needs to break.” Avoid applying page-break-inside: avoid indiscriminately to large containers that cannot fit on one page.

<style>
/* Start a major section on a new printed page. */
.report-section {
  page-break-before: always;
}

/* Keep a table row together when it can fit on a page. */
tr {
  page-break-inside: avoid;
}

/* PDFCrowd-specific conversion styles can be scoped this way. */
.pdfcrowd-body .report-title {
  margin-top: 0;
}
</style>

PDFCrowd documents the .pdfcrowd-body prefix for conversion-specific styles; its examples explain that this prefix does not change the page in a visitor’s browser. For current CSS behavior, consult the PDFCrowd HTTP API examples.

6. Review page size, margins, headers, and footers

The content area is affected by paper size, margins, and header and footer height. If text is clipped or overlapped after changing the fit mode, inspect those settings together. A large header or footer, or generous margins, leaves less room for the body. If the source HTML’s own spacing contributes to the problem, PDFCrowd’s FAQ shows how to suppress default html and body margins and padding.

Change one factor at a time and compare the affected page in the generated PDF. This makes it easier to tell whether the issue came from scaling, page geometry, or source CSS.

7. Run a conversion and inspect the result

  1. Choose paginated or continuous output.
  2. For paginated output, select a fitting mode suited to the clipping: check width-oriented modes for horizontal overflow and automatic or smart scaling for generally oversized content.
  3. If content appears only after scrolling, adjust content_viewport_height from auto to large or a considered custom height.
  4. If you own the HTML, add targeted print page-break rules and review margins and headers or footers.
  5. Generate the PDF and inspect the first affected page, the final page, wide content such as tables, and any section that previously disappeared.
  6. Change one setting at a time until the output matches the intended layout.

PDFCrowd’s option names and availability depend on the interface and version. Its command-line reference says content_fit_mode and content_viewport_height require API client version 6.0 or later and converter version 24.04 or later. Check the reference for the client you use before troubleshooting an unrecognized option.

8. Troubleshooting

Symptom Likely cause What to try
Right side of a table or layout is missing Content is wider than the available page area, or scaling is disabled. Review content_fit_mode; compare viewport-width, content-width, auto, or smart-scaling. Check page size and margins.
Lower sections are absent Content is lazy-loaded or did not appear within the conversion viewport. Try content_viewport_height=large or a custom height, then verify the missing section is present.
Content is clipped despite changing the fit mode Page geometry, headers, footers, or source CSS may still leave insufficient content area. Inspect margins, paper size, header/footer heights, and source html/body spacing. Check the rendered HTML if possible.
Sections break in awkward places No deliberate print break rules, or an unbreakable block is too tall. Add targeted page-break rules to sections or table rows. Let content taller than a page split.
The whole document is on one huge page A single-page setting or page_height=-1 is in use. Use ordinary paginated output if the PDF is for standard reading or printing; keep tall-page output only when it is specifically needed.
PDF will not open in some viewers The page may be extremely tall. For custom single-page height, stay within PDFCrowd’s documented 200-inch safe maximum or use paginated output.
Client reports an unknown option The client or converter version may not support the parameter. Check the applicable API reference and version requirements; the command-line reference specifies client 6.0+ and converter 24.04+ for the fitting and viewport options.

9. Performance, reliability, and cost considerations

Increasing the content viewport can make more lazy content available, but it is not a substitute for checking whether the page actually loads that content. Very tall single-page PDFs can also be less convenient and may exceed some viewers’ practical limits. Prefer the smallest viewport adjustment that resolves the missing content, and use paginated output when it serves the reader better.

The cited PDFCrowd references describe settings and expected behavior; they do not provide a conversion success rate or guarantee a particular mode will repair a given page. The research for this guide does not establish PDFCrowd pricing, so check its current plan and pricing information directly before estimating conversion cost.

Or skip the browser setup

If you need a clean image capture of a webpage rather than a PDF, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF, with options including full-page capture and waiting for a selector, a delay, or network idle. It is not a replacement for PDFCrowd’s long-page PDF layout controls when your goal is a paginated document.

For example, this cURL request captures a page as WebP:

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Equivalent examples in Python and Node.js:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
const bytes = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

See the ScreenshotNeo API documentation for parameters and output options. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed; an MCP server lets AI agents take screenshots; and 1,000 screenshots a month are free with no card, with paid plans starting at $5 for 3,000. Sign up for free.

FAQ

Should I use a single-page PDF for a long article?

Only if one continuous page is the intended result. For normal reading and printing, use paginated output and fit the content to the page area.

Is 10000px the right content viewport height?

It is an example in PDFCrowd’s HTTP reference, not a universal setting. Start with auto, then choose a suitable larger value only if lazy-loaded content is missing.

Can page-break CSS prevent every split?

No. It can guide breaks, but content taller than a page still needs to break.

Where should I check parameter support?

Use the documentation for your PDFCrowd interface and installed version. The command-line reference lists the version requirements for fitting and viewport options.

References