How to Convert a Long Webpage to PDF Without Cutting Off Content in PDFShift
Control page breaks, wait for dynamic content, and reduce resource delays to make long PDFShift conversions more complete and readable.
To keep a long webpage from being cut off in a PDFShift PDF, address three separate causes: control pagination with CSS, wait for asynchronous content to finish rendering, and reduce avoidable resource-loading delays. Then inspect the generated PDF and adjust the rules for the actual page. These controls can improve layout and completeness, but they do not guarantee that every page will render without clipping.
This guide covers PDFShift’s documented page-break CSS, wait_for readiness checks, timeouts, performance options, and troubleshooting. If you need a visual record rather than a document, ScreenshotNeo also offers a one-call webpage capture; its options and example are in the ScreenshotNeo documentation.
1. Identify what is being cut off
First determine whether content is present but split badly across pages, missing because it had not rendered yet, or absent because loading or conversion timed out. The fixes differ:
| Symptom | Likely cause | Start with |
|---|---|---|
| A heading or card is split across pages | Pagination rules do not suit the element | break-inside and page-break rules |
| A chart, font, or widget is missing | Asynchronous rendering was incomplete | wait_for readiness check |
| The conversion fails after a delay | Loading and processing exceeded the account timeout | Reduce resource work; verify the timeout |
| Some area is clipped at the page edge | The content’s print layout or sizing does not fit the page | Inspect print flow and adjust page-specific CSS |
Use the browser developer console to identify the element involved. PDFShift also recommends using print preview to see how CSS changes affect document flow. Make one change at a time so you can tell whether it fixed the observed problem.
2. Control page breaks with CSS
PDFShift documents the CSS properties break-before, break-after, and break-inside for controlling where content breaks. If you own the page, put the rules in its print stylesheet. If you cannot edit the source page, PDFShift documents sending CSS in the conversion request and targeting the source page’s elements.
/* Keep a short card, figure, or section together where possible. */
.keep-together {
break-inside: avoid;
}
/* Start a major section on a fresh PDF page. */
.start-new-page {
break-before: page;
}
/* End a section before the next section begins. */
.end-page-after {
break-after: page;
}
Apply a rule only to the elements that need it. Avoid forcing every paragraph, row, or small element to start a new page: excessive forced breaks can create blank space or awkwardly short pages. A long element that is taller than a page cannot be kept intact on one page; it must flow across pages or be redesigned for print.
For an external URL, the CSS can be supplied with the PDFShift request. The selector below is deliberately illustrative: inspect the actual page and replace .section-heading with a selector that matches the content you need to move.
curl --request POST \
--url https://api.pdfshift.io/v3/convert/pdf \
--header 'X-API-Key: YOUR_PDFSHIFT_API_KEY' \
--header 'Content-Type: application/json' \
--data '{
"source": "https://example.com/long-report",
"css": ".section-heading { break-before: page; } .card { break-inside: avoid; }"
}' \
--output report.pdf
Check the current PDFShift API documentation for request authentication and response details before using this example in an application. The example shows the documented idea of passing CSS with a URL source; the selector and rules must match the target page. PDFShift’s page-break guidance is in its PDF Page Breaks help article.
3. Wait for asynchronous content
A page can be considered loaded before its chart, custom element, or font has finished rendering. PDFShift’s wait_for parameter can defer conversion until a named global function on the page returns a truthy value. PDFShift polls that function until it succeeds or the conversion’s remaining overall timeout expires.
For a page you control, expose a readiness function after the relevant content is ready. For example, a chart can set a flag when rendering finishes:
<script>
window.reportReady = false;
renderReportChart().then(() => {
window.reportReady = true;
});
</script>
Then name that function in the conversion request’s wait_for option. The following is a Python request pattern; check PDFShift’s current API reference for the exact request schema and authentication method in use by your account.
import requests
response = requests.post(
"https://api.pdfshift.io/v3/convert/pdf",
auth=("YOUR_PDFSHIFT_API_KEY", ""),
json={
"source": "https://example.com/long-report",
"wait_for": "reportReady",
},
timeout=110,
)
response.raise_for_status()
with open("report.pdf", "wb") as pdf:
pdf.write(response.content)
PDFShift’s wait guide also demonstrates waiting for document.fonts.ready and checking whether a chart has reached its rendered height. If there is no suitable readiness function on the page, the guide describes injecting JavaScript in the request to create the needed check. See PDFShift’s wait-for-element guide and its Node.js custom-element example.
A readiness check must eventually become truthy. If it never does, the conversion waits until it runs out of time. Also make sure the function exists in the page’s global scope and that it represents the content you actually need; waiting for a generic page-load event may not mean a chart or third-party embed is ready.
4. Stay within PDFShift’s conversion timeout
PDFShift’s wait guide and FAQ describe an overall conversion timeout of 30 seconds for free accounts and 100 seconds for premium or paid accounts. The total includes loading and processing the document as well as converting it to PDF. The FAQ says a timeout response uses HTTP 408 with JSON. These account limits can change, so verify the current terms before relying on them. The guide says premium customers can contact PDFShift about extending the timeout.
A longer client-side HTTP timeout does not extend PDFShift’s own conversion limit. It only determines how long your application waits for a response. Treat a 408 as a signal to check whether the page is slow, whether a readiness function is waiting indefinitely, or whether unnecessary resources are delaying processing.
5. Reduce resource-loading delays
Network requests are a major source of conversion time. PDFShift’s performance guidance recommends reducing work the converter must perform. These measures can improve conversion time, but they do not independently fix every pagination or clipping problem.
- Send raw HTML instead of a URL when you can provide the document content directly.
- Inline CSS and JavaScript where practical so the conversion does not need extra requests to fetch them.
- Remove unnecessary scripts, such as trackers that do not affect the PDF.
- Embed image data where appropriate to avoid external image requests.
- Resize or optimize images for their intended PDF page size instead of loading oversized originals.
Keep scripts and resources that produce content the PDF needs. For example, removing a chart library may speed up conversion but leave the chart missing. After changing resources, inspect the PDF as well as the conversion response. Read PDFShift’s conversion-time guidance.
6. A repeatable conversion workflow
- Save a baseline PDF. Record the page URL and conversion settings so you can compare changes.
- Inspect the affected page boundary. Identify the specific element and whether it is split, clipped, or absent.
- Adjust pagination CSS. Use the appropriate break property on the smallest relevant set of elements.
- Add a readiness condition if content arrives late. Confirm the global function returns truthy after the needed content is actually rendered.
- Reduce unnecessary resource work. Remove only resources that are not required for the output.
- Convert again and inspect every affected page. If the break remains poor, revise the selector or rule for that page’s actual layout.
This inspection loop is a practical recommendation based on the documented controls. PDFShift’s sources do not publish a clipping success rate or promise that one setting prevents every kind of cut-off.
7. Troubleshooting
| Error or symptom | Cause to check | Fix |
|---|---|---|
| HTTP 408 or timeout JSON | Loading, processing, readiness polling, or PDF generation exceeded the overall account timeout | Check the page’s load cost and readiness function, remove unneeded resources, and verify the current account limit. |
| Wait condition never finishes | The named function is missing, not global, or never returns a truthy value | Expose it on window, confirm it reflects the required content, and make sure all success and failure paths update it. |
| Chart or custom element is blank | Conversion began before asynchronous rendering completed | Use a readiness function that checks the rendered result, such as the chart’s dimensions or completion flag. |
| Font differs or text wraps differently | Fonts may not have loaded before conversion, changing line wrapping and page flow | Wait for document.fonts.ready when appropriate, then review pagination again. |
| Heading appears at the bottom of a page | No break rule moves it with the section that follows | Apply break-before: page to the heading or a suitable wrapper, then inspect surrounding whitespace. |
| Card or row is split across pages | The element is allowed to fragment | Try break-inside: avoid on the element. It cannot keep content together if the element itself exceeds the available page height. |
| Large blank gaps after adding rules | Break rules are applied too broadly or conflict with the content’s natural flow | Narrow the selector and remove unnecessary forced page breaks. |
| Images are missing | External resources may be unavailable, slow, or removed during optimization | Check that required image URLs load; where suitable, embed image data or reduce image size instead of removing needed images. |
| Conversion is slow but completes | Many network requests or large assets add loading and processing work | Use raw HTML where practical, inline assets where appropriate, and remove scripts that do not affect the PDF. |
8. Reliability, performance, and cost considerations
Reliability: A readiness check makes a conversion depend on a condition you define. Keep that condition specific and bounded by the overall timeout. Inspect output after changes because CSS rules can improve one page while creating excess whitespace elsewhere.
Performance: Reduce avoidable requests and oversized assets first. A shorter conversion is easier to keep within the documented account timeout, but a fast conversion is not proof that asynchronous content rendered correctly.
Cost: This research dossier does not establish current PDFShift pricing, quotas, or the cost of retries. Check PDFShift’s current plan terms for those details. The documented timeout values are account-tier limits, not a statement about the price of a particular conversion.
9. Or skip the browser setup
If you need a screenshot of the page rather than a paginated PDF, ScreenshotNeo is a website screenshot API and MCP server. It accepts one GET request with a URL and returns a PNG, JPEG, WebP, or PDF. The PDF options include paper size, margins, landscape orientation, and page ranges. This is an alternative capture workflow; it does not guarantee that arbitrary long-page content will fit without pagination adjustments.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/long-report -o report.pdf
See the ScreenshotNeo API documentation for the request options. Cookie and consent banners are accepted as a visitor and removed before capture, along with 60+ known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the page verdict and billing status in headers. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Create a free ScreenshotNeo account and get 1,000 screenshots a month with no card.
10. FAQ
Can CSS guarantee that no content will be cut off?
No. The documented break properties let you control page flow, but the correct rules depend on the page’s actual structure and print layout. Inspect the generated PDF.
Can I use wait_for to wait an arbitrary number of seconds?
It is documented as a check for a named global function that returns a truthy value. For content readiness, prefer a function that checks the content itself over an arbitrary delay.
What if I cannot change the webpage?
PDFShift documents passing CSS in the conversion request so you can target elements in the source page. You still need to inspect the page and choose matching selectors.
Does optimizing resources fix page breaks?
It can reduce conversion time. Pagination still needs appropriate CSS when the page’s flow breaks poorly.


