How to Archive a Webpage as PDF When It Uses Shadow DOM
Save a webpage with Shadow DOM as a PDF using browser print or headless Chrome, then check the output for missing content and layout issues.
To archive a webpage that uses Shadow DOM, open it in a browser and use Print → Save to PDF. Check the print preview and the saved PDF to confirm that the component content appears and that pagination, margins, orientation, and backgrounds are suitable. For repeatable captures, Chrome Headless can print a page to PDF, but its timeout only limits how long it waits; it does not confirm that the page’s asynchronous components have finished loading.
Shadow DOM encapsulates a component’s internal tree from ordinary document-level selectors and styles. That can make it harder to inspect or restyle component internals with page scripts, but the fact that content is visible on screen does not guarantee that every browser will print it identically. There is no universal Shadow DOM PDF repair switch.
1. Save a page as PDF in Firefox
- Open the webpage in Firefox and choose Print.
- In the destination list, select Save to PDF.
- Review the preview. Confirm that the text and content inside the web components appear.
- Adjust page selection, orientation, margins, and print backgrounds if needed.
- Save the PDF, then open the file and inspect it. Check more than the first page if the page is long.
Firefox documents these print controls in its web page printing help. If component content is missing or badly formatted, try the site’s own print or export feature, or compare the result in another browser. The result can depend on the page and browser.
2. Print to PDF with Chrome Headless
For a repeatable command-line capture, Chrome Headless supports --print-to-pdf. For example, with Chrome installed and available as google-chrome:
google-chrome --headless --print-to-pdf=archive.pdf "https://example.com/page"
On systems where the executable is named chromium or chromium-browser, use that command name instead. The output path can be changed to an absolute path if you want to control where the PDF is saved.
You can set a maximum wait with --timeout, in milliseconds:
google-chrome --headless --timeout=10000 --print-to-pdf=archive.pdf "https://example.com/page"
The timeout is an upper bound on waiting before capture, not evidence that a specific component has finished rendering. A slow network request, client-side rendering, or lazy content can still leave the PDF incomplete. Open the resulting file and verify the intended content.
Chrome’s Headless command-line reference describes PDF output and timeout behavior. Browser print output also uses paged-media layout. Print CSS, pagination rules, and browser support can affect the result; see Chrome’s discussion of printed page margins and paged media.
3. Understand what Shadow DOM changes
A component can attach a shadow tree to a host element. The tree encapsulates its internals: ordinary selectors such as document.querySelectorAll() and document-level CSS do not simply cross the shadow boundary. This is why a script or stylesheet that works on the rest of a page may not inspect or modify a component’s internal markup.
For an open shadow root, page JavaScript can access it through the host element’s shadowRoot property. A closed root is not exposed through that property. Chrome has an openOrClosedShadowRoot() API for extension code, but that specialized extension API is not a general page-script method for inspecting closed roots.
These access limits are distinct from what the browser renders. Shadow DOM content may be visible on screen even though ordinary document scripts cannot select it. The browser’s print pipeline then lays out the page for paper, and the result can differ from the screen. Inspect the print preview and saved PDF rather than assuming that screen appearance guarantees identical output.
For the platform details, see MDN’s guide to using Shadow DOM and Chrome’s extension DOM API reference.
4. If you own the component or page
If the page is yours and its component content is missing or unsuitable in print, add print behavior where the component is defined. Because component styles are encapsulated, a page-level print stylesheet may not be enough to style internals. Adjust the component’s own print styles and test the result in the browsers your users rely on.
- Check whether the component needs a print-specific layout, colors, or visibility rules.
- Make sure important content is not only available after an interaction that does not occur during printing.
- Check long content across page breaks, including headings, tables, and component boundaries.
- Test the actual saved PDF, not only the screen or print preview.
There is no single CSS change that can be prescribed for every component: the fix depends on its markup, styles, and how the page loads its content. MDN describes how component styles live inside the shadow tree in its Shadow DOM guide.
5. Choose a capture route
| Route | Best for | What to check |
|---|---|---|
| Browser Print → Save to PDF | A one-time archive | Component content, page selection, margins, orientation, and backgrounds |
Chrome Headless --print-to-pdf |
Repeatable command-line capture | Page readiness, timeout behavior, print CSS, and the generated file |
| Changes to the page or component | A page you control | Print styles inside the component and output in target browsers |
Choose based on whether the task is one-off or repeatable and whether you control the page. No cited source establishes a universally best browser for printing Shadow DOM pages.
6. Troubleshoot missing content and poor output
| Symptom | Likely cause | What to try |
|---|---|---|
| A component is missing in the PDF | The page may not have finished rendering, or its print behavior may omit or hide content. | Wait for the page to settle, check preview again, try the site’s export feature or another browser, and inspect the saved PDF. If you own the component, review its print styles. |
| The PDF has blank space or awkward page breaks | Print layout and pagination differ from the screen layout. | Check orientation and margins, and test a different print route. For a page you control, adjust its print CSS and verify pagination. |
| Background colors or images are absent | The print settings may omit backgrounds. | Enable print backgrounds if the browser offers that setting, then inspect the saved PDF. |
| Headless Chrome captures stale or incomplete content | The fixed timeout elapsed before client-side or lazy content was ready. | Increase the wait only as a practical measure, and verify output. The timeout does not guarantee readiness; if you need a page-specific readiness condition, use an automation setup that can wait for it before printing. |
| A page-level selector or style has no effect inside a component | Shadow DOM encapsulates component internals. | For an open root, inspect through the host’s shadowRoot where appropriate. If you own the component, change its internal styles. A closed root is not exposed through the ordinary shadowRoot property. |
| PDF looks different from the on-screen page | Printing uses paged layout and print-specific behavior. | Review print preview, margins, orientation, and the page’s print styles. Compare another supported print path if needed. |
7. Or skip the browser setup
If you need a clean screenshot of the page rather than a PDF archive, ScreenshotNeo is a website screenshot API and MCP server for developers. A single request returns an image or PDF, and its PDF options include paper size, margins, orientation, and page ranges. See the API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/page -o archive.pdf
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/page", "format": "pdf"}, timeout=90)
r.raise_for_status()
open("archive.pdf", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/page', format: 'pdf' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await Bun.write('archive.pdf', bytes);
The supplied Node.js example uses Bun to write the response bytes to a file. With Node.js, replace the final two lines with:
import { writeFile } from 'node:fs/promises';
const bytes = Buffer.from(await res.arrayBuffer());
await writeFile('archive.pdf', bytes);
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture, and each of those steps can be turned off. Bot checks or CAPTCHAs, blank pages, failed loads, timeouts, and cache hits are not billed; response headers report the page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Every feature is available on every plan. These are capture options, not a guarantee that a particular site’s Shadow DOM will print exactly as it appears on screen, so inspect the PDF output.
Start with 1,000 free screenshots a month, no card required.
8. Performance, reliability, and cost
Browser printing avoids a per-request screenshot API charge, but automated capture still depends on having a browser installed and deciding when the page is ready. A fixed headless timeout can bound the wait, yet a short timeout risks incomplete content and a long one can delay a batch. Inspect sample output before relying on an automated archive job.
Printing can change pagination, margins, and styling, so verify important documents and retain the source URL and capture date alongside the archive if later identification matters. The reviewed browser sources do not promise preservation of interactive state, animations, video, lazy-loaded content, or every component’s screen styling.
For API capture costs, ScreenshotNeo’s published plans are Free: 1,000 shots per month; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; and Business: $249 for 1,000,000. Yearly billing gives two months free. Each response includes verdict and billing headers, so you can distinguish a clean capture from a failed or non-billable result. Use the API when its capture workflow and controls fit the job; for a one-off PDF, the browser’s built-in print command may be sufficient.
FAQ
Can I save a closed Shadow DOM component as a PDF?
Closed roots are not accessible through the host’s ordinary shadowRoot property. Try the browser’s print output and inspect it; if you own the component, adjust its own print behavior.
Does Chrome’s headless timeout wait for Shadow DOM to finish?
No. It sets a maximum wait before capture. It does not confirm that a particular page’s asynchronous components are ready.
Will the PDF preserve everything I can see and interact with?
Not necessarily. Print output uses paged layout, and the reviewed sources do not guarantee preservation of interactive state or every screen effect. Check the generated file.
Can I fix a component’s print style from ordinary page CSS?
Not reliably across the shadow boundary. If you maintain the component, put the needed behavior in its own styles and test the PDF in your target browsers.


