How to Extract Content from a Shadow DOM
Learn to extract text and markup from open Shadow DOM roots with browser JavaScript, Playwright, and Selenium, and diagnose closed roots and timing issues.

To extract content from an open Shadow DOM, first find its custom-element host, then query inside host.shadowRoot. A document-level document.querySelector() does not cross into a shadow tree. If the root is closed, ordinary page JavaScript cannot read it through host.shadowRoot; use a suitable automation, extension, or DevTools Protocol context instead. [MDN: Using shadow DOM]
1. Understand the host, root, and query scope
A web component’s custom element is its host. The host may own a shadow root containing the component’s internal elements. That root is a separate query scope: find the host from the document, then find descendants from the root.

const host = document.querySelector('my-component');
const root = host?.shadowRoot;
const target = root?.querySelector('.target');
const text = target?.textContent?.trim();
console.log(text);
This pattern returns undefined when the host or target is missing, or when the root is closed. Check those cases explicitly in reusable code:
function readShadowText(hostSelector, targetSelector) {
const host = document.querySelector(hostSelector);
if (!host) throw new Error(`Host not found: ${hostSelector}`);
const root = host.shadowRoot;
if (!root) throw new Error('No accessible open shadow root found');
const target = root.querySelector(targetSelector);
if (!target) throw new Error(`Target not found: ${targetSelector}`);
return target.textContent.trim();
}
console.log(readShadowText('my-component', '.target'));
textContent reads descendant text, including text that may not be visibly rendered. If you need only visible, user-facing content, consider the element’s rendered state and use a browser automation locator that reflects visibility. To inspect markup, use innerHTML for the contents of a node or outerHTML for that node’s serialization. These are serializations, not a complete representation of runtime state, event listeners, or every nested shadow tree.
2. Find the component in DevTools
- Open the page in Chrome and inspect the visible content with DevTools’ element picker.
- Identify the custom-element host around the content. It may have a tag name such as
user-cardormy-component. - With the relevant node selected in the Elements panel, switch to the Console. Chrome exposes that selected node as
$0. - Check the root and query from it.
$0.shadowRoot
$0.shadowRoot?.innerHTML
$0.shadowRoot?.querySelector('.target')?.textContent
DevTools displays shadow roots in the DOM tree, which helps identify the host and nested structure. The Console expression is still ordinary page JavaScript: it can read an open root, while a closed root yields null. [Chrome DevTools: View and change the DOM]
3. Handle nested shadow roots
A selector does not jump through multiple shadow boundaries. For nested components, locate each inner host from the current root, then enter that host’s root:

const outerHost = document.querySelector('account-panel');
const outerRoot = outerHost?.shadowRoot;
const innerHost = outerRoot?.querySelector('profile-card');
const innerRoot = innerHost?.shadowRoot;
const name = innerRoot?.querySelector('.name')?.textContent?.trim();
console.log(name);
For a longer chain, make traversal explicit so failures identify the level that was not found:
function findInOpenShadowRoots(start, selectors) {
let scope = start;
for (const selector of selectors) {
const match = scope.querySelector(selector);
if (!match) return null;
scope = match.shadowRoot ?? match;
}
return scope;
}
const result = findInOpenShadowRoots(document, [
'account-panel',
'profile-card',
'.name'
]);
console.log(result?.textContent?.trim());
For chains where an intermediate selector identifies a normal element rather than a component host, keep the distinction in mind: the next query runs on that element, while crossing to another component’s internals requires entering its shadowRoot.
4. Extract content with Playwright
Playwright locators pierce open Shadow DOM by default, so a normal CSS locator can target an element inside an open root. XPath locators do not pierce shadow roots, and closed-mode roots are unsupported. [Playwright: Locators]
Runnable Node.js example (install with npm install playwright; install the browser with npx playwright install chromium):
const { chromium } = require('playwright');
(async () => {
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage();
try {
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
const value = await page.locator('my-component .target').textContent();
console.log(value?.trim());
} finally {
await browser.close();
}
})();
Replace the URL and selectors with the target page and component. If the component renders asynchronously, wait for the target locator:
const target = page.locator('my-component .target');
await target.waitFor({ state: 'visible', timeout: 10000 });
console.log((await target.textContent())?.trim());
Use role or text locators when they describe the user-visible element better than a CSS selector. Avoid XPath for shadow-tree content. Playwright handles open-root traversal; for application-specific nested components, a scoped locator may make the intended host relationship easier to verify.
5. Extract content with Selenium JavaScript
Selenium exposes a ShadowRoot search context with findElement and findElements. Get the host element, obtain its shadow root, and search within that root. Check the API reference for the methods supported by your language binding and Selenium version. [Selenium JavaScript: ShadowRoot]
Example for Selenium’s JavaScript API (install with npm install selenium-webdriver; run with a compatible Chrome and driver available):
const { Builder, By } = require('selenium-webdriver');
(async () => {
const driver = await new Builder().forBrowser('chrome').build();
try {
await driver.get('https://example.com');
const host = await driver.findElement(By.css('my-component'));
const root = await host.getShadowRoot();
const target = await root.findElement(By.css('.target'));
console.log((await target.getText()).trim());
} finally {
await driver.quit();
}
})();
For nested roots, repeat the host lookup and getShadowRoot() sequence from the current shadow search context. If your installed binding does not expose the same method names, use its version-specific API documentation rather than assuming the JavaScript interface applies.
6. Choose the right output
| Need | Approach | Watch for |
|---|---|---|
| Plain text | textContent in page JavaScript; locator text in automation |
Text may include hidden descendants or whitespace; trim or normalize for your use. |
| Element attributes | Find the element in its root, then read the attribute | Missing attributes return no value; distinguish that from an empty string if it matters. |
| Markup of one element | outerHTML |
Does not include the element’s own shadow tree as a complete recursive serialization. |
| Markup inside an open root | shadowRoot.innerHTML |
Nested shadow roots are separate; serialize or traverse each explicitly. |
| Automation target | Playwright locator or Selenium ShadowRoot search context | Use APIs that cross open roots; Playwright XPath does not. |
| Shadow-inclusive protocol markup | Chrome DevTools Protocol DOM serialization | This is a protocol-level method, not a standard page script. |
For Chrome DevTools Protocol users, the DOM getOuterHTML method has an includeShadowDOM option for including shadow roots in returned markup. Treat it as a protocol workflow: attach to the relevant target, resolve the node, and call the protocol method. The exact setup depends on how you connect to CDP. [Chrome DevTools Protocol: DOM]
7. Open roots, closed roots, and execution context
For an open root, page JavaScript can access host.shadowRoot. For a closed root, that property returns null; ordinary page code cannot use the same traversal pattern. Closed mode is an encapsulation boundary, though MDN cautions that it should not be treated as a strong security mechanism. [MDN: Shadow DOM]
Do not confuse privileged inspection paths with page JavaScript. Chrome extensions can use chrome.dom.openOrClosedShadowRoot(element), documented as available from Chrome 88. This is an extension API, not a standard web-page method. CDP also provides shadow-inclusive serialization. Choose based on where your code runs:
- Page script or console: open roots through
shadowRoot. - Playwright: locators traverse open roots by default; closed roots are unsupported.
- Selenium: use the binding’s shadow-root search context for accessible roots.
- Chrome extension: the documented extension API can access open or closed roots.
- CDP client: use protocol operations such as shadow-inclusive
getOuterHTML.
Sources: Chrome extension DOM API and CDP DOM API.
8. Wait for rendering without guessing
A host or its content may not exist yet when your script runs. A null result can mean the selector missed, rendering has not completed, or the root is closed. Diagnose the host and root separately before changing selectors.
In a browser console, rerun the query after the component appears. In Playwright, wait for a locator with a bounded timeout. In Selenium, use the wait facility from your binding to wait for the host and then its shadow descendant. Avoid fixed sleeps as the default: they can waste time on fast pages and still be too short on slow ones. If a page replaces the component during navigation or interaction, reacquire the host and root instead of assuming an earlier element handle remains valid.
9. Troubleshooting checklist
| Symptom | Likely cause | Fix |
|---|---|---|
document.querySelector('.target') returns null |
The target is inside a shadow tree. | Find the host, then query from host.shadowRoot. |
host is null |
The host selector is wrong or the component has not rendered. | Inspect the DOM in DevTools, confirm the custom-element tag, and wait for rendering. |
host.shadowRoot is null |
The root is closed, the root is not attached yet, or the selected node is not the expected host. | Verify the host and timing; if closed, use an execution context with an applicable extension API or CDP. |
| Playwright CSS locator times out | Wrong selector, missing host, late content, or hidden target. | Inspect the page and wait for the correct target state; confirm the root is open. |
| Playwright XPath finds nothing | XPath does not pierce shadow roots. | Use a CSS, role, or text locator supported for the target. |
| Selenium cannot find a descendant from the driver | The search is still scoped to the document. | Get the host’s shadow root and search within that context. |
| Serialized HTML omits nested component internals | Nested roots are separate from ordinary element markup. | Traverse into each nested host or use CDP’s shadow-inclusive serialization. |
| Text is empty or unexpected | The selected node may be a wrapper, content may not have rendered, or text includes whitespace/hidden descendants. | Inspect the exact node, wait for content, and choose/normalize the desired text. |
10. Performance, reliability, and operating cost
For a single extraction, direct page JavaScript has little setup: locate one host and query its root. Automation adds browser startup and page navigation, so reuse a browser process when collecting from multiple pages and close contexts reliably. Scope selectors to the relevant host/root to make intent clear and avoid repeatedly querying an entire document.
Reliability depends on the target component’s structure and rendering behavior. Custom element names and internal class names can change; prefer stable attributes, roles, or user-visible text where appropriate. Wait for a meaningful condition, use a timeout, and report whether the host, root, or target was missing. Closed-root access depends on execution context and browser tooling, so choose that context before designing the extraction flow.
Browser automation has compute and maintenance costs: browser binaries, execution time, concurrency, and keeping automation dependencies compatible. The dossier provides no universal runtime or cost benchmark, so measure against the pages and infrastructure you actually use. For screenshot capture rather than DOM text extraction, ScreenshotNeo offers a one-request API; screenshots do not expose the underlying Shadow DOM text or markup.
11. Or skip the browser setup
If your goal is a visual record of a page rather than extracting its DOM text or markup, ScreenshotNeo captures a URL as an image or PDF. It does not replace Shadow DOM inspection when you need text or HTML. Its API accepts a URL in one GET request; see the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await require('node:fs/promises').writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
Cookie banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off. Bot checks, blank pages, and failed loads are never billed, and response headers report the page verdict and billing status. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots.
Sign up free for 1,000 screenshots a month, no card required.
12. Frequently asked questions
Can I get a ShadowRoot from an element found inside another root?
Yes, if that element is itself a component host with an open root. Find it from the current root, then read its shadowRoot.
Does Shadow DOM hide content from screenshots?
Shadow DOM changes DOM query scope; it does not by itself mean the rendered content is absent from a visual capture. A screenshot records rendered pixels, not accessible markup or extracted text.
Can I use a CSS selector that includes the whole nested path?
Not with ordinary document querying across boundaries. Traverse each host/root step, or use an automation locator that supports open-root traversal.
Is closed Shadow DOM secure storage?
No. Closed mode limits ordinary access through the host’s standard property, but it should not be treated as a security boundary for sensitive data.


