How to Find Text on a Page with Playwright
Use Playwright’s modern text locators to match, scope, read, and assert page text reliably—even inside iframes and dynamic UIs.
Use page.getByText() to find non-interactive text in Playwright. It supports substring matching by default, exact matching with { exact: true }, and regular expressions. For buttons and links, prefer getByRole(); for repeated text, scope the locator with a parent or filter({ hasText }). Use web-first assertions such as toHaveText() and toContainText() so Playwright waits for dynamic content. For an iframe, use frameLocator(...).getByText().
This guide covers matching rules, complete TypeScript examples, text extraction, lists and cards, shadowed or dynamic content, iframes, troubleshooting, performance, and a browser-free ScreenshotNeo option.
1. Find visible text with getByText()
The modern text locator is documented in the Playwright locator guide. A substring match is the default:
import { test, expect } from '@playwright/test';
test('find text', async ({ page }) => {
await page.goto('https://example.com');
await expect(page.getByText('Example Domain')).toBeVisible();
});
getByText('Example Domain') can match an element whose text contains that phrase. Locators are resolved when an action or assertion runs, so they work with content that appears after navigation or an API request.
Substring, exact, and regular-expression matches
await expect(page.getByText('Welcome, John')).toBeVisible();
await expect(
page.getByText('Welcome, John', { exact: true })
).toBeVisible();
await expect(
page.getByText(/welcome, [A-Z a-z]+$/i)
).toBeVisible();
| Form | Use it when |
|---|---|
getByText('Saved') |
You want a substring match. |
getByText('Saved', { exact: true }) |
The whole normalized text must equal the string. |
getByText(/saved|complete/i) |
The text varies, or you need case-insensitive or anchored matching. |
2. How Playwright normalizes text
Text matching normalizes whitespace, including line breaks and leading or trailing spaces. This means markup such as:
<div>
Welcome
back
</div>
can be found with:
await expect(page.getByText('Welcome back')).toBeVisible();
With exact: true, Playwright still trims and normalizes whitespace before comparing. If text is split across several elements, target a suitable ancestor or use a role locator instead of assuming one text node contains the entire sentence.
3. Use roles for buttons, links, and other controls
Text locators are mainly for non-interactive content. For an interactive element, a role and accessible name usually describe the user-facing contract more reliably than incidental visible text.
await page.getByRole('button', { name: 'Sign in' }).click();
await expect(page.getByText('Welcome, John!')).toBeVisible();
This approach continues to work if the button’s internal markup changes while its accessible name remains the same. See the official locator guide for role, label, placeholder, and text locators.
Common role examples
await page.getByRole('link', { name: 'Documentation' }).click();
await page.getByRole('heading', { name: 'Settings' }).isVisible();
await page.getByRole('checkbox', { name: 'Receive updates' }).check();
await page.getByRole('textbox', { name: 'Email' }).fill('dev@example.com');
4. Disambiguate repeated text with scope and filters
A page can contain the same product name, status, or action in several cards. First locate the container, then find text or a control inside it.
const product = page
.getByRole('listitem')
.filter({ hasText: 'Product 2' });
await expect(product).toHaveCount(1);
await product.getByRole('button', { name: 'Add to cart' }).click();
filter({ hasText }) keeps the match tied to the intended container. You can also chain locators:
const settings = page.getByRole('region', { name: 'Account settings' });
await expect(settings.getByText('Two-factor authentication')).toBeVisible();
When multiple matches are expected
If a locator intentionally matches several elements, assert the count or inspect all values rather than using first() without a reason:
const statuses = page.getByText('Pending');
await expect(statuses).toHaveCount(3);
await expect(statuses.nth(0)).toBeVisible();
first(), last(), and nth() are useful when order is part of the UI contract. A semantic container or unique accessible name is usually less fragile.
5. Assert text on dynamic pages
Use web-first assertions instead of reading text and comparing it manually. Playwright assertions retry until they pass or the assertion timeout is reached, as described in the assertions guide.
Exact text with toHaveText()
await expect(page.locator('.title')).toHaveText('Dashboard');
await expect(page.locator('.title')).toHaveText(/dashboard/i);
Substring text with toContainText()
await expect(page.locator('.status')).toContainText('Submitted');
Ordered lists of text
await expect(page.getByRole('listitem')).toHaveText([
'apple',
'banana',
'orange'
]);
An array checks the matching elements in order. Use a regular expression for values that contain IDs, timestamps, or other changing portions.
Configure a suitable timeout
await expect(page.getByText('Report ready')).toBeVisible({
timeout: 15_000
});
Prefer a condition tied to the UI state over an arbitrary sleep. A delay can make a test slower and still fail when the page takes longer than expected.
6. Read text when your code needs the value
Assertions are preferable for verification, but extraction is appropriate when you need to parse, log, or send a value elsewhere.
const links = await page.getByRole('link').allInnerTexts();
const raw = await page.locator('.message').textContent();
const rendered = await page.locator('.message').innerText();
const allRaw = await page.locator('.message').allTextContents();
| API | Returns |
|---|---|
innerText() |
Rendered, user-visible text according to layout and visibility rules. |
textContent() |
Raw text content, including text that may not be visibly rendered. |
allInnerTexts() |
An array of rendered text for every matched element. |
allTextContents() |
An array of raw text content values. |
If you only need to verify a value, prefer toHaveText() or toContainText() because those assertions automatically wait for the expected state.
7. Find text inside an iframe
Content in an iframe belongs to a separate document. Use frameLocator(), then call the same text or role locators on the frame.
const frame = page.frameLocator('#payment-frame');
await expect(frame.getByText('Card number')).toBeVisible();
await frame.getByRole('textbox', { name: 'Card number' }).fill('4242424242424242');
You can also select an iframe by title or another stable attribute:
const checkout = page.frameLocator('iframe[title="Checkout"]');
await expect(checkout.getByText('Pay now')).toBeVisible();
If the iframe is cross-origin, do not try to reach into it with page-level CSS. Use frameLocator() and ensure the frame has loaded before asserting its content.
8. A complete TypeScript example
import { test, expect } from '@playwright/test';
test('locate and verify an order', async ({ page }) => {
await page.goto('https://example.com/orders');
await page.getByRole('button', { name: 'Sign in' }).click();
await page.getByRole('textbox', { name: 'Email' }).fill('dev@example.com');
await page.getByRole('button', { name: 'Continue' }).click();
const order = page
.getByRole('listitem')
.filter({ hasText: 'Order #1007' });
await expect(order).toHaveCount(1);
await expect(order).toContainText('Shipped');
await expect(order.getByText('Order #1007', { exact: true })).toBeVisible();
});
Replace the example URL and selectors with your application’s accessible names and stable containers.
9. cURL, Python, and Node.js alternatives for page text
Playwright is the right tool when you need browser interaction, JavaScript execution, or assertions. For a simple server-side HTML response, an HTTP client plus an HTML parser can be faster, but it will not see content rendered only by browser JavaScript.
cURL
curl -L https://example.com
Python with Playwright
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page()
page.goto("https://example.com")
print(page.get_by_text("Example Domain").inner_text())
browser.close()
Node.js with Playwright
import { chromium } from 'playwright';
const browser = await chromium.launch();
const page = await browser.newPage();
await page.goto('https://example.com');
console.log(await page.getByText('Example Domain').innerText());
await browser.close();
10. Troubleshooting text locators
| Symptom | Likely cause | Fix |
|---|---|---|
| Strict mode violation | The locator matches multiple elements. | Use getByRole(), { exact: true }, a parent locator, or filter({ hasText }); assert the expected count. |
| Timeout waiting for text | The text is rendered later, differs from the expected value, or is in a frame. | Check the exact rendered text, wait on a meaningful state with a web-first assertion, and use frameLocator() for iframe content. |
| Text appears in the browser but is not found | The text is inside a shadow DOM, iframe, canvas, or a different page state. | Inspect the DOM context. Use a frame locator for iframes and a supported locator or component API for shadow DOM. Canvas pixels are not DOM text. |
| Exact match fails unexpectedly | Whitespace, hidden descendants, punctuation, or dynamic text differs. | Use toContainText(), a regular expression, or inspect innerText(); remember whitespace is normalized. |
| Button text locator is flaky | Visible text is incidental or duplicated. | Use getByRole('button', { name: ... }) and verify the accessible name. |
| Manual sleep still fails | Fixed delays do not model network or rendering completion. | Replace the sleep with a retrying assertion such as toBeVisible() or toHaveText(). |
Debug a locator
const message = page.getByText('Submitted');
console.log('matches:', await message.count());
console.log('text:', await message.allInnerTexts());
Run headed or use Playwright’s trace and inspector tools when you need to see the DOM state at the failure point.
11. Performance, reliability, and maintainability
- Prefer semantic locators. Roles, labels, and stable containers survive presentational markup changes better than long CSS or XPath selectors.
- Scope early. Searching inside a card, dialog, or list item reduces ambiguity and the amount of DOM Playwright must inspect.
- Use one assertion for one state. Small, meaningful assertions make failures easier to diagnose.
- Avoid unnecessary extraction. A retrying assertion is usually more reliable than a read followed by a language-level comparison.
- Keep timeouts proportional. Increase a targeted assertion timeout for a known slow operation instead of raising every test timeout.
- Wait for the UI contract. Assert the text or state your user needs, rather than sleeping for an estimated render time.
- Separate browser work from parsing. If the text is present in the original HTML, a direct HTTP request and parser may cost less than launching a browser. If JavaScript creates it, use Playwright.
12. Or skip the browser setup
If your goal is a clean page image rather than an interactive assertion, ScreenshotNeo returns a screenshot or PDF from one GET request. Its capture options include full-page shots, CSS-element capture, custom JavaScript and CSS, waits, device presets, and iframe-aware browser rendering.
With the API, cookie and consent banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are never billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. ScreenshotNeo also provides an MCP server so Claude, Cursor, and other MCP clients can call take_screenshot, get_page_info, and capture_pdf.
See the ScreenshotNeo API documentation for all options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', data));
There is a free tier of 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots, and every feature is included on every plan. Create a free ScreenshotNeo account.
13. FAQ
Is getByText() case-sensitive?
String matching is case-sensitive. Use a regular expression with the i flag when case should not matter.
Should I use CSS, XPath, or getByText()?
Use semantic locators first. Use CSS or XPath when you have a specific structural need that cannot be expressed through the user-facing role, label, or text.
Can getByText() find text hidden with CSS?
It can match DOM text, but visibility assertions such as toBeVisible() require the element to be rendered and visible. Use textContent() when you intentionally need raw hidden text.
How do I match text that includes a changing number?
Use a regular expression, for example page.getByText(/items: \\d+/), or assert a stable substring with toContainText().
Why does text in a canvas not match?
Canvas content is pixels rather than DOM text. Assert the application state or use an accessibility representation; a text locator cannot inspect pixels.


