How to Fetch Every Page Title with Pyppeteer
Use Pyppeteer’s async title API to read one page, collect titles from open tabs, handle dynamic documents, and troubleshoot common failures.
Use await page.title() after navigation. To read titles from every visible page currently open in one Pyppeteer browser, call await browser.pages() and then read each page’s title. Pyppeteer’s page enumeration does not include every possible browser target, such as some background pages.
Read one page title
Install Pyppeteer in a Python 3.8-or-newer environment, then navigate before reading the title:
python -m pip install pyppeteer
import asyncio
from pyppeteer import launch
async def main():
browser = await launch()
try:
page = await browser.newPage()
await page.goto("https://example.com")
title = await page.title()
print(title)
finally:
await browser.close()
asyncio.run(main())
page.title() is asynchronous and returns the document title as a string. If the document has no title, the result can be an empty string.
Fetch titles from every open page
If your script has several tabs or pages in the same browser, enumerate them and gather their titles concurrently:
import asyncio
from pyppeteer import launch
async def main():
browser = await launch()
try:
first = await browser.newPage()
second = await browser.newPage()
await asyncio.gather(
first.goto("https://example.com"),
second.goto("https://www.python.org"),
)
pages = await browser.pages()
titles = await asyncio.gather(*(page.title() for page in pages))
for page, title in zip(pages, titles):
print(f"{page.url}\t{title}")
finally:
await browser.close()
asyncio.run(main())
browser.pages() returns page objects known to that browser instance. It describes visible pages, so it is not a universal inventory of tabs, workers, extensions, or background targets.
Fetch titles for a list of URLs
For a crawler-like job, create one page per URL, navigate each page, read its title, and close the page when finished. A semaphore prevents an unbounded number of Chromium tabs:
import asyncio
from pyppeteer import launch
URLS = [
"https://example.com",
"https://www.python.org",
"https://www.djangoproject.com/",
]
async def title_for(browser, url, limit):
async with limit:
page = await browser.newPage()
try:
await page.goto(url)
return {"url": page.url, "title": await page.title(), "error": None}
except Exception as exc:
return {"url": url, "title": None, "error": str(exc)}
finally:
await page.close()
async def main():
browser = await launch()
try:
limit = asyncio.Semaphore(4)
results = await asyncio.gather(
*(title_for(browser, url, limit) for url in URLS)
)
for result in results:
print(result)
finally:
await browser.close()
asyncio.run(main())
Use a concurrency limit based on available memory and the sites’ capacity. More tabs do not guarantee higher throughput.
Handle titles that change after navigation
Some applications replace the title after JavaScript runs. A title read immediately after goto() can therefore be an earlier value. Wait for a condition that matches the site you are automating:
await page.goto("https://example.com")
await page.waitForFunction(
"document.title && document.title !== 'Loading…'",
{"timeout": 10000},
)
title = await page.title()
The correct condition is application-specific. If the site changes its title after an API response, wait for a selector or another observable state associated with that response rather than choosing an arbitrary delay.
You can inspect the DOM directly when you need custom title-like data:
title = await page.evaluate("document.title")
page.title() remains the simpler choice for the normal document title. When passing an expression to evaluate(), Pyppeteer’s README documents force_expr=True for cases where an expression is interpreted as a function.
Choose navigation and title options deliberately
| Need | Approach |
|---|---|
| Wait for the initial document | Call await page.goto(url), then await page.title(). |
| Wait for a known application state | Use waitForSelector() or a site-specific waitForFunction(). |
| Collect open tabs | Call await browser.pages(); remember its visibility limitation. |
| Preserve page-to-title mapping | Keep the page objects and titles together, as in zip(pages, titles). |
| Read a title without navigation | Call await page.title() on an already loaded page. |
Pass navigation settings supported by the Pyppeteer version you installed when a site needs a particular wait condition or timeout. Check the installed version’s API reference rather than assuming options from another browser automation library.
Common errors and fixes
| Error or symptom | Likely cause | Fix |
|---|---|---|
ModuleNotFoundError: pyppeteer |
The package is not installed in the active environment. | Run python -m pip install pyppeteer with the same Python interpreter that runs the script. |
| Chromium download or launch failure | No suitable browser binary is available, or the runtime cannot download one. | The project README documents pyppeteer-install for downloading Chromium in advance. You can also provide an installed Chrome binary through the launch configuration. |
TimeoutError during goto() |
The server is slow, unavailable, or never reaches the expected navigation state. | Check the URL, network access, redirects, and the navigation timeout. Record the URL and exception, then continue or retry according to your job’s policy. |
| Title is empty | The document has no <title>, or the title is assigned later by JavaScript. |
Confirm the DOM, then wait for a site-specific condition before calling title(). |
| Title is “Loading…” or another placeholder | The application updates the title after an asynchronous operation. | Wait for the relevant selector, state, or title predicate instead of using a universal sleep. |
| Fewer pages than expected | browser.pages() reports page objects, not every browser target or background page. |
Track pages your script creates. If you need other target types, inspect the browser’s target APIs for your installed version. |
| Browser processes remain after failure | An exception occurred before cleanup. | Put browser shutdown in a finally block and close temporary pages in their own finally blocks. |
Performance and reliability
- Reuse one browser: Launching Chromium for every URL adds startup overhead. Keep one browser open and create or reuse pages.
- Bound concurrency: A semaphore limits memory use and avoids overwhelming the destination.
- Close pages: Always close pages created for a single URL, even when navigation fails.
- Keep results structured: Store the requested URL, final URL, title, and error so redirects and failures are diagnosable.
- Retry selectively: A retry can help with transient network failures, but repeated retries against a consistently failing URL waste time. Use a maximum attempt count and backoff.
- Expect site differences: Consent dialogs, bot checks, client-side routing, and delayed title updates can change what a browser sees. Your wait condition must reflect the target application.
Pyppeteer’s current project README describes the project as unmaintained and points readers toward Playwright Python for new work. Existing Pyppeteer scripts can continue to use the APIs shown here, but pin and review the version used in production.
Or skip the browser setup
For a rendered screenshot rather than a title-only crawler, ScreenshotNeo provides a single GET request. Its capture flow accepts cookie and consent banners before the shot and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response reports the result in X-Page-Verdict and X-Billed headers.
See the ScreenshotNeo API documentation for options such as full-page capture, element selectors, waits, custom JavaScript and CSS, device presets, PDFs, caching, signed links, asynchronous jobs, bulk capture, and usage reporting.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also includes an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
FAQ
Does browser.pages() find every tab?
It returns the browser’s page objects and excludes certain non-visible pages. It is not a complete list of every browser target.
Can I call page.title() before goto()?
You can call it on an existing page, but before navigation it will describe that page’s current document, which may be blank or a previous URL.
Why is my title different from the text in the page header?
page.title() reads the HTML document’s <title>. A visible heading such as <h1> is separate content and must be selected or evaluated independently.
How do I process thousands of URLs?
Reuse a browser, cap concurrent pages, close each page, record failures, and apply bounded retries. For screenshot workloads, a managed API can remove Chromium installation and cleanup from your worker.


