ScreenshotNeo

BlogHow-to

How to Find and Follow Links with Puppeteer

Learn how to extract link URLs, click links reliably, and wait for navigation with Puppeteer, including single-page app and selector edge cases.

By the ScreenshotNeo team4 October 20267 min read

Use page.$$eval('a', anchors => anchors.map(anchor => anchor.href)) to collect links, and use a Puppeteer locator such as page.locator('a.next-page').click() to follow one. If the click may navigate, start page.waitForNavigation() and the click together with Promise.all() so the navigation wait is already active when the click happens.

These are different jobs: extraction reads destination URLs without interacting with the page; clicking activates a link as a browser user would. Puppeteer recommends locators for interaction because they wait for an element and check action preconditions. See the Puppeteer page interactions guide.

1. Set up a runnable Puppeteer script

Install Puppeteer in a project, then save this as links.js. The example opens a page, prints its links, follows a known link, and reports the resulting URL.

npm install puppeteer
const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch({ headless: true });
  try {
    const page = await browser.newPage();
    await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });

    const links = await page.$$eval('a', anchors =>
      anchors.map(anchor => ({
        text: anchor.innerText.trim(),
        href: anchor.href,
      }))
    );
    console.log('Links:', links);

    // Replace this selector with a link that exists on the page.
    const selector = 'a.more-information';
    const [response] = await Promise.all([
      page.waitForNavigation(),
      page.locator(selector).click(),
    ]);

    console.log('Current URL:', page.url());
    console.log('Navigation response:', response);
  } finally {
    await browser.close();
  }
})().catch(error => {
  console.error(error);
  process.exitCode = 1;
});

The selector in this example is deliberately a placeholder; replace it with a selector that matches the target page. Puppeteer APIs can change between versions, so check the official guide and API reference for the version installed in your project.

Use $eval for one matching element and $$eval for a collection. The functions run against the page’s DOM and return serializable values such as strings and objects.

const href = await page.$eval('a', anchor => anchor.href);
console.log(href);

page.$eval() passes the first match to the callback and throws if nothing matches. See the Puppeteer $eval API reference.

const links = await page.$$eval('a', anchors =>
  anchors.map(anchor => ({
    text: anchor.innerText.trim(),
    href: anchor.href,
  }))
);

The anchor.href DOM property returns the browser-resolved URL. For example, a relative attribute such as /docs is resolved against the page’s base URL. If you need the literal markup value instead, read anchor.getAttribute('href'); that can return a relative URL or null.

$$eval returns an empty array when there are no matches, making it suitable for optional collections. For one optional link, use a query and check for a missing result:

const href = await page.evaluate(() => {
  const anchor = document.querySelector('a.optional-link');
  return anchor ? anchor.href : null;
});
if (href === null) {
  console.log('No matching link found');
}

For normal interaction, use a locator. Locators wait for the target to be available and check conditions such as visibility, enabled state, viewport position, and a stable bounding box before acting.

await page.locator('a.next-page').click();

Use a selector that identifies the intended link uniquely and remains stable. Puppeteer supports CSS selectors by default and documents additional selector syntax for text, accessibility, XPath, and shadow-root traversal. When selecting by visible wording or accessible name is more reliable than a styling class, use the appropriate documented selector syntax and verify it against the page.

Click and wait for a document navigation

Start the navigation wait before the click can trigger it. Running both promises concurrently avoids the race where a fast navigation begins before the wait is listening.

const [response] = await Promise.all([
  page.waitForNavigation(),
  page.locator('a.next-page').click(),
]);

console.log('URL after click:', page.url());
console.log('Response:', response);

The response can be null, including when the URL changes through a hash or the History API. Puppeteer treats History API URL changes as navigation, but a navigation wait does not promise that a new document response exists. See the Puppeteer waitForNavigation reference.

Wait for the state that matters in a single-page app

In a single-page app, clicking may update application content without loading a new document. If success means that a particular view appeared, wait for that view rather than assuming a response object will be present:

await page.locator('a.account').click();
await page.locator('[data-page="account"] h1').wait();
console.log('Account view is visible at:', page.url());

Choose a target element or URL condition that represents the outcome you need. A click completing only means the click action completed; it does not prove the application reached the intended state.

4. Choose the right selector and interaction API

Need Use Behavior to account for
Read one destination page.$eval(selector, fn) Uses the first match and throws if none exists.
Read a list of destinations page.$$eval(selector, fn) Returns values from all matches; map to strings or plain objects.
Activate a link page.locator(selector).click() Recommended interaction path; waits for readiness and action preconditions.
Use a lower-level click page.click(selector) Clicks the first match after scrolling it into view; throws if no match exists.
Wait for a visible target Locator wait or a selector wait waitForSelector() is lower-level and does not automatically retry a later action that fails.

For links inside an iframe, first select the corresponding frame and query or click within that frame. For links inside a shadow root, use Puppeteer’s documented shadow-root selector syntax where applicable. A selector that matches multiple links should be narrowed by a stable parent, link text, accessible name, or other page structure so the intended destination is clear.

5. Common errors and fixes

Symptom Likely cause Fix
No node found for selector or a locator timeout The link is absent, rendered later, or the selector is wrong. Confirm the selector in the page DOM, wait for the relevant content, and check whether the link appears after an interaction.
The wrong link is clicked The selector matches several anchors and the first one is not the target. Narrow the selector to a unique link or stable containing section; inspect matches before clicking.
Navigation wait times out The click caused an in-page update, did not activate the link, or the page did not navigate. For SPA changes, wait for the destination element or expected URL condition. Confirm that the click target is actionable.
The navigation response is null The URL changed through a hash or History API without a new document response. Check page.url() or wait for the resulting page state; do not require a response when no document was loaded.
The link cannot be found in the main page It is inside a frame or shadow root. Query the correct frame, or use supported shadow-root selector syntax.
Extracted URL differs from the literal href anchor.href resolves relative URLs using the page’s base URL. Use getAttribute('href') when the raw attribute is what you need.
Click finishes but expected content is missing The click action succeeded, but the app state or destination did not meet the task’s success condition. Wait for a meaningful destination element or URL and handle the case where it never appears.

6. Performance, reliability, and operating cost

Extract only the fields you need. Returning resolved URL strings and link text keeps the result small and avoids transferring page nodes or handles into Node.js. If the page has many anchors, narrow the selector to the relevant navigation region instead of mapping every link.

For reliable clicks, use locators and wait for the outcome that defines success. Avoid arbitrary delays when an element or URL condition can express readiness directly. Always close the browser in a finally block so failures do not leave Chromium processes running.

Browser automation consumes the resources needed to run Chromium and load the target page. Page weight, scripts, network conditions, and the chosen load condition affect run time and resource use. Puppeteer itself does not set a per-screenshot API price in this workflow; account for your own browser hosting and execution costs.

7. Or skip the browser setup

If the task is to capture a page rather than click through its links, ScreenshotNeo returns a screenshot or PDF from one GET request. Its API documentation covers the available parameters.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const fs = require('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

Cookie banners are accepted and removed before capture, along with known newsletter popups and chat widgets. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Sign up for 1,000 free screenshots a month, with no card.

8. FAQ

No. anchor.href returns the browser-resolved URL. Read getAttribute('href') for the literal attribute value.

The examples above handle the current page. If the target opens a new page, listen for the browser’s new-page event and work with that page; the current page’s navigation wait does not represent a separate tab.

Should I use page.click() or a locator?

Use a locator for ordinary interactions because Puppeteer recommends it and it checks readiness. Use lower-level selector or element-handle flows when your automation needs their specific control.

Why can a URL change without a navigation response?

Hash and History API changes can update the URL without loading a new document. Check the URL or the application state that matters to your workflow.