How to Extract Paginated Search Results with Browse AI
Train a Browse AI robot to capture search-result lists across pages, load-more buttons, or infinite scroll, then check for missing and duplicate records.
To extract paginated search results with Browse AI, train a robot to reach the results page, choose Capture Text → From a list, select the repeating result cards or rows, pick the fields you need, and configure how the robot advances: click next, click load more, scroll down, or stop because all results are already visible. Run the robot, then check later-page records and compare the captured count with what you expected.
Browse AI describes “From a list” as suitable for repeating information such as product listings or search results. These steps follow its help guidance; they do not guarantee that every site exposes every field or that every run will capture every record. Browse AI: Capture text from a list.
1. Prepare the search-results page
- Identify the URL where the robot should start. If results require a search, use the site’s search field during training rather than assuming a direct results URL will work.
- Train the robot to focus the search field, enter the query, and submit it. If filters are needed, apply them as part of the trained interaction.
- Wait for the page to finish updating before selecting results. Browse AI’s search guide recommends waiting for loading indicators to disappear and for counts, filters, and result elements to render.
- If a control such as “Show all” or a filter toggle reveals the results or pagination, train that click first. Then capture the list after it appears.
Search-first extraction differs from starting at a known results URL: the robot must reproduce the search interaction and wait for its results. See Browse AI’s search-results guide.
2. Capture the repeating results as a list
- In Robot Studio, choose Capture Text → From a list.
- Hover over the repeating result cards or rows. Select the outlined group when it matches the results you intend to collect, rather than a surrounding container that includes unrelated page content.
- Give the list a descriptive name, such as “Search results.”
- Choose the number of items to capture. Set a limit that reflects the amount of data you need and the time you can spend reviewing it.
- Review the suggested dataset structure and customize it if the fields are missing or grouped incorrectly.
Depending on what the site visibly exposes, useful fields may include a title, description, URL, price, or rating. These are examples, not a promise that every result page contains them. Browse AI says the list capture automatically structures repeated items into a dataset that can be customized. Source.
3. Choose the pagination behavior
Match the robot’s pagination setting to the control or behavior the site actually uses:
| What the page does | Browse AI behavior | What to check |
|---|---|---|
| Shows a Next button, arrow, or numbered pages | Click next | Confirm the control advances to a new set of results and that the robot continues through the intended pages. |
| Has a “Load more” or “Show more” button that appends results | Click load more | Check that each click adds records and that the robot waits for them to render. |
| Loads additional results when you scroll, with no button | Scroll down / infinite scroll | Demonstrate the scroll and allow time for newly loaded content to appear. |
| All results are already visible | No more items | Capture the displayed list once and stop. |
These mappings follow Browse AI’s documented list-pagination choices. When there is an explicit load-more button, its guide recommends using that behavior rather than relying on infinite scroll. Source.
4. Save, run, and review the robot
- Save the captured list and finish the robot setup.
- Name the robot so its purpose and query are clear.
- Run it and review the resulting records, especially items from later pages.
- Compare the captured count with the expected result count. Check for missing pages, repeated records, and fields that are blank or attached to the wrong item.
- Adjust the list selection, item limit, wait behavior, or pagination interaction if the output does not match the page.
For repeated runs, treat the result as data that needs validation. Sites can change their layout or loading behavior, so a successful training run is not proof that future runs will remain complete.
5. Handle detail pages separately when needed
Pagination collects the list of search results. If you also need information that appears only after opening each result, list extraction alone does not visit those detail pages. Browse AI describes connecting robots in a workflow: one robot gathers result items or URLs, and another visits the detail pages. Its documented patterns include a site search followed by results and then a page. Search guide and deep-scraping guide.
Train and validate the detail-page step independently. Confirm that the links collected from the list are the intended destinations and that the detail robot extracts the fields you need.
6. Troubleshoot incomplete or inconsistent results
| Symptom | Likely cause | What to try |
|---|---|---|
| Only the first page appears | The robot is not using the site’s actual pagination interaction, or the list capture is not configured to continue. | Check whether the site uses a Next control, load-more button, or scrolling, then configure the matching behavior. |
| Items are missing after scrolling | New content may load asynchronously and was not rendered when the robot captured it. | Demonstrate scrolling and allow the new items time to appear before capture. Check whether an explicit load-more button is available. |
| The captured list includes navigation or unrelated content | The selected outline may include a broad page container instead of the repeating result rows. | Return to list selection and choose the repeated cards or rows whose outline matches the intended results. |
| Fields are missing or arranged incorrectly | The suggested dataset structure may not match the page’s visible layout. | Customize the dataset and use manual selection where needed; inspect several records, not only the first one. |
| Some runs take too long | The requested page or item limit may be larger than practical for one run. | Reduce the limit or divide the work into batches, then check that the batches cover the needed results. |
| Some records appear more than once | Repeated results may be exposed across pages or runs. | Check the page transitions and deduplicate the exported data using a stable field such as the result URL, where available. |
| Search results are empty or stale | The search interaction may not have completed before the list was captured. | Revisit the search steps and wait for loading indicators, result counts, filters, and result elements to update before capture. |
These are troubleshooting approaches described in Browse AI’s help material, not independently measured guarantees. If the site’s structure changes, revisit the training steps and validate a new run. Pagination troubleshooting.
7. Plan limits, batches, and data checks
- Set a realistic item limit. Capture only as many results as the task needs, then expand after validating a smaller run.
- Batch large collections. If a run is too slow, reduce its page or item limit and split the work. Keep track of which pages or ranges each batch covers.
- Check both completeness and duplicates. A plausible row count alone does not prove that every page was reached; inspect records from later pages and compare them with the expected range.
- Keep the query and filters reproducible. Record the search terms and filters used so that later runs can be compared meaningfully.
- Revalidate after site changes. Search pages and pagination controls can change. Recheck the robot when its output starts missing fields or pages.
The cited Browse AI guidance recommends reducing limits or splitting work into batches when runs are slow, and deduplicating when repeated records occur. The research available for this guide does not establish extraction accuracy, speed, or a guaranteed completion rate.
Or skip the browser setup
Browse AI extracts structured, repeated search-result data. If the job is to capture a visual record of a results page or a particular state, ScreenshotNeo is a website screenshot API and MCP server. A single GET request returns a PNG, JPEG, WebP, or PDF capture. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits are never billed; response headers say the page verdict and billing status.
- An MCP server gives AI agents tools to take screenshots, inspect page information, and capture PDFs.
- 1,000 screenshots a month are free with no card. Paid plans start at $5 for 3,000 screenshots; every feature is on every plan.
Sign up free for 1,000 screenshots a month, with no card required.
FAQ
Can Browse AI search first and then paginate?
Yes. Train the robot to enter and submit the query, wait for the results to render, and then capture the repeated list with its pagination behavior.
Does “From a list” extract the content of each result page?
No. It captures repeated items in the list. To gather fields from each linked detail page, add a separate detail-page robot in a workflow.
Which pagination option should I choose for numbered pages?
Use Click next for a Next control, arrow, or numbered pages, then verify later-page records in the run output.
Is a screenshot API a replacement for search-result extraction?
No. A screenshot records how a page looks; it does not turn all result cards into a structured dataset. Use a list-extraction workflow when you need rows and fields, and a screenshot when you need a visual capture.


