ScreenshotNeo

BlogHow-to

How to Choose Fields for a Browse AI Data Extraction Robot

Choose fields that fit your page structure and the dataset you need. Learn when to use lists, text capture, or Table Studio, and how to check coverage.

By the ScreenshotNeo team4 October 20269 min read

Choose fields from the dataset you need, then match the capture method to the page. Use From a list when the page repeats similar records and you want one row per record. Use Just text for individual values scattered across a page. If the data is already visible and does not require interaction, Browse AI currently recommends starting in Table Studio, then reviewing its proposed columns.

Before building the robot, write down each value you need, what it means, where it appears, how many records or pages are in scope, and whether runs will repeat. This gives you a testable schema and helps prevent collecting fields that do not serve the output. See Browse AI’s best practices for web scraping, extraction, and monitoring.

1. Choose a capture method that fits the page

Page and task Starting method Why it fits
Search results, directory entries, product cards, reviews, or another repeated pattern From a list Capture consistent data points into one row per item, with pagination settings for lists that span more content.
A single page with scattered values, such as a title, price, contact detail, or specification Just text Select individual page elements and give each a meaningful field name.
Data is visible on the page without clicks or other actions Table Studio It proposes a structured table that you can review, add columns to, or trim before saving.
Values appear only after clicking, choosing a dropdown, entering a value, or logging in Robot Studio interaction, then capture Train the required action first, then capture the revealed data. A complex journey across page types may need a workflow or multiple robots.

Browse AI describes the distinction between From a list and Just text, and its robot-building guide covers the initial setup. Its current getting-started guidance recommends Table Studio for data already visible without interaction.

2. Define the output before selecting anything

Sketch the result as a small schema. Make each field name describe one value, not the whole page. For example, if a pricing page shows distinct monthly and annual prices, use separate fields such as monthly_price and annual_price. Those names are an editorial example; the key is that the labels make the values distinguishable to the person or system using the table.

  • Purpose: What decision, report, or downstream process will use the data?
  • Fields: Which exact values are required for that purpose?
  • Meaning: What does each value represent, including units or whether a price is monthly or annual?
  • Scope: Which records, pages, or page types should be included?
  • Run inputs: Does each run need a variable input such as a search term?
  • Refresh needs: Will the robot run once or repeatedly, and how will you compare results?

These are planning questions drawn from Browse AI’s best-practices guidance. There is no universal ideal number of fields: capture the fields needed for the task and validate them against the source.

3. Select fields for a repeated list

  1. Open the page that contains the actual records you want. For example, when extracting competitor pricing, use the pricing page rather than a general homepage.
  2. Choose From a list for records that share a repeated structure.
  3. Wait until the dotted outline covers exactly the repeated items you mean to collect, then select it.
  4. Inspect the suggested data points. Check that each proposed field describes the same kind of value across the records.
  5. If automatic structure misses a required field or groups the wrong elements, use manual field selection and select the fields you need.
  6. Set the item count and pagination behavior to match the scope of your dataset.
  7. Preview representative records, check labels and values, then save.

Browse AI notes that a robot can combine list extraction with text capture or screenshot capture when the page needs more than one kind of output. For example, the repeated list could be the main dataset while a separate page-level value or a screenshot is also useful. See How to extract data From a list.

4. Select individual text fields or review a proposed table

Use Just text for individual page values

Choose Just text when the desired values are separate elements rather than repeated records. Select each intended value and give it a clear label. This works well for a page-level title, a contact value, or a few specifications that do not form a list of similar items.

Use Table Studio when visible data already has a table-like shape

When the information is visible and does not need clicks or typing, start in Table Studio under Browse AI’s current guidance. Review the proposed table before saving: add a column by naming it and describing the content you want, remove columns that do not belong in your output, and preview the result. Point the robot at the page that actually contains the data you need. The first-robot guide describes this review flow.

For a web page with an HTML or visually styled table, Browse AI says it can detect many table formats. If rows expand to show hidden details, decide whether visible values are sufficient or whether the robot must click to reveal more. Nested or complex structures may call for manual selection. See the table extraction guide.

5. Set pagination to match how records load

Pagination is part of the dataset definition: selecting the right fields on the first view does not collect records that never load. For list extraction, match the robot setting to the page behavior:

What the site does Pagination choice
Shows a next button, arrows, or numbered pages Click next
Has a “Load more” or “Show more” button Click load more
Loads additional records as you scroll Scroll down
Already shows all records in scope No more items

Some sites use JavaScript navigation that looks like ordinary pagination but behaves more like load more. If a run stops early, inspect what the page does when you advance it and try the matching setting. Browse AI’s pagination guide covers list extraction; its scope does not include other capture modes or traversal of individual detail pages.

6. Validate the table before saving

  • Check a few representative records, including records near the start and later in the list.
  • For every column, confirm that the value matches the label and means the same thing across rows.
  • Check that the selected list boundary includes the intended records and excludes neighboring page content.
  • Confirm the robot reaches the intended last record through the configured pagination behavior.
  • Look at blank cells in context: determine whether the source genuinely omits the value for some records.
  • Review context fields such as extraction date and input parameters when interpreting or comparing runs.

This checklist follows Browse AI’s preview and review steps. A missing value can reflect variation in the source rather than a robot failure. For example, not every product may have a rating. Browse AI’s list guidance also says that when automatic structure omits a necessary field, the documented workaround is to choose manual selection; it does not offer a mixed automatic-and-manual selection for that list. See the list extraction article and the data structure guide.

7. Common problems and fixes

Symptom Likely cause What to check or change
One row contains several records, or list boundaries look wrong The selected outline does not match one repeated item, or it includes surrounding content. Revisit the repeated group and select the outline that encloses the intended items exactly; inspect the suggested fields and sample rows.
A required field is missing from the proposed list structure Automatic structure did not detect that data point. Switch to full manual field selection and select all required list fields, as described in Browse AI’s list guidance.
Some cells are blank The source may not provide that value for every record, or the selected field may be wrong. Compare several records in the source. Keep blanks if the value is genuinely absent; correct the selection if the value is present but mapped incorrectly.
The run stops before all records are collected The page’s advance behavior and pagination choice may not match. Observe whether the site uses next, load more, or scrolling, then configure the corresponding option. Check the intended item count as well.
The page has data, but capture returns only what is initially visible Some values require an interaction, such as expanding a row or choosing a control. Train the click, input, or other required action before capturing. For complex journeys, consider a workflow or multiple robots.
A table has nested rows or inconsistent formatting The page structure may be too complex for automatic table detection. Decide which values are required, inspect whether details need to be expanded, and use manual selection when the detected structure is unsuitable.

These fixes are based on the documented behaviors and limitations in the Browse AI help articles linked above; they are not claims of independent software testing.

8. Reliability, repeated runs, and data cost

For a robot that runs more than once, make the schema and scope stable enough to compare between runs. Keep field names tied to meaning, include only needed input parameters, and retain relevant context such as extraction date and run inputs when assessing changes. Preview the output again after changing the target page, list boundary, or pagination behavior.

Plan for source variation: fields can be absent on some records, tables can be nested, and a site can change how it reveals or paginates content. Those cases affect what the output contains, so define how your downstream use should treat blanks and verify representative values after setup changes. The supplied Browse AI guidance does not establish a universal field-count target, extraction-speed benchmark, or reliability rate, so do not use one as a planning assumption.

9. Where ScreenshotNeo fits

Browse AI is for structuring data extraction into fields and rows. If a screenshot is also part of the work—for example, you need a visual record alongside extracted values—ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. A screenshot can document a page state, but it does not replace choosing and validating the extraction schema.

Or skip the browser setup

For the screenshot portion, one GET request returns an image or PDF. For the full set of capture options, see the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners are accepted and removed before capture; 60+ known consent platforms, newsletter popups, and chat widgets can be removed, and each step can be turned off.
  • Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers report the page verdict and billing status.
  • An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
  • 1,000 screenshots a month are free with no card. Paid plans start at $5 for 3,000 screenshots.

Sign up for ScreenshotNeo’s free 1,000 screenshots a month, with no card required.

10. FAQ

Should I capture every visible field?

No. Select the values needed for the intended dataset. Unneeded fields add columns to review without helping the task.

Can one robot collect a list and a separate page value?

Browse AI says a robot can combine list extraction with text capture or screenshot capture when the page calls for multiple output types.

Does a blank cell always mean the robot failed?

No. The source can omit a value for some records. Check the page for those records before changing the selection.

The reviewed official guidance does not provide a universal optimal count. Choose fields according to the task and validate their values in the preview.