ScreenshotNeo

BlogComparisons

Choosing Between Search, Fetch, and Browser APIs for Web Data

Use search to discover sources, fetch to retrieve known URLs, and browser APIs when rendering or interaction is part of the job.

By the ScreenshotNeo team30 September 20269 min read

Choosing Between Search, Fetch, and Browser APIs for Web Data

Use search when you do not know the URL, fetch when you already know the resource, and a browser API when the task depends on rendering, page state, or interaction. These are workflow roles rather than mutually exclusive technologies. A production pipeline often searches for candidates, fetches straightforward pages, and opens a browser only for pages that require JavaScript or controls.

Search, fetch, and browser APIs at a glance

Question Search API Direct fetch or extraction Browser automation
Do you know the URL? Usually no; discovery is central Yes; your code supplies it Usually yes, often from an earlier step
Main output Ranked candidates, URLs, titles, snippets, and sometimes citations or extracted content An HTTP response or transformed page content Rendered page state, screenshots, PDFs, or interaction results
Does the task require controls or navigation? Usually not No in an ordinary request Yes, when state or controls matter
Typical workflow role Discover Retrieve Render or interact
Validate before choosing Coverage, freshness, ranking, filters, citations, result limits Status, content type, authentication, parsing, size, access policy Selectors, browser runtime, sessions, latency, access constraints

Search providers differ in whether they include page content, citations, geographic filters, or scraping. “Fetch API” can mean the browser-standard JavaScript interface or a commercial extraction service that adds parsing, Markdown conversion, or rendering. Confirm the exact contract for the provider you select.

Search discovers a candidate, fetch retrieves a known resource, and a browser handles rendered interaction.
Search discovers a candidate, fetch retrieves a known resource, and a browser handles rendered interaction.

When to use a search API

Start with search when the information location is unknown. A search request accepts a query and returns candidate resources such as URLs, titles, snippets, or structured records. OpenAI describes web search as a way for models to access up-to-date information and return sourced citations; its documentation also describes URL citation annotations. That behavior is specific to OpenAI’s product, so do not assume every search API returns full page text or citations.

Search is a good fit for

  • Finding documentation pages, products, or news articles from a natural-language question.
  • Building a candidate list before a later fetch or browser step.
  • Agent workflows that need sources with citations.
  • Queries where freshness, ranking, geography, category, or time filters matter.

Search is not a substitute for retrieval

A result title and snippet may be insufficient for extraction. Treat search output as discovery data unless the provider explicitly promises scraped or extracted content. Browserless, for example, documents a Search API that can optionally scrape each result into structured, LLM-ready data, but its endpoint is marked beta and its parameters and response shapes may change. It also documents plan-dependent result caps and cloud-plan availability; verify current limits before relying on them.

When to use direct fetch or an extraction API

Use direct HTTP fetch when the URL is already known and the server response contains what you need. The browser Fetch API provides request and response interfaces for obtaining resources and is grounded in the WHATWG Fetch standard. It is not a search engine, a browser automation framework, or a cross-origin bypass.

A direct request is usually the simplest choice for a first-party JSON endpoint, a static HTML page, an RSS feed, an object in storage, or a public document. A commercial fetch or extraction service may additionally parse HTML, remove navigation, convert content to Markdown, or render JavaScript. Those capabilities vary by vendor and must be checked individually.

Minimal direct fetch examples

curl -L --fail --max-time 30 https://example.com/ -o page.html
python - <<'PY'
import requests

url = "https://example.com/"
r = requests.get(url, timeout=30)
r.raise_for_status()
print(r.headers.get("content-type"))
print(r.text[:500])
PY
const res = await fetch('https://example.com/', { redirect: 'follow' });
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const contentType = res.headers.get('content-type');
const text = await res.text();
console.log(contentType, text.slice(0, 500));

Fetch checklist

  • Set a finite connect and read timeout.
  • Follow redirects deliberately and record the final URL.
  • Check status and content type before parsing.
  • Limit response size to avoid unbounded memory use.
  • Send authentication only to the intended host.
  • Respect robots, terms, rate limits, and access controls that apply to your use.
  • Retry transient failures with capped exponential backoff and jitter.

When to use browser automation

Choose a browser when the job actually needs a browser context: JavaScript rendering, navigation, cookies, local storage, login state, clicking controls, infinite scroll, or a visual result. Browser automation is slower and operationally heavier than an HTTP request, but it observes the page after scripts and layout have run.

Browser APIs are appropriate when accuracy depends on rendered state or interaction. They do not guarantee access to every site, defeat bot checks, or remove the need to handle authentication and policy constraints. Browserbase’s published guidance frames Search as discovery, Fetch as checking a known URL, and Browser as the option for interaction, login, JavaScript, or workflows where accuracy matters more than speed; verify its current documentation before implementing against that positioning.

Runnable Playwright example

npm install playwright
npx playwright install chromium
import { chromium } from 'playwright';

const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
await page.goto('https://example.com/', { waitUntil: 'networkidle', timeout: 60000 });
console.log(await page.title());
console.log((await page.locator('body').innerText()).slice(0, 1000));
await page.screenshot({ path: 'page.png', fullPage: true });
await browser.close();

Browser workflow controls

  • Use explicit waits for a selector or state instead of arbitrary long sleeps.
  • Set viewport, locale, timezone, and user agent to match the use case.
  • Persist only the session data you are allowed to retain.
  • Close pages and contexts promptly; browser processes consume substantially more resources than HTTP clients.
  • Capture console, network, and page errors so failures are diagnosable.
  • Make selectors resilient to layout changes and provide fallbacks for missing controls.

Compose the three approaches

A practical architecture routes each URL through the least expensive method that can satisfy the task:

  1. Search: find candidate URLs and retain ranking, snippets, and citation metadata.
  2. Classify: decide whether each candidate is a static resource, a known API, or an interactive page.
  3. Fetch: retrieve static HTML, JSON, feeds, or documents directly.
  4. Escalate: open a browser only when rendering, state, or interaction is required.
  5. Normalize: store the source URL, final URL, timestamp, status, content type, and extraction method with the result.
async function retrieve(candidate) {
  if (candidate.kind === 'api' || candidate.kind === 'static-html') {
    const response = await fetch(candidate.url);
    if (!response.ok) throw new Error(`Fetch failed: ${response.status}`);
    return { method: 'fetch', url: response.url, body: await response.text() };
  }

  // Route interactive pages to your browser worker.
  return { method: 'browser', url: candidate.url };
}

How to choose a provider

Compare providers against the actual workflow rather than labels alone. Ask these questions during evaluation:

  • Discovery: What indexes are covered? How fresh are results? Can you constrain geography, time, category, or domain?
  • Citations: Are source URLs and citation metadata returned in a stable format?
  • Rendering: Which JavaScript features, browsers, and page states are supported?
  • Output: Do you receive HTML, Markdown, structured fields, a screenshot, a PDF, or raw browser events?
  • Access: How are cookies, headers, authentication, proxies, and consent dialogs handled?
  • Operations: What are rate limits, concurrency limits, data-retention rules, regional controls, and failure semantics?
  • Cost: Is billing based on requests, results, browser time, bandwidth, or successful outputs?
  • Change risk: Are the endpoint and response schema stable, or explicitly beta?

Browserless currently documents its Search API as beta, cloud-plan-only, and subject to plan-dependent result limits (3 on Free, 5 on Prototyping, 10 on Starter, and 20 on Scale and above; omitted limit defaults to 10 or the plan maximum, whichever is lower). These are provider-specific limits captured from its documentation and may change.

Screenshot capture without managing a browser

If the browser step is only needed to produce a clean image or PDF, ScreenshotNeo provides a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled.

Only clean shots are billed. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the outcome with X-Page-Verdict and X-Billed headers. The service also supports full-page capture with lazy images, CSS-selector element capture, dark mode, device presets, custom viewports, retina scale, PDF options, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, async jobs, bulk capture, usage data, and an OpenAPI specification.

Or skip the browser setup

See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots, and 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Troubleshooting

Search returns irrelevant or stale candidates

Tighten the query, add domain or time filters, and keep the original ranking metadata. Validate the final page with a fetch before sending it to downstream extraction. Search coverage and freshness are provider-specific.

A screenshot cleanup step removes common overlays before the image is returned.
A screenshot cleanup step removes common overlays before the image is returned.

Fetch returns HTML instead of the content you expected

Inspect the status, redirect chain, and Content-Type. You may have reached a login page, an error document, or a JavaScript shell. If the content appears only after scripts run, escalate to a browser.

A browser page never reaches network idle

Some sites keep analytics or streaming connections open. Wait for a specific selector or application state, and set a hard timeout. Capture console and network errors to identify failing resources.

Selectors fail intermittently

Wait for the element, use stable attributes, and handle optional dialogs. Avoid selectors tied to generated class names or exact positions.

Authentication works in fetch but not in the browser

HTTP credentials and browser session state are different. Load cookies or storage state explicitly, confirm the correct origin, and avoid logging secrets in traces.

Use a cleanup-capable screenshot service or click and hide the relevant selectors before capture. ScreenshotNeo removes supported consent platforms, newsletter popups, and chat widgets before capture.

Performance, reliability, and cost

  • Latency: Direct fetch is generally the smallest path; search adds discovery work; browser startup and rendering add the most moving parts. Measure on your target domains instead of relying on cross-provider claims.
  • Reliability: Record status, timeout, final URL, and method. Retry only transient failures and make browser jobs idempotent.
  • Cost: Cache stable resources, deduplicate URLs, and reserve browser execution for pages that need it. Check whether a provider charges per result, browser time, request, or successful output.
  • Capacity: Bound concurrency for both HTTP and browser workers. Browser contexts can exhaust CPU and memory faster than fetch workers.
  • Freshness: Search indexes and caches can be stale. Store retrieval timestamps and choose cache TTLs that match the data’s change rate.

Decision checklist

  • Do I know the exact URL? If no, begin with search.
  • Is the required content in the HTTP response? If yes, use fetch.
  • Must I click, log in, wait for JavaScript, or inspect rendered layout? If yes, use a browser.
  • Can I compose methods so only a small fraction of pages use a browser?
  • Do the provider’s limits, citations, output, access rules, retention, and pricing match my workload?
  • Have I evaluated the design against representative target pages and explicit success criteria?

FAQ

Do I need a browser API for every JavaScript-heavy site?

No. First determine whether the data is available from an underlying JSON endpoint or server-rendered response. Use a browser when the required state is produced only through rendering or interaction.

Can a search API replace a crawler?

Usually not. Search discovers indexed candidates; crawling and fetching are separate steps unless the provider explicitly bundles extraction.

Is the browser Fetch API the same as a fetch service?

No. The browser Fetch API is a request and response interface. Commercial fetch services may add parsing, extraction, Markdown conversion, or rendering.

Should I always choose the fastest option?

Choose the least complex option that meets the success criteria. A fast fetch that misses rendered content is not successful, while a browser for every URL wastes resources.

Are provider limits permanent?

No. Beta status, plan caps, parameters, and response formats can change. Recheck current documentation before shipping and monitor responses in production.