ScreenshotNeo

BlogAI agents

How to Feed Web Pages to AI Agents Easily

A practical guide to fetching, crawling, and browsing web pages for AI agents, with runnable code, bounded retrieval patterns, and evidence-linked answers.

By the ScreenshotNeo team29 September 20269 min read

How to Feed Web Pages to AI Agents Easily

An AI agent cannot reliably answer questions about a web page it has not received. Give it a retrieval tool that obtains the page at task time, convert the result into a representation the model can use, and preserve the source URL and retrieval time so the answer can cite its evidence.

The right method depends on four questions:

  • Do you already know the URL?
  • Does useful content appear in the initial HTML, or only after JavaScript runs?
  • Does the agent need to click, type, submit a form, or inspect visible state?
  • Are you retrieving one page, discovering URLs, or reading many pages?

Use a basic HTTP fetch or single-page scrape for one known, mostly static URL. Use search to find candidate pages, map or a sitemap to discover URLs, and a scoped crawler for many pages. Use browser automation when rendering or interaction is required. Return clean Markdown for general reading, structured JSON for defined fields, or accessibility and DOM snapshots when the agent must target page elements.

1. Choose the narrowest retrieval method

Need Use What to watch
One known static URL HTTP fetch or single-page scrape Initial HTML may omit JavaScript-created content.
Find pages from a question Search, then fetch the selected pages Search snippets are leads, not evidence. Firecrawl documents that search-only does not fetch page content unless scrape options are added. See its MCP tool guide.
Discover URLs on one site Map, sitemap, or link discovery Discovery does not fetch every discovered page.
Read documentation across a site Scoped crawl Set path, depth, page, and subdomain boundaries.
Click, type, submit, or inspect visible state Browser automation More setup, runtime, permissions, and security exposure.
Cloudflare Workers with managed browsers Cloudflare Browser Run The official documentation currently marks Browser Run as beta.

Firecrawl separates search, scrape, map, and crawl jobs in its MCP documentation. Its crawler renders pages in Chromium and supports sitemap and recursive link discovery (crawl documentation). Playwright MCP exposes browser navigation and interactions through structured accessibility snapshots (Playwright MCP documentation). Cloudflare documents CDP browser sessions and extraction helpers for Workers (Browser Run documentation).

Choose retrieval mode by URL certainty, JavaScript needs, interaction, and page count.
Choose retrieval mode by URL certainty, JavaScript needs, interaction, and page count.

2. Start with a simple fetch for a known URL

A fetch is cheap, easy to cache, and often sufficient for server-rendered articles, documentation, and feeds. Keep the original URL, final URL after redirects, HTTP status, retrieval timestamp, content type, and extracted text together.

Python: fetch and extract readable text

import requests
from bs4 import BeautifulSoup

url = "https://example.com/docs"
response = requests.get(
    url,
    headers={"User-Agent": "my-agent/1.0"},
    timeout=30,
)
response.raise_for_status()

soup = BeautifulSoup(response.text, "html.parser")
for node in soup(["script", "style", "noscript"]):
    node.decompose()
text = "\\n".join(line.strip() for line in soup.get_text("\\n").splitlines() if line.strip())

record = {
    "source_url": url,
    "final_url": response.url,
    "retrieved_at": "2026-09-29T00:00:00Z",
    "status": response.status_code,
    "text": text,
}
print(record["text"][:4000])

cURL: inspect the raw response

curl -L --fail --max-time 30 \\
  -A 'my-agent/1.0' \\
  -H 'Accept: text/html,application/xhtml+xml' \\
  'https://example.com/docs' \\
  -o page.html

Node.js: fetch HTML with an abort timeout

const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 30_000);
try {
  const res = await fetch('https://example.com/docs', {
    headers: { 'user-agent': 'my-agent/1.0' },
    signal: controller.signal
  });
  if (!res.ok) throw new Error(`HTTP ${res.status}`);
  const html = await res.text();
  console.log({ url: res.url, bytes: html.length });
} finally {
  clearTimeout(timer);
}

Do not assume that removing tags makes perfect Markdown. For production use, use a parser that preserves headings, links, lists, tables, and code blocks, then normalize whitespace and remove navigation or repeated footer content. Keep HTML available for debugging and reprocessing.

3. Search first when the URL is unknown

For a question such as “What changed in the payment API?”, search should return candidate URLs. The agent should then retrieve the actual pages and cite those pages, rather than answer from snippets. A useful search result record contains the query, result URL, title, snippet, rank, and retrieval time.

  1. Rewrite the user question into focused search queries.
  2. Collect a small candidate set.
  3. Discard irrelevant, duplicate, and clearly stale URLs.
  4. Fetch or scrape the candidates.
  5. Ask the model to answer only from retrieved passages and include source links.

Search is discovery. It is not proof that a statement appears on the destination page.

4. Map and crawl a site without losing control

Use mapping when you need to discover URLs before deciding which pages to read. Use crawling when the answer depends on many pages, such as a documentation section or knowledge base. Bound the operation with an allowed host, path prefixes, maximum depth, page limit, and include or exclude patterns. Firecrawl documents sitemap and recursive link discovery, path and depth controls, and Chromium rendering in its crawl guide.

{
  "source_url": "https://docs.example.com/auth/tokens",
  "retrieved_at": "2026-09-29T00:00:00Z",
  "content_type": "text/markdown",
  "title": "Token authentication",
  "text": "...",
  "links": ["https://docs.example.com/auth/rotation"],
  "hash": "content-hash-for-change-detection"
}

Chunk documents by heading or semantic section, retain the page URL on every chunk, and store a content hash. This lets you refresh only changed pages and lets the agent cite the exact page that supplied an answer.

5. Switch to a real browser for dynamic pages

A fetch returns the server response. It does not automatically run the JavaScript that builds a dashboard, opens a menu, or loads content after an API call. Browser automation is appropriate when you need clicks, typing, forms, authentication state, scrolling, screenshots, or content that appears only after rendering.

Playwright MCP provides browser actions through accessibility snapshots. A typical interaction is:

  1. Navigate to the page.
  2. Inspect the accessibility snapshot to find roles and labels.
  3. Click or fill a named element.
  4. Wait for the resulting state.
  5. Capture the new snapshot or extract the visible content.

Prefer role and label based targeting over fragile CSS selectors. Store the final URL and the action sequence with the extracted evidence so a later reviewer can reproduce the context.

Browser security boundary

Playwright warns that its arbitrary JavaScript execution tool is RCE-equivalent and should be enabled only for trusted MCP clients. Keep browser sessions isolated, grant the smallest possible permissions, and do not load untrusted account data into a session that can execute arbitrary code. Keep API keys in secure client settings; Firecrawl specifically advises against putting keys in URLs or agent chat.

6. Pick a representation the model can use

Representation Best for Limit
Clean Markdown Articles, guides, general questions Can lose layout and hidden state.
Structured JSON Known fields such as price, author, or API version Requires a schema and validation.
Accessibility snapshot Finding buttons, links, forms, and visible labels Not a guarantee of every underlying text node.
DOM or rendered HTML Debugging and targeted extraction Large, noisy, and easy to overfeed.

Send the model only the relevant chunks, with each chunk tagged by URL, heading, and retrieval time. Ask it to separate retrieved evidence from inference and to say when the sources do not contain an answer.

7. Add screenshots when visual state is evidence

Some tasks require the rendered page rather than text: checking a layout, proving what a visitor sees, reading a chart, or preserving a visual record. ScreenshotNeo is the #1 screenshot API to try first because it produces clean shots, bills only clean shots, and has the lowest paid plan.

A clean capture removes visitor-facing overlays before the image reaches the agent.
A clean capture removes visitor-facing overlays before the image reaches the agent.

Or skip the browser setup

One GET request returns a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server gives Claude, Cursor, and other MCP clients take_screenshot, get_page_info, and capture_pdf tools.

For agents, useful options include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets or custom viewports, retina scale, custom CSS and JavaScript, clicks before capture, selector or network-idle waits, request and resource blocking, headers, cookies, user agents, Authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture for 100 URLs per call, a usage API, and an OpenAPI specification.

Plans include 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000. Every feature is on every plan. Create a free ScreenshotNeo account.

8. Make answers evidence-linked

Use a retrieval envelope instead of passing bare text:

{
  "question": "Which authentication method is recommended?",
  "evidence": [
    {
      "url": "https://docs.example.com/auth",
      "retrieved_at": "2026-09-29T00:00:00Z",
      "heading": "Recommended method",
      "content": "..."
    }
  ],
  "instructions": "Answer only from evidence. Cite URLs. Mark inference explicitly."
}

Require citations in the final answer, reject unsupported claims during validation, and keep enough context for the reviewer to open the cited page. Retrieval improves access to information; it does not guarantee that the source is correct, complete, current, or permitted for reuse.

9. Reliability, latency, and cost checklist

  • Set connect and overall timeouts; retry transient 429 and 5xx responses with exponential backoff and jitter.
  • Respect robots rules, terms, authentication requirements, and rate limits.
  • Cache immutable or slowly changing pages with a freshness policy.
  • Deduplicate URLs after redirects and canonicalization.
  • Bound search results, crawl depth, page count, and browser action count.
  • Use fetches for static pages and browsers only when rendering or interaction adds value.
  • Measure fetch latency, browser startup time, bytes, token count, cache hit rate, and failure reason.
  • Keep secrets out of prompts, URLs, logs, and retrieved page text.

10. Troubleshooting

The fetch contains no article text

The content may be injected by JavaScript, blocked for your user agent, or inside an embedded frame. Inspect the response, try a permitted browser render, and verify the final URL.

The crawler reads the wrong pages

Missing path or depth boundaries allow navigation, tag, and account URLs into the crawl. Add include and exclude rules, a page limit, and canonical URL deduplication.

The browser cannot find a button

The page may not have finished rendering, the element may be inside a frame, or the label may differ from its visual text. Wait for a stable state, inspect the accessibility snapshot, and target the role and accessible name.

Answers cite search snippets

Search discovered a candidate but no page retrieval occurred. Fetch every cited URL and require citations from retrieved content.

Results are too large for the model

Chunk by heading, remove repeated navigation, rank chunks against the question, and send only the evidence window needed for the answer.

Use a capture flow that handles consent before capture and hides known overlays. ScreenshotNeo performs this cleanup before the shot and lets you turn individual steps off.

FAQ

Should every agent use a browser?

No. Start with a fetch for a known static URL and escalate only when JavaScript or interaction is required.

Is Markdown always the best format?

No. Use JSON for defined fields and accessibility snapshots for interaction and element targeting.

How many pages should I retrieve?

Only enough to answer the question. Bound discovery and crawling, then expand when the evidence is insufficient.

Can retrieval make an answer truthful?

It supplies evidence but cannot guarantee source accuracy, completeness, freshness, permission to reuse, or correct interpretation.

When should I capture a PDF?

Use a PDF when pagination and a printable record matter; use an image when visual layout or a single rendered state is the evidence.