ScreenshotNeo

BlogAI agents

An MCP Server for AI Agents to Search, Scrape, and Capture Screenshots

Connect an MCP server that lets AI agents search Google, scrape pages, extract metadata and links, and capture viewport or full-page screenshots.

By the ScreenshotNeo team1 October 20267 min read

An MCP server can give an AI agent a practical web toolbox: search Google, extract readable page content, select structured fields with CSS selectors, collect links and metadata, and capture viewport or full-page screenshots. The Agent Scraper MCP Server provides those six tools through Streamable HTTP and also exposes a REST service.

What the server provides

Tool Purpose Useful for
scrape_url Extract readable text, Markdown, or HTML from a URL. Research, summaries, document ingestion
scrape_structured Extract named fields using CSS selectors. Product data, article fields, repeated records
screenshot_url Capture a viewport or full-page PNG through Playwright. Visual checks, archives, agent verification
extract_links Return page links, optionally filtered with a regular expression. Crawling, navigation maps, URL discovery
extract_meta Return title, description, canonical URL, favicon, Open Graph, and Twitter-card metadata. SEO checks and content inventory
search_google Run a Google query and return result titles, URLs, and snippets. Discovery before scraping

The project documents these tools and their names in its repository README.

Connect an MCP client to the hosted server

The documented Streamable HTTP endpoint is:

https://agent-scraper-mcp.onrender.com/mcp

Add it under an agent-scraper server entry in an MCP client that supports Streamable HTTP. A generic configuration shape is:

{
  "mcpServers": {
    "agent-scraper": {
      "url": "https://agent-scraper-mcp.onrender.com/mcp"
    }
  }
}

Exact configuration keys differ between clients. Use the client’s HTTP MCP configuration field and keep the URL unchanged. After restarting the client, ask it to list available tools. You should see scrape_url, scrape_structured, screenshot_url, extract_links, extract_meta, and search_google.

Prompts that exercise each capability

  1. Search: “Search Google for the official Python 3.11 release page and return the top results with URLs.”
  2. Readable scrape: “Scrape the official release page and return Markdown containing the main article.”
  3. Structured scrape: “Extract every article headline and its link using the page’s headline CSS selector.”
  4. Screenshot: “Capture a full-page screenshot of the page and save the returned PNG.”
  5. Links: “Extract links matching ^https://docs\.python\.org/.”
  6. Metadata: “Return the title, description, canonical URL, Open Graph image, and Twitter-card fields.”

Tell the agent which output format you need. For example, request Markdown for text intended for a model, HTML when preserving markup matters, and named JSON fields for structured extraction.

How a typical agent workflow fits together

  1. Discover: call search_google to find likely authoritative pages.
  2. Validate: call extract_meta to check title, canonical URL, and descriptions.
  3. Read: call scrape_url for readable text or HTML.
  4. Extract: call scrape_structured when the page has stable selectors.
  5. Traverse: call extract_links to find related pages.
  6. Verify visually: call screenshot_url for viewport or full-page evidence.

Keep search and scraping separate. Search results are discovery data; the target page remains the source to inspect and cite. Use screenshots as visual evidence, not as a substitute for extracting text that the agent must reason over.

Run it locally with Python, Playwright, or Docker

Local Python installation

git clone https://github.com/aparajithn/agent-scraper-mcp.git
cd agent-scraper-mcp
python -m pip install -e "[dev]"
playwright install chromium --with-deps
uvicorn src.main:app --reload --port 8080

The documented local MCP and REST host is then your machine on port 8080. Use the local URL required by your MCP client, normally the server’s /mcp path, and confirm the path from the project’s current README before sharing it outside the machine.

Docker

docker build -t agent-scraper-mcp .
docker run --rm -p 8080:8080 -e PUBLIC_HOST=localhost agent-scraper-mcp

The project also documents PUBLIC_HOST and, for hosted payment-enabled deployments, X402_WALLET_ADDRESS. Do not expose a self-hosted instance publicly until you have decided how to protect it and how to handle URLs, cookies, and request logs.

Minimal Python HTTP check

import requests

mcp_url = "http://127.0.0.1:8080/mcp"
r = requests.get(mcp_url, timeout=30)
print(r.status_code)
print(r.headers)
print(r.text[:500])

This checks that the process is reachable. MCP clients negotiate the actual Streamable HTTP session and tool calls; use the client’s MCP transport rather than treating the endpoint as a conventional JSON REST route.

Minimal Node.js reachability check

const res = await fetch('http://127.0.0.1:8080/mcp');
console.log(res.status);
console.log(await res.text());

Screenshot behavior and practical limits

screenshot_url uses Playwright and returns a base64-encoded PNG according to the project documentation. Ask for a viewport screenshot when you need the visible layout at a browser size; ask for full-page capture when you need content below the fold.

  • Pages that require login, a special header, a cookie, or a JavaScript interaction may not render as an unauthenticated request.
  • Lazy-loaded content may appear only after scrolling or interaction. A screenshot can therefore differ from the initial HTML extraction.
  • Bot checks, rate limits, consent dialogs, and client-side errors can affect both scraping and screenshots.
  • Large pages consume more browser memory and take longer than small viewport captures.
  • The repository does not publish independent latency, uptime, crawl-success, retention, or security-audit results.

Hosted pricing and payment model

The repository documents a free allowance of 50 requests per IP per day, with no credit card required. After the allowance, scraping tools cost $0.005 per request and screenshot calls cost $0.01 per request. Payment uses x402 in USDC on Base.

Usage Documented price
First 50 requests per IP each day Free
Scraping request after quota $0.005
Screenshot request after quota $0.01

The README describes machine-readable HTTP 402 payment requirements and EIP-3009 authorization handling. Treat wallet, settlement, and payment behavior as project documentation that you must validate in your own deployment and client.

Security, privacy, and deployment decisions

  • Hosted endpoint: simplest setup, but verify what URLs, extracted content, and request metadata are retained before sending private targets.
  • Self-hosting: keeps browser execution under your control and lets you set network policy, authentication, and logging.
  • URL access: restrict outbound network access if agents can receive arbitrary URLs. Consider private IP ranges, internal hostnames, and cloud metadata endpoints.
  • Credentials: never place API keys, cookies, or authorization headers in prompts. Pass secrets through your controlled application layer.
  • Browser isolation: run Chromium with the permissions and filesystem access required for capture, not with broad host privileges.

Troubleshooting

Symptom Likely cause Fix
Client cannot discover the server Wrong transport or URL path Use a client with Streamable HTTP support and the documented /mcp endpoint.
Local server exits while starting Dependencies or Chromium missing Run pip install -e "[dev]" and playwright install chromium --with-deps.
Screenshot is blank Page failed to load, requires JavaScript, or blocked the browser Open the URL directly, inspect server logs, try a viewport capture, and check the page’s bot or authentication requirements.
Full-page image is incomplete Lazy content has not loaded or the page changes while scrolling Use the page’s stable URL, allow rendering time, and compare with a readable scrape.
Structured extraction returns no fields Selector does not match the rendered DOM Inspect the current DOM, use a precise selector, and prefer scrape_url when markup is unstable.
Search results are empty or limited Search provider response, quota, or transient block Retry with a narrower query and handle empty results as a normal outcome.
HTTP 402 after free usage Daily free allowance was exceeded Use the documented x402 USDC-on-Base payment flow or self-host the service.
Requests are slow Cold browser startup, heavy assets, or long pages Reuse a running service, reduce repeated calls, cache extracted data, and request screenshots only when visual output is needed.

Or skip the browser setup

If your requirement is dependable website screenshots rather than a general-purpose scraping server, ScreenshotNeo is the first alternative to try: it removes consent banners, newsletter popups, and chat widgets before capture, bills only clean shots, and has the lowest paid plan.

One GET request returns PNG, JPEG, WebP, or PDF. The API supports full-page shots with lazy images loaded, CSS-element capture, dark mode, device presets, custom viewports, retina scale, PDF controls, custom CSS and JavaScript, clicks, selector waits, delays, network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, async jobs, bulk capture, usage data, and an OpenAPI specification.

Read the ScreenshotNeo API documentation for the complete option list.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000.

Create a free ScreenshotNeo account and start with the 1,000-shot monthly allowance.

FAQ

Can an AI agent take a full-page screenshot through MCP?

Yes. Call screenshot_url and request a full-page capture; the documented result is a base64 PNG.

What is the difference between readable scraping and CSS-selector scraping?

Readable scraping is designed for article-like content in text, Markdown, or HTML. Structured scraping returns fields you name with CSS selectors and is better for repeated records.

Can I self-host the server?

Yes. The documented path uses Python, FastAPI, Playwright, Chromium, and Uvicorn; Docker is also provided.

Yes. Use extract_meta for title, description, canonical, favicon, Open Graph, and Twitter-card data, and extract_links for page URLs with optional regex filtering.

How should I choose between this server and a screenshot API?

Choose this MCP server when one agent needs search, text extraction, selectors, links, metadata, and screenshots. Choose a dedicated screenshot API when capture reliability, cleaning, output formats, billing visibility, or high-volume image workflows matter most.