An MCP Server for AI Agents to Search, Scrape, and Capture Screenshots
Connect an MCP server that lets AI agents search Google, scrape pages, extract metadata and links, and capture viewport or full-page screenshots.
An MCP server can give an AI agent a practical web toolbox: search Google, extract readable page content, select structured fields with CSS selectors, collect links and metadata, and capture viewport or full-page screenshots. The Agent Scraper MCP Server provides those six tools through Streamable HTTP and also exposes a REST service.
What the server provides
| Tool | Purpose | Useful for |
|---|---|---|
scrape_url |
Extract readable text, Markdown, or HTML from a URL. | Research, summaries, document ingestion |
scrape_structured |
Extract named fields using CSS selectors. | Product data, article fields, repeated records |
screenshot_url |
Capture a viewport or full-page PNG through Playwright. | Visual checks, archives, agent verification |
extract_links |
Return page links, optionally filtered with a regular expression. | Crawling, navigation maps, URL discovery |
extract_meta |
Return title, description, canonical URL, favicon, Open Graph, and Twitter-card metadata. | SEO checks and content inventory |
search_google |
Run a Google query and return result titles, URLs, and snippets. | Discovery before scraping |
The project documents these tools and their names in its repository README.
Connect an MCP client to the hosted server
The documented Streamable HTTP endpoint is:
https://agent-scraper-mcp.onrender.com/mcp
Add it under an agent-scraper server entry in an MCP client that supports Streamable HTTP. A generic configuration shape is:
{
"mcpServers": {
"agent-scraper": {
"url": "https://agent-scraper-mcp.onrender.com/mcp"
}
}
}
Exact configuration keys differ between clients. Use the client’s HTTP MCP configuration field and keep the URL unchanged. After restarting the client, ask it to list available tools. You should see scrape_url, scrape_structured, screenshot_url, extract_links, extract_meta, and search_google.
Prompts that exercise each capability
- Search: “Search Google for the official Python 3.11 release page and return the top results with URLs.”
- Readable scrape: “Scrape the official release page and return Markdown containing the main article.”
- Structured scrape: “Extract every article headline and its link using the page’s headline CSS selector.”
- Screenshot: “Capture a full-page screenshot of the page and save the returned PNG.”
- Links: “Extract links matching
^https://docs\.python\.org/.” - Metadata: “Return the title, description, canonical URL, Open Graph image, and Twitter-card fields.”
Tell the agent which output format you need. For example, request Markdown for text intended for a model, HTML when preserving markup matters, and named JSON fields for structured extraction.
How a typical agent workflow fits together
- Discover: call
search_googleto find likely authoritative pages. - Validate: call
extract_metato check title, canonical URL, and descriptions. - Read: call
scrape_urlfor readable text or HTML. - Extract: call
scrape_structuredwhen the page has stable selectors. - Traverse: call
extract_linksto find related pages. - Verify visually: call
screenshot_urlfor viewport or full-page evidence.
Keep search and scraping separate. Search results are discovery data; the target page remains the source to inspect and cite. Use screenshots as visual evidence, not as a substitute for extracting text that the agent must reason over.
Run it locally with Python, Playwright, or Docker
Local Python installation
git clone https://github.com/aparajithn/agent-scraper-mcp.git
cd agent-scraper-mcp
python -m pip install -e "[dev]"
playwright install chromium --with-deps
uvicorn src.main:app --reload --port 8080
The documented local MCP and REST host is then your machine on port 8080. Use the local URL required by your MCP client, normally the server’s /mcp path, and confirm the path from the project’s current README before sharing it outside the machine.
Docker
docker build -t agent-scraper-mcp .
docker run --rm -p 8080:8080 -e PUBLIC_HOST=localhost agent-scraper-mcp
The project also documents PUBLIC_HOST and, for hosted payment-enabled deployments, X402_WALLET_ADDRESS. Do not expose a self-hosted instance publicly until you have decided how to protect it and how to handle URLs, cookies, and request logs.
Minimal Python HTTP check
import requests
mcp_url = "http://127.0.0.1:8080/mcp"
r = requests.get(mcp_url, timeout=30)
print(r.status_code)
print(r.headers)
print(r.text[:500])
This checks that the process is reachable. MCP clients negotiate the actual Streamable HTTP session and tool calls; use the client’s MCP transport rather than treating the endpoint as a conventional JSON REST route.
Minimal Node.js reachability check
const res = await fetch('http://127.0.0.1:8080/mcp');
console.log(res.status);
console.log(await res.text());
Screenshot behavior and practical limits
screenshot_url uses Playwright and returns a base64-encoded PNG according to the project documentation. Ask for a viewport screenshot when you need the visible layout at a browser size; ask for full-page capture when you need content below the fold.
- Pages that require login, a special header, a cookie, or a JavaScript interaction may not render as an unauthenticated request.
- Lazy-loaded content may appear only after scrolling or interaction. A screenshot can therefore differ from the initial HTML extraction.
- Bot checks, rate limits, consent dialogs, and client-side errors can affect both scraping and screenshots.
- Large pages consume more browser memory and take longer than small viewport captures.
- The repository does not publish independent latency, uptime, crawl-success, retention, or security-audit results.
Hosted pricing and payment model
The repository documents a free allowance of 50 requests per IP per day, with no credit card required. After the allowance, scraping tools cost $0.005 per request and screenshot calls cost $0.01 per request. Payment uses x402 in USDC on Base.
| Usage | Documented price |
|---|---|
| First 50 requests per IP each day | Free |
| Scraping request after quota | $0.005 |
| Screenshot request after quota | $0.01 |
The README describes machine-readable HTTP 402 payment requirements and EIP-3009 authorization handling. Treat wallet, settlement, and payment behavior as project documentation that you must validate in your own deployment and client.
Security, privacy, and deployment decisions
- Hosted endpoint: simplest setup, but verify what URLs, extracted content, and request metadata are retained before sending private targets.
- Self-hosting: keeps browser execution under your control and lets you set network policy, authentication, and logging.
- URL access: restrict outbound network access if agents can receive arbitrary URLs. Consider private IP ranges, internal hostnames, and cloud metadata endpoints.
- Credentials: never place API keys, cookies, or authorization headers in prompts. Pass secrets through your controlled application layer.
- Browser isolation: run Chromium with the permissions and filesystem access required for capture, not with broad host privileges.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Client cannot discover the server | Wrong transport or URL path | Use a client with Streamable HTTP support and the documented /mcp endpoint. |
| Local server exits while starting | Dependencies or Chromium missing | Run pip install -e "[dev]" and playwright install chromium --with-deps. |
| Screenshot is blank | Page failed to load, requires JavaScript, or blocked the browser | Open the URL directly, inspect server logs, try a viewport capture, and check the page’s bot or authentication requirements. |
| Full-page image is incomplete | Lazy content has not loaded or the page changes while scrolling | Use the page’s stable URL, allow rendering time, and compare with a readable scrape. |
| Structured extraction returns no fields | Selector does not match the rendered DOM | Inspect the current DOM, use a precise selector, and prefer scrape_url when markup is unstable. |
| Search results are empty or limited | Search provider response, quota, or transient block | Retry with a narrower query and handle empty results as a normal outcome. |
| HTTP 402 after free usage | Daily free allowance was exceeded | Use the documented x402 USDC-on-Base payment flow or self-host the service. |
| Requests are slow | Cold browser startup, heavy assets, or long pages | Reuse a running service, reduce repeated calls, cache extracted data, and request screenshots only when visual output is needed. |
Or skip the browser setup
If your requirement is dependable website screenshots rather than a general-purpose scraping server, ScreenshotNeo is the first alternative to try: it removes consent banners, newsletter popups, and chat widgets before capture, bills only clean shots, and has the lowest paid plan.
One GET request returns PNG, JPEG, WebP, or PDF. The API supports full-page shots with lazy images loaded, CSS-element capture, dark mode, device presets, custom viewports, retina scale, PDF controls, custom CSS and JavaScript, clicks, selector waits, delays, network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, async jobs, bulk capture, usage data, and an OpenAPI specification.
Read the ScreenshotNeo API documentation for the complete option list.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000.
Create a free ScreenshotNeo account and start with the 1,000-shot monthly allowance.
FAQ
Can an AI agent take a full-page screenshot through MCP?
Yes. Call screenshot_url and request a full-page capture; the documented result is a base64 PNG.
What is the difference between readable scraping and CSS-selector scraping?
Readable scraping is designed for article-like content in text, Markdown, or HTML. Structured scraping returns fields you name with CSS selectors and is better for repeated records.
Can I self-host the server?
Yes. The documented path uses Python, FastAPI, Playwright, Chromium, and Uvicorn; Docker is also provided.
Does the server extract metadata and links?
Yes. Use extract_meta for title, description, canonical, favicon, Open Graph, and Twitter-card data, and extract_links for page URLs with optional regex filtering.
How should I choose between this server and a screenshot API?
Choose this MCP server when one agent needs search, text extraction, selectors, links, metadata, and screenshots. Choose a dedicated screenshot API when capture reliability, cleaning, output formats, billing visibility, or high-volume image workflows matter most.


