ScreenshotNeo

BlogAI agents

Extract Website Markdown with an MCP Server

Learn how MCP Fetch converts URLs to Markdown, when browser rendering is needed, and how to run a reliable, secure extraction workflow.

By the ScreenshotNeo team1 October 20267 min read

To extract a website as Markdown with MCP, start with the official Model Context Protocol Fetch server. Its fetch tool accepts a URL, converts HTML to Markdown, and supports max_length, start_index, and raw-content responses. For JavaScript-heavy or bot-protected pages, add a browser-backed server or hosted renderer.

1. Choose the right extraction path

Situation Best first step Why
Static HTML or server-rendered page Official Fetch server Simple local setup and Markdown conversion
Client-rendered page with an empty HTML shell Browser-backed MCP server Chromium executes JavaScript before extraction
Bot protection, regional delivery, or large crawls Hosted MCP/API service Managed rendering, proxies, crawling, or structured output
Private or regulated content Local deployment with restricted egress More control over credentials and data movement

Use the plain Fetch server first. Escalate only when the returned Markdown is missing content, blocked, or incomplete.

2. Install the official MCP Fetch server

The official project describes itself as an MCP server that provides web content fetching and documents the prompt “Fetch a URL and extract its contents as markdown.”

uvx mcp-server-fetch

Alternatively:

python -m pip install mcp-server-fetch

The documented package requires MCP Python SDK 1.x (mcp>=1.29.0,<2). Confirm the current package requirements before pinning a production environment.

3. Configure it in an MCP client

Add the server to an MCP client such as Claude Desktop. The exact configuration file location depends on the client and operating system; use that client’s documented MCP configuration path.

{
  "mcpServers": {
    "fetch": {
      "command": "uvx",
      "args": ["mcp-server-fetch"]
    }
  }
}

Restart the client, then ask it to fetch a URL and extract its contents as Markdown. Keep requests bounded when pages are large.

4. Call the Fetch tool with paging

A typical tool call contains a URL and an optional maximum output length:

{
  "url": "https://example.com/docs",
  "max_length": 12000
}

If the response is truncated, request the next section with start_index:

{
  "url": "https://example.com/docs",
  "start_index": 12000,
  "max_length": 12000
}

Request raw content when you need the fetched representation before Markdown conversion. Exact argument names and response fields should be checked against the version you install.

5. Use a browser-backed MCP server for JavaScript pages

Plain HTTP cannot see content that is created only after JavaScript runs. The open-source web-to-markdown-mcp project documents a three-tier strategy:

  1. Request native text/markdown when the site provides it.
  2. Try plain HTTP and extraction.
  3. Fall back to Chromium when the page needs a real browser.

Its fetch_url_as_markdown tool documents controls for navigation timing, timeout, headless mode, and post-navigation polling. Use a browser fallback when the plain response is an empty shell, omits article text, or requires interaction before content appears.

{
  "url": "https://example.com/app",
  "timeout": 30000,
  "wait_until": "networkidle",
  "post_navigation_wait_ms": 1500,
  "headless": true
}

Use the parameter names supported by the server version you install; the example shows the kinds of controls browser-backed implementations commonly expose.

6. Hosted MCP and API alternatives

Hosted services remove browser and proxy operations from your deployment. HasData documents an MCP service that can fetch public URLs through managed proxies, render JavaScript, and return Markdown, text, HTML, or JSON. Its documented controls include proxy country and type, waiting, CSS selectors, link extraction, screenshots, and browser scenarios.

Context.dev documents URL-to-Markdown conversion, full-site crawling, structured extraction, and an MCP wrapper pattern built with the official SDK. Its example tool, scrape_web_markdown, requires a URL and optionally includes images; the result contains a title, resolved URL, and Markdown body. You.com documents an MCP server that combines web search with page extraction and can return full page content in Markdown or HTML.

Recheck vendor pricing, credit allowances, package versions, and feature availability before publication or procurement because those values change.

7. Build a small MCP client workflow

A robust agent workflow is:

  1. Validate and normalize the URL.
  2. Fetch with the official server and a bounded max_length.
  3. Inspect the result for missing headings, an empty shell, or an access challenge.
  4. Retry later chunks with start_index when truncation occurred.
  5. Escalate to a browser-backed or hosted renderer only when needed.
  6. Store the resolved URL and extraction metadata with the Markdown.

Keep the Markdown as the source document and perform summarization or chunking afterward. This makes extraction failures easier to distinguish from model-context limits.

8. Direct HTTP examples before MCP

A plain request is useful for diagnosing whether the target is server-rendered. These examples fetch HTML; an MCP server then performs the conversion and paging.

cURL

curl -L --max-time 30 \
  -H 'Accept: text/html,application/xhtml+xml' \
  'https://example.com/docs' -o page.html

Python

import requests

url = "https://example.com/docs"
r = requests.get(url, timeout=30, headers={"Accept": "text/html,application/xhtml+xml"})
r.raise_for_status()
open("page.html", "wb").write(r.content)

Node.js

const res = await fetch('https://example.com/docs', {
  headers: { accept: 'text/html,application/xhtml+xml' }
});
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
await Bun.write('page.html', await res.arrayBuffer());

9. Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. Use its MCP tools take_screenshot, get_page_info, and capture_pdf from Claude, Cursor, or another MCP client when the task needs a visual capture or page inspection alongside extracted content. For a direct capture, see the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

10. Troubleshooting

Symptom Cause Fix
Markdown is empty or only contains a root element Content is rendered by JavaScript Use a browser-backed server or hosted renderer and wait for navigation to settle
Output stops halfway through a page Response length or model context limit Lower max_length per request and continue with start_index
HTTP 403 or an access challenge Bot protection, rate limiting, or an untrusted user agent Respect site rules, slow requests, and use an approved browser or managed proxy where appropriate
Headings or tables are malformed Complex HTML, script-generated tables, or converter limitations Compare raw HTML with converted Markdown; use a renderer that preserves the required structures
Images are missing Lazy loading or relative URLs Use a browser fallback, resolve URLs against the final page URL, and retain image links separately
Requests reach internal addresses Unrestricted outbound fetching Allowlist destinations, block private and loopback ranges, and never expose internal URLs to untrusted prompts
Local server will not start Python, uvx, or MCP SDK version mismatch Check the installed interpreter, reinstall the documented package, and verify the SDK range

11. Security controls

The official Fetch documentation warns that the server can access local or internal IP addresses and may present a security risk. Treat every URL as untrusted input.

  • Allowlist schemes and destinations; normally permit only HTTPS.
  • Block loopback, link-local, private, and cloud metadata addresses.
  • Do not pass internal credentials or cookies into prompts.
  • Run the server with least privilege and restrict egress at the network layer.
  • Set timeouts, response-size limits, and concurrency limits.
  • Review proxy credentials and logs for accidental secret disclosure.
  • Check robots.txt and applicable site terms before crawling.

12. Performance, reliability, and cost

  • Latency: Plain HTTP is usually faster than launching Chromium. Browser rendering adds startup and page-settle time.
  • Reliability: Retry transient network failures with exponential backoff, but do not repeatedly retry permanent 4xx responses.
  • Context usage: Bound output with max_length and page large documents with start_index.
  • Scale: Use bounded concurrency and cache unchanged pages. Hosted services can be preferable when you need managed proxies, crawling, or browser capacity.
  • Cost: Local software shifts cost to your compute and operations. Hosted services charge according to their current plans or credits; verify those terms before budgeting.

13. Comparison checklist

  • Does it handle static HTML and client-rendered JavaScript?
  • Can it survive bot protection without violating site rules?
  • Are headings, tables, links, code blocks, and images preserved?
  • Can you page or chunk long results?
  • Can you run it locally for privacy?
  • Does it support crawling, selectors, or structured extraction?
  • What are the measured latency, concurrency, and current usage costs?

14. FAQ

Can MCP Fetch crawl an entire site?

It is primarily a URL fetch tool. For sitemap discovery and full-site crawling, use a service that documents those capabilities, such as Context.dev.

When should I use raw content?

Use raw content when you need to inspect the fetched representation or apply a specialized converter. Use Markdown output for normal agent reading.

Why does a browser fallback matter?

It executes page JavaScript and can wait for content that does not exist in the initial HTTP response.

How do I prevent one page from consuming the model context?

Set a maximum length, retrieve later chunks with start_index, and summarize only after the complete source has been collected.

Is local deployment automatically safe?

No. Local Fetch servers can still reach internal networks unless you enforce destination allowlists and network egress restrictions.

Can ScreenshotNeo replace Markdown extraction?

ScreenshotNeo is for screenshots, page information, and PDFs. Use it when an agent also needs a clean visual capture or page inspection; use an MCP Fetch or scraping server for Markdown text.