ScreenshotNeo

BlogEngineering

Website to Markdown API: Convert URLs for LLMs and RAG

Compare Jina Reader and Firecrawl, then build a reliable URL-to-Markdown pipeline for LLM prompts, RAG indexing, and site crawls.

By the ScreenshotNeo team1 October 20267 min read

A website-to-Markdown API fetches a URL, removes navigation and other page chrome, and returns clean Markdown or structured data for an LLM prompt, RAG index, or knowledge base. For one public URL, Jina Reader is the shortest path: prepend https://r.jina.ai/ to the URL. For JavaScript-heavy pages, structured extraction, screenshots, or a site-wide corpus, use Firecrawl Scrape or Crawl.

Choose the right API

Need Best fit Reason
One static or mostly static page Jina Reader A single URL pattern returns Markdown with little setup.
JavaScript-rendered page Firecrawl Scrape It renders pages in real Chromium and can return Markdown, JSON, HTML, links, or screenshots.
Many pages from a site Firecrawl Crawl It follows subpages from a starting URL and returns a consistent corpus.
Auditable RAG records Either, with your own metadata layer Store the source URL, retrieval time, and page metadata beside every document.

Jina documents 20 requests per minute without a key, 500 RPM with free or paid keys, 5,000 RPM on premium, and approximately 7.9 seconds average latency. Firecrawl bills by credits: Scrape and Crawl cost one credit per page, Map costs one credit per call, Search costs two credits per 10 results, and JSON extraction adds four credits per page.

See the Jina Reader documentation and Firecrawl documentation for current limits and request options.

Minimal URL-to-Markdown request with Jina Reader

The Reader pattern is GET https://r.jina.ai/<URL>. URL-encode the target when using a client that requires it.

curl --fail --show-error --location \
  "https://r.jina.ai/https://example.com/docs/getting-started" \
  -o page.md

Pass the returned Markdown directly to a model for a small prompt, or send it to a chunker for retrieval. A key gives higher rate limits; keep it in an environment variable rather than source control when your account requires one.

export JINA_API_KEY="YOUR_API_KEY"
curl --fail --show-error --location \
  -H "Authorization: Bearer $JINA_API_KEY" \
  "https://r.jina.ai/https://example.com/docs/getting-started" \
  -o page.md

Python

import os
import requests

url = "https://example.com/docs/getting-started"
headers = {}
if os.getenv("JINA_API_KEY"):
    headers["Authorization"] = f"Bearer {os.environ['JINA_API_KEY']}"

response = requests.get(
    f"https://r.jina.ai/{url}",
    headers=headers,
    timeout=60,
)
response.raise_for_status()
open("page.md", "w", encoding="utf-8").write(response.text)
print(f"saved {len(response.text)} characters")

Node.js

const target = 'https://example.com/docs/getting-started';
const headers = process.env.JINA_API_KEY
  ? { Authorization: `Bearer ${process.env.JINA_API_KEY}` }
  : {};

const res = await fetch(`https://r.jina.ai/${target}`, { headers });
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
await Bun.write('page.md', await res.text());

Build a RAG ingestion record

Do not index only the Markdown body. Preserve provenance so an answer can be traced to the page that produced it.

from datetime import datetime, timezone
import hashlib
import requests

source_url = "https://example.com/docs/getting-started"
markdown = requests.get(
    f"https://r.jina.ai/{source_url}", timeout=60
).text

record = {
    "id": hashlib.sha256(source_url.encode()).hexdigest(),
    "source_url": source_url,
    "retrieved_at": datetime.now(timezone.utc).isoformat(),
    "content_type": "text/markdown",
    "text": markdown,
}
# Send record["text"] to your chunker and vector index.
print(record["source_url"], record["retrieved_at"])

Chunking and metadata checklist

  • Split on headings where possible so a section remains understandable.
  • Keep the source URL and retrieval timestamp on every chunk.
  • Carry page title and other returned metadata into chunk metadata when available.
  • Hash the canonical URL to make repeat crawls idempotent.
  • Re-fetch on a schedule appropriate to how often the source changes.
  • Show the source URL in generated answers for traceability.

Use Firecrawl Scrape for difficult pages

Firecrawl Scrape is intended for pages that need browser rendering or more than plain Markdown. It runs the page in Chromium and can return Markdown, JSON, HTML, links, or a screenshot. Choose the output that matches the downstream job: Markdown for general RAG, JSON for a known schema, links for discovery, and HTML when your parser needs the original structure.

Use the request examples in the Firecrawl API documentation with your Firecrawl key. A typical Scrape job supplies a URL and requests one or more output formats. Keep the response’s source URL and retrieval time in your own record exactly as you would with Jina.

Crawl an entire site

When one page is insufficient, Firecrawl Crawl starts at a URL, follows subpages, and returns a consistent Markdown or JSON corpus. Define the scope before running it: decide which host and path are allowed, how many pages you need, and whether you want Markdown or structured JSON. Store each returned page as an independent record so failed or changed pages can be retried without rebuilding the whole index.

  1. Start with a representative URL.
  2. Run a small crawl and inspect titles, links, and duplicate content.
  3. Set your scope and page limit in the Crawl request.
  4. Persist each page with its canonical URL and retrieval time.
  5. Chunk and index successful pages; record failures for retry.

Crawl costs one credit per page. If you add JSON extraction, budget four additional credits per page.

Rendering, access, and edge cases

  • JavaScript content: A plain fetch can miss content inserted after load. Use Firecrawl Scrape for a Chromium-rendered page.
  • Authentication: Public Reader requests cannot access a private application unless the provider’s supported authentication flow is configured. Do not place private credentials in a URL.
  • Robots, bot checks, or CAPTCHAs: The fetch may return an interstitial instead of the article. Detect very short or clearly blocked output and keep the page out of the index.
  • Infinite scroll: A single URL may expose only the initially rendered items. Prefer a source API or a crawler with explicit pagination handling.
  • PDFs and binary files: Treat them as a separate ingestion path; do not assume an HTML-to-Markdown converter will preserve layout.
  • Duplicate URLs: Normalize fragments, trailing slashes, and tracking parameters before hashing.
  • Very large pages: Cap document size, split by headings, and reject navigation-only or error pages.
  • Freshness: Save retrieval time and re-fetch changed sources instead of blindly appending duplicates.

Reliability and performance

Rate limits and concurrency

Respect provider limits with a queue, bounded concurrency, exponential backoff, and jitter. Jina’s documented limits are 20 RPM without a key, 500 RPM with free or paid keys, and 5,000 RPM on premium. Approximately 7.9 seconds average latency means a high-volume pipeline should process asynchronously rather than block a request thread.

Retries

  • Retry transient 5xx responses and timeouts.
  • Do not retry a deterministic 4xx response until the request or credentials change.
  • Use an idempotency key based on normalized URL and extraction options in your job store.
  • Keep the last successful document while a refresh is failing.

Cost and token control

Fetch only pages you will index. For Firecrawl, count one credit per Scrape or Crawl page, one credit per Map call, two credits per 10 Search results, and four extra credits per page for JSON extraction. Chunking reduces prompt size but does not reduce provider credits, so filter boilerplate before embedding.

Troubleshooting

Symptom Likely cause Fix
401 or 403 Missing, invalid, or misplaced API key Set the provider’s documented authorization header and verify the key outside source control.
Markdown contains a bot-check page The target blocked automated access Detect the interstitial, do not index it, and use a supported browser-rendered workflow where appropriate.
Important text is missing Content is rendered by JavaScript after the initial response Use Firecrawl Scrape or another browser-capable fetch.
429 responses Rate limit exceeded Queue requests, lower concurrency, honor retry timing, or use a key/tier with a higher documented limit.
RAG answers cite the wrong page Chunks lost provenance Attach source URL and retrieval time to every chunk and display them in the answer.
Duplicate documents after refresh URL variants were treated as separate sources Normalize URLs and use a deterministic content or URL hash.

Or skip the browser setup

When your RAG system benefits from a visual page alongside Markdown, ScreenshotNeo returns a screenshot or PDF from one GET request. It can accept consent banners, remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture, with each step switchable. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers report the page verdict and billing status. Its MCP server gives Claude, Cursor, and other MCP clients take_screenshot, get_page_info, and capture_pdf tools.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests; r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90); open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for full-page capture, CSS selectors, device presets, custom headers and cookies, waiting rules, blocking controls, caching, signed links, async jobs, bulk capture, and PDF options. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000.

Create a free ScreenshotNeo account.

FAQ

Is a website-to-Markdown API the same as a web scraper?

It is a specialized scraper that returns cleaned, model-friendly content and often metadata instead of the original page chrome.

Should I store Markdown or embeddings?

Store the original Markdown and metadata as the source of truth, then derive chunks and embeddings from it.

When should I crawl instead of scrape?

Use a single scrape for a known page and a crawl when you need a linked set of pages from one site.

How do I keep answers current?

Schedule refreshes, retain retrieval timestamps, replace changed documents, and keep the last successful version when a refresh fails.

Can screenshots replace Markdown?

No. Screenshots preserve visual context; Markdown is usually better for text retrieval. Store both when layout or charts matter.