Convert Any Website to Markdown with an API
Compare URL-to-Markdown APIs, get runnable cURL, Python, and Node.js examples, and learn how rendering, selectors, crawling, and retries affect results.

To convert a website to Markdown with an API, send the page URL to a service that fetches the page, renders JavaScript when needed, removes navigation and other boilerplate, and returns Markdown. For a quick prototype, prepend https://r.jina.ai/ to a public URL. For rendered pages with browser controls, use a browser automation API such as Browserless. For a whole domain, use a crawler such as Firecrawl Crawl rather than repeatedly guessing page URLs.
The key decision is not just which HTML-to-Markdown converter to use. The result depends on whether the service can access the page, render its content, wait for it to appear, and select the useful part of the page. This guide covers single-page extraction, site crawls, runnable requests, controls, failure handling, and when a screenshot is the more useful output.
1. Choose an API for the job
First decide what you need back and how much of the site you need. A single article, a JavaScript-heavy page, and an entire documentation site are different workloads.
| Approach | Best fit | What to consider |
|---|---|---|
| Jina Reader URL prefix | Prototype or fetch one URL as readable content | Simple request; browser fetching, selectors, waits, exclusions, and output modes are documented controls. |
| Browserless GraphQL | Applications already using GraphQL or needing browser-level steps | Navigate, inspect status, then request Markdown from the rendered page; supports selector, timeout, and visibility options. |
| Firecrawl Scrape | One page with Markdown, structured data, links, or screenshots | Browser rendering and clean output are part of the product workflow. |
| Firecrawl Crawl | Discover and process multiple pages on a domain | Plan URL discovery, deduplication, rate limits, storage, and incremental refreshes. |
| ScreenshotNeo | A visual screenshot or PDF rather than Markdown text | One GET request captures a URL; it is not a Markdown extraction API. |
For direct URL-to-content conversion, start with a single-page reader. Choose a browser API when you need to control the rendered page state. Choose a crawl endpoint for site-wide ingestion. If the downstream task needs layout, charts, or a visual record, capture an image or PDF instead of expecting Markdown to preserve presentation.
2. Convert a URL with Jina Reader
The simplest request uses the target URL after the Reader prefix. This cURL example saves the response as Markdown text:

curl "https://r.jina.ai/https://www.example.com" \
-H "Accept: text/plain" \
-o page.md
Python equivalent:
import requests
url = "https://r.jina.ai/https://www.example.com"
response = requests.get(url, headers={"Accept": "text/plain"}, timeout=60)
response.raise_for_status()
with open("page.md", "w", encoding="utf-8") as f:
f.write(response.text)
Node.js equivalent using built-in fetch:
const url = "https://r.jina.ai/https://www.example.com";
const response = await fetch(url, {
headers: { Accept: "text/plain" },
signal: AbortSignal.timeout(60_000),
});
if (!response.ok) throw new Error(`Reader failed: ${response.status}`);
await Bun.write("page.md", await response.text());
In Node versions without Bun.write, write the response with node:fs/promises:
import { writeFile } from "node:fs/promises";
const body = await response.text();
await writeFile("page.md", body, "utf8");
Check the service documentation for current authentication and request headers before moving a prototype into production. Jina documents Markdown, HTML, text, screenshot, frontmatter, and markdown+frontmatter output modes, plus browser-engine fetching, selector scoping, wait-for selectors, exclusions, and cache controls. Set only the controls your page needs: broad waits slow down batches, and an overly narrow selector can produce an empty result.
3. Scope extraction and wait for dynamic content
A page-wide conversion may include a site header, footer, cookie notice, recommendations, and navigation. If your target is an article, scope extraction to its main container where the provider supports a CSS selector. Exclude known noise when a selector is more reliable than guessing the main region.

For pages that render content after the initial document load, use a browser-based fetch mode and wait for a meaningful selector. A fixed delay can work for a predictable page, but it wastes time on fast responses and can still be too short on slow ones. Waiting for the content element is usually a better signal. Confirm that the selector is stable across the pages in your batch.
Jina’s Reader controls include browser fetching, target selector, wait-for selector, exclusions, output format, and cache-related options. Use the provider’s current documentation for exact header names and accepted values; configuration can change. The design pattern is consistent: select the article body, wait for a real content marker, and exclude repeated elements only when they pollute the result.
Browser rendering and Markdown conversion solve separate problems. Rendering makes client-generated DOM available; extraction decides which parts become output. If a page requires a logged-in session, custom cookies, or a particular region, verify that the service supports those controls and that you are authorized to access the content.
4. Use Browserless when you need browser control
Browserless documents a GraphQL pattern that navigates to a URL and then converts the page to Markdown. Its Markdown operation accepts a selector, timeout, and visibility setting; the documented default timeout is 30,000 milliseconds.
mutation Markdownify {
goto(url: "https://example.com") {
status
}
markdown {
markdown
}
}
Send the mutation as the GraphQL request body to the Browserless endpoint configured for your account, with the authentication method and content type required by your Browserless setup. A generic cURL shape is:
curl "YOUR_BROWSERLESS_GRAPHQL_ENDPOINT" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_TOKEN" \
--data '{"query":"mutation Markdownify { goto(url: \"https://example.com\") { status } markdown { markdown } }"}'
Replace the endpoint and authentication header with the values from your account documentation. Inspect the navigation status before trusting the conversion. If your application already drives a browser, using the same browser session to retrieve Markdown can reduce duplicate fetching and lets you make a selector or visibility decision based on the rendered DOM.
5. Scrape one page or crawl a site with Firecrawl
Firecrawl Scrape is intended for a single URL and can return Markdown, structured data, links, or screenshots. Its product description says it renders pages in a real browser and removes navigation, footers, ads, and tracking. Firecrawl Crawl extends the pattern to discover and process subpages across a domain, producing a corpus suited to search or retrieval-augmented generation (RAG).
Use Scrape when you already know the URL. Use Crawl when page discovery is part of the task. Check the current Firecrawl API reference for the exact endpoint, authentication, request schema, crawl limits, and output fields; those operational details can change and are not reproduced here. Store the source URL and retrieval time alongside each result so you can trace and refresh documents later.
Whole-site ingestion checklist
- Define the allowed hostnames and URL paths before crawling.
- Decide how query parameters, trailing slashes, redirects, and canonical links map to one document.
- Deduplicate URLs and content so tracking variants do not create repeated chunks.
- Set concurrency and retry limits to fit provider limits and the source site’s access rules.
- Persist crawl status and failures; resume from failed URLs instead of restarting everything.
- Record content hashes or modification metadata where available to avoid unnecessary reprocessing.
- Plan refresh cadence and deletion handling so removed source pages do not remain in your index indefinitely.
6. Select output and preserve useful metadata
Markdown is convenient for plain text and headings, but it is not a complete web archive. Tables, code blocks, image references, links, and metadata may be represented differently by different providers. Validate representative pages before ingesting thousands of URLs.
Choose plain Markdown when downstream systems need readable text. Use frontmatter when you want metadata such as source URL or title carried with the document. Choose JSON when your application needs a stable structure, such as named fields or extracted entities. Use HTML when you need more of the original markup. A screenshot preserves visual appearance but does not provide editable semantic text.
For RAG, keep provenance outside the text body as structured fields: canonical URL, fetch timestamp, title, content hash, and crawl job. Split the Markdown on headings where possible, preserve heading paths in each chunk, and avoid splitting code blocks or tables mid-element. Conversion output should be treated as source material to validate, not automatically as verified truth.
7. Reliability, performance, and cost
Every remote extraction has at least three variable stages: URL fetch, page rendering, and conversion. JavaScript rendering and selector waits generally add work, while caching can reduce repeated requests when the page has not changed. A single average latency does not predict tail latency for a slow or blocked target.
Jina’s published operational table in the research dossier lists 20 requests per minute without an API key, 500 RPM with a free key, and up to 5,000 RPM with a premium key, plus 7.9 seconds average latency. These figures are volatile; verify the current Jina documentation and your account limits before launch. Do not use a published average to set a hard timeout or promise a batch completion time.
Use bounded concurrency, exponential backoff with jitter for transient errors, and a maximum attempt count. Respect Retry-After where returned. Retry network failures, rate limits, and selected server errors; do not repeatedly retry permanent authorization failures or a page that explicitly denies access. Cache based on the freshness your use case permits. Measure cost per successfully processed page, not just requests submitted, and account for crawl discovery, browser execution, storage, and reprocessing.
For source access, follow the website’s terms and access controls. Jina says Reader does not actively circumvent anti-bot systems or access controls; do not treat any extractor as a way to bypass those restrictions. You remain responsible for rights in content you fetch and store.
8. Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Markdown is empty or only has a title | Content is rendered later, selector does not match, or access is denied. | Try the browser rendering mode, wait for a content selector, verify the selector on the rendered DOM, and inspect response status. |
| Navigation and repeated page chrome dominate | Conversion is page-wide or the target selector is too broad. | Scope to the article/main region and exclude stable navigation or ad containers. |
| Text is stale | A cached response is being reused. | Review cache controls and TTL; disable or refresh cache for pages requiring current content. |
| Intermittent timeout | Slow origin, delayed scripts, heavy assets, or an unrealistic wait. | Set a bounded timeout appropriate to the page, wait for a content signal, reduce unnecessary browser work, and retry transient failures with backoff. |
| HTTP 401 or 403 | Missing/invalid API credentials or the target denies access. | Check provider authentication separately from target-site authorization. Use supported credentials only for content you may access. |
| HTTP 429 | Provider rate limit or too much concurrency. | Reduce request rate, honor retry timing, queue work, and verify the current account quota. |
| Duplicate pages in a crawl | Query strings, redirects, or URL variants map to the same content. | Normalize URLs, follow canonical metadata where available, and deduplicate by normalized URL and content hash. |
| Broken tables or code blocks | Source markup is irregular or conversion formats differ. | Test a sample set, choose HTML/JSON if Markdown loses required structure, or post-process with format-aware parsing. |
9. Or skip the browser setup
If the deliverable is a screenshot or PDF rather than Markdown, ScreenshotNeo captures a URL with one GET request. The API supports PNG, JPEG, WebP, and PDF; see the ScreenshotNeo API documentation for parameters and response details.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners, popups, and chat widgets are removed before the shot; each step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers report the page verdict and billing status.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for AI agents. - The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.
Sign up for ScreenshotNeo’s free 1,000 screenshots a month, with no card required.
10. Frequently asked questions
Can I convert a whole website from one URL?
Usually not with a single-page reader. Use a crawl workflow that discovers links, tracks completed URLs, deduplicates variants, and observes provider and source-site limits.
Does Markdown preserve the original layout?
No. Markdown captures semantic text and some structure, but not the page’s exact visual layout, styling, or interactive state. Use a screenshot or PDF when appearance matters.
Should I render every page in a browser?
No. Rendering can recover client-generated content, but adds latency and resource use. Start with the lighter fetch path and enable rendering for pages that need it.
Can an API extract content behind a login?
Only if the provider supports the required authenticated session and you have permission to access and process that content. Confirm credential handling and retention terms before sending secrets.
Which output should I store for an AI knowledge base?
Store Markdown or structured fields with URL and retrieval metadata. Keep the original source reference so chunks can be audited and refreshed; retain HTML only when its extra structure is useful.


