How to Convert Any Webpage to Markdown for Your LLM
Convert webpages into clean Markdown for ChatGPT and other LLMs with Jina Reader, Pandoc, browser rendering, validation, and troubleshooting.
Use a two-stage pipeline: first fetch the page and isolate its meaningful content, then convert that cleaned HTML to Markdown. For a quick hosted conversion, prepend https://r.jina.ai/ to the page URL. For a reproducible local workflow, save the HTML and run Pandoc. JavaScript-heavy pages need a browser-capable fetcher before conversion.
Choose the right conversion route
| Route | Best for | Strength | Limitation |
|---|---|---|---|
| Jina Reader URL prefix | One-off conversions and simple integrations | Fetches, extracts article content and returns Markdown in one request | Hosted behavior, access and limits can change |
| Pandoc | Repeatable local builds from saved HTML | Documented HTML input and Markdown output | Does not fetch URLs or identify the article region |
| Browser plus Pandoc | Client-rendered applications | Executes JavaScript before saving HTML | Slower and more resource-intensive than a raw HTTP request |
| ReaderLM-v2 or schema extraction | Structured fields or JSON for downstream systems | Can produce Markdown or structured output | Model output requires validation |
Fastest method: Jina Reader
Jina Reader documents a URL-prefix pattern that converts a page into LLM-friendly input:
https://r.jina.ai/https://example.com/article
Open that URL in a browser, or request it from code. The response is Markdown, so you can save it directly.
cURL
curl -L "https://r.jina.ai/https://example.com/article" -o page.md
Python
import requests
url = "https://r.jina.ai/https://example.com/article"
response = requests.get(url, timeout=90)
response.raise_for_status()
with open("page.md", "w", encoding="utf-8") as file:
file.write(response.text)
Node.js
const response = await fetch(
"https://r.jina.ai/https://example.com/article"
);
if (!response.ok) {
throw new Error(`HTTP ${response.status}`);
}
const markdown = await response.text();
await Bun.write("page.md", markdown);
// For Node.js without Bun, use fs.promises.writeFile("page.md", markdown).
What the hosted pipeline does
A useful conversion is more than replacing HTML tags with Markdown syntax. The documented pipeline can choose between a lightweight curl-impersonate fetch and headless Chrome. Chrome executes JavaScript, while the lightweight path is faster when the content is already present in the raw HTML. It then applies Mozilla Readability-style extraction to remove navigation and boilerplate before producing Markdown.
- Fetch the URL with a suitable engine.
- Identify the main article or content region.
- Remove repeated navigation, ads, cookie notices and unrelated recommendations.
- Convert headings, paragraphs, lists, links, tables and code blocks to Markdown.
- Inspect the result before putting it in an LLM prompt.
Local conversion with Pandoc
Pandoc is a documented command-line option when you already have the HTML. It converts formats; it does not fetch the URL or decide which part of a page is the article.
Download HTML and convert it
curl -L "https://example.com/article" -o page.html
pandoc -f html -t gfm page.html -o page.md
The gfm target is convenient for GitHub-compatible Markdown. Use the standard Markdown target when your consumer expects Pandoc Markdown.
Drop noisy div and span wrappers
pandoc -f html-native_divs-native_spans -t markdown page.html -o page.md
This option is useful when layout-only div and span elements are polluting the output. It does not replace article extraction; for that, save only the meaningful content or run a Readability-style extractor first.
Handling JavaScript-rendered pages
A raw HTTP request can return an HTML shell while the article is inserted later by JavaScript. In that case, use a real browser, wait for a meaningful selector or network idle, then save the rendered DOM.
Playwright example
import { chromium } from "playwright";
import { writeFile } from "node:fs/promises";
const browser = await chromium.launch();
const page = await browser.newPage();
await page.goto("https://example.com/article", { waitUntil: "networkidle" });
await page.locator("main, article").first().waitFor({ timeout: 15000 });
await writeFile("rendered.html", await page.content(), "utf8");
await browser.close();
Then convert the saved file:
pandoc -f html -t gfm rendered.html -o page.md
Prefer a specific article selector when one exists. Waiting only for networkidle can hang on pages with analytics or long-lived connections; a selector plus a bounded timeout is more predictable.
Extract the meaningful content before conversion
Readability-style extraction removes common boilerplate, but unusual layouts can still cause omissions. When accuracy matters, compare the extracted result with the source page.
- Keep the title and publication metadata.
- Keep the article’s headings and section order.
- Preserve ordered and unordered lists.
- Preserve links and link text.
- Preserve tables when they carry meaning.
- Keep code blocks with their language labels when available.
- Decide whether comments, related posts and author bios belong in your use case.
Validate Markdown before sending it to an LLM
Inspect the generated file instead of assuming a successful HTTP response means a complete conversion.
sed -n '1,120p' page.md
rg -n '^#{1,6} |^[-*] |^```|^\| ' page.md
| Check | What to look for |
|---|---|
| Title | The expected page title appears once |
| Headings | Heading levels and order match the source |
| Content | Important paragraphs are present and not duplicated |
| Lists | Nested and ordered lists retain their structure |
| Links | Destination URLs are still attached to the correct text |
| Tables | Headers, rows and cell relationships remain understandable |
| Code | Fences, indentation and language labels are intact |
| Images | Useful alt text or source URLs are retained when needed |
Put the source URL and retrieval date in front matter so a model can cite or refresh the source later:
---
source_url: https://example.com/article
retrieved_at: 2026-10-01
---
Preparing Markdown for an LLM prompt
Remove content unrelated to the question: repeated navigation, cookie banners, recommendation rails and comments when they are not part of the task. Keep headings because they provide structure. For long pages, split at heading boundaries and retain the source URL in every chunk.
Answer the question using only the source below.
If the source does not contain the answer, say so.
Source URL: https://example.com/article
# Extracted page
[page.md contents here]
Do not claim that conversion automatically reduces tokens by a fixed percentage; the sources provide no general benchmark. Measure your own pages if token cost matters.
Or skip the browser setup
If you need a visual capture for an LLM workflow, ScreenshotNeo provides a website screenshot API. It is useful when the page’s visual state matters alongside extracted text; it does not turn HTML into Markdown by itself.
One GET request returns a PNG, JPEG, WebP or PDF. See the ScreenshotNeo API documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners, newsletter popups and chat widgets are removed before the shot.
- Bot checks, blank pages and failed loads are never billed.
- An MCP server lets AI agents take screenshots with Claude, Cursor or another MCP client.
- The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Create a free ScreenshotNeo account and start with 1,000 screenshots a month without a card.
Troubleshooting
The Markdown is empty or nearly empty
Cause: the server returned a JavaScript shell, blocked the request, or the extractor could not identify the content. Fix: render the page with Playwright, wait for the article selector, save the rendered HTML and run Pandoc; also inspect the HTTP status and response body.
Headings or paragraphs are missing
Cause: Readability-style extraction removed an unusual layout region. Fix: compare the extracted HTML with the original, target the article container explicitly, or convert a cleaned fragment locally.
Navigation and cookie text remain
Cause: the page structure does not match common boilerplate rules. Fix: remove those nodes before Pandoc, or use a selector-based extraction step.
Tables look broken
Cause: complex tables, merged cells or responsive layouts do not map cleanly to Markdown tables. Fix: keep the table as HTML, simplify it before conversion, or extract its rows into structured JSON.
Code indentation changed
Cause: the source used layout elements instead of semantic code blocks. Fix: preserve pre/code nodes during extraction and inspect the generated fences.
The browser job never finishes
Cause: long-lived analytics, WebSockets or streaming requests prevent network-idle detection. Fix: wait for a specific selector with a finite timeout, then continue.
Images are absent
Cause: lazy loading, CSS backgrounds or inaccessible image URLs. Fix: scroll or trigger lazy loading in the browser, preserve alt text, and decide whether the LLM needs the image itself or only its description.
Performance, reliability and cost
- Use a raw HTTP fetch when the content is present in initial HTML; it is generally faster and cheaper than browser rendering.
- Use a browser only for pages that need JavaScript, interaction or lazy loading.
- Cache fetched HTML and generated Markdown when the source does not change frequently.
- Set bounded request and browser timeouts, and retry transient failures with backoff.
- Store the source URL and retrieval timestamp with every document.
- Validate content before paying for downstream LLM tokens.
- Respect the site’s terms, robots policies and access controls.
FAQ
What is the easiest URL-to-Markdown converter?
For a one-off conversion, prepend https://r.jina.ai/ to the page URL. Use a local browser plus Pandoc when you need control and reproducibility.
Can Pandoc convert a URL directly?
Pandoc converts input files and streams; fetch the URL first, save the HTML, then run Pandoc.
Why does a page work in my browser but not with curl?
Your browser executes JavaScript and may provide cookies or headers that a raw HTTP request lacks. Use a browser-capable fetcher when the content is client-rendered.
Should I send HTML or Markdown to an LLM?
Markdown is usually easier to inspect and chunk, but validate the conversion first. Keep HTML when table structure or visual semantics would otherwise be lost.
Is model-assisted extraction always accurate?
No. Structured or model-generated output can be useful, but compare it with the source whenever omissions would change the answer.
Sources
- Jina Reader URL-prefix workflow and reader output.
- Jina Reader repository architecture and extraction details.
- Pandoc manual for HTML and Markdown conversion options.


