ScreenshotNeo

BlogHow-to

How to Convert a Website to Markdown

Convert a single webpage to Markdown with Pandoc, a browser based extractor, or an API. Learn which method fits dynamic pages, how to save the output, and what to check.

By the ScreenshotNeo team4 October 20269 min read

To convert one webpage to Markdown, either give its HTML to Pandoc, use a browser based URL converter, or call an extraction API and save the returned Markdown as a .md file. Use Pandoc for local HTML or a simple URL; use a browser rendering extractor when the page depends on JavaScript; use an API when you need a repeatable pipeline. “A website” usually means one URL in these examples. Converting an entire site requires discovering and processing multiple pages separately.

Choose a conversion method

Situation Good starting point What to watch
You have an HTML file Pandoc locally Review tables and complex layout after conversion.
You have one public URL and want a quick result Browser based converter Check whether the page is rendered in a browser and whether access is supported.
The page fills in after JavaScript runs Browser rendering or a wait-for-element extractor Wait for the actual content, then inspect for missing sections.
You need repeatable conversion in an application An extraction API Check current syntax, limits, pricing, and output quality.
You need a site archive or migration A crawler and a per-page conversion workflow Plan URL discovery, pagination, duplicates, failures, and link rewriting.

Markdown conversion preserves useful structure, but it is not necessarily lossless. Pandoc notes that its intermediate document model cannot express every feature of every input format, and complex tables may not fit its simple model. Always compare the result with the source page. Pandoc User’s Guide

Convert an HTML file with Pandoc

Install Pandoc using the official installation instructions, save the page source as page.html, and run:

pandoc -f html -t markdown page.html -o page.md

-f html selects the HTML reader, -t markdown selects Markdown output, and -o writes the result to a file. To choose a Markdown flavor supported by your publishing destination, use a Pandoc output format such as gfm where appropriate:

pandoc -f html -t gfm page.html -o page.md

Use GitHub Flavored Markdown only if the target platform expects it. Check the Pandoc format options for available writers and extensions.

Convert a URL directly with Pandoc

Pandoc’s official demo shows reading a web page as HTML and writing a text output. Adapt the URL and filename:

pandoc -s -r html https://example.com/article -o article.md

This is convenient when the required content is present in the fetched HTML response. It may miss content inserted only after client side JavaScript executes. If the output lacks the article body, use a browser rendering extractor or an API with a wait-for-content control. Pandoc demos

Use a browser based converter for a one-off page

  1. Choose a converter that accepts a URL and says whether it renders pages in a browser.
  2. Enter the public page URL and start conversion.
  3. Copy or download the Markdown output.
  4. Save it as UTF-8 text with a .md extension.
  5. Compare the output with the source, especially headings, links, images, code, and tables.

For example, Firecrawl’s website-to-Markdown converter describes a URL workflow that fetches, renders, extracts, and lets you copy or download Markdown. Its free tool is for publicly accessible pages; its FAQ says login protected or paywalled content is not accessible through that free tool. Do not treat a converter as a way to bypass access controls. Use authenticated access only when you are authorized, and check the service’s current policies and capabilities.

Automate conversion with an API

An API pipeline usually takes a URL, fetches or renders the page, extracts its content, returns Markdown, and writes that text to a file or downstream system. Firecrawl’s vendor tutorial demonstrates the Python SDK and returns the Markdown in document.markdown:

from firecrawl import Firecrawl

app = Firecrawl(api_key="YOUR_FIRECRAWL_API_KEY")
document = app.scrape(
    "https://example.com/article",
    formats=["markdown"],
)
markdown = document.markdown

with open("article.md", "w", encoding="utf-8") as output:
    output.write(markdown or "")

Install the SDK and confirm the current method signature in the Firecrawl documentation before using this in production; API syntax can change. The cited Firecrawl tutorial stated a free allowance of 1,000 credits per month and one credit per page scraped when checked on 2026-10-03. Those are vendor published plan details, not a guarantee of current pricing or limits; verify them before choosing a plan. Firecrawl Python tutorial

Cloudflare Browser Run also documents a Markdown endpoint that accepts a URL or raw HTML and returns Markdown. Its raw HTML path is useful when your own system already fetched the page, though it is a developer API rather than the shortest route for a one off conversion. Check the Browser Run documentation for current request details.

Extract Markdown with Jina Reader

Jina Reader’s documented pattern prefixes a page URL with https://r.jina.ai/ to request LLM friendly page content. Its interface documents controls for waiting for selected elements, extracting selected elements, and removing selectors such as navigation or footers. Inspect the returned text for your target page; selectors and wait conditions depend on the site. See the Jina Reader documentation.

curl -L "https://r.jina.ai/https://example.com/article" -o article.md

Build a small conversion pipeline

For an API driven workflow, keep fetching and file writing separate from quality checks. This Python example uses the Firecrawl SDK shape shown in its tutorial; install the current SDK as documented by Firecrawl and provide an API key through an environment variable in your application:

import os
from pathlib import Path
from firecrawl import Firecrawl

url = "https://example.com/article"
client = Firecrawl(api_key=os.environ["FIRECRAWL_API_KEY"])
document = client.scrape(url, formats=["markdown"])
markdown = document.markdown
if not markdown:
    raise RuntimeError("The extractor returned no Markdown")

Path("article.md").write_text(markdown, encoding="utf-8")

For a site rather than one page, first define how URLs are discovered, then process each permitted URL independently. Store the source URL alongside each result, use deterministic filenames or a URL-to-path map, record failures for retry, and prevent the same canonical URL from being processed repeatedly. Do not assume that a single page conversion endpoint crawls a whole domain.

Handle JavaScript, clutter, and access constraints

  • JavaScript rendered content: If the source HTML has only a shell, choose a browser rendering tool. When supported, wait for a selector that appears only after the main content loads rather than relying on a fixed delay.
  • Navigation and repeated boilerplate: Use an extractor’s main content or selector controls where available. Check that these do not remove article headings, code blocks, or important callouts.
  • Cookie prompts and overlays: A prompt may obscure a browser capture or interfere with extraction. Prefer a workflow that can handle consent and remove common overlays, then inspect the text for omissions.
  • Login or paywall: Use only access you are entitled to use. A public converter may not support authenticated pages; an API’s support for headers or cookies does not grant permission to access restricted content.
  • Images and links: Markdown may contain remote image links rather than embedded image data. Check whether those URLs remain reachable and whether relative links need to be made absolute.
  • Tables: Inspect wide, nested, or merged cell tables manually. Conversion may flatten them or lose relationships.

Quality check the Markdown

  1. Confirm the title, main headings, and main body are present in order.
  2. Check links point to useful destinations and resolve relative URLs if you are moving the file.
  3. Review images: decide whether to keep references, download assets, or omit them.
  4. Check code blocks, lists, tables, footnotes, and special characters.
  5. Look for navigation, cookie notices, repeated footer text, or other content that drowned out the page.
  6. If sections are missing, determine whether they were absent from the initial HTML or hidden behind interaction, pagination, or lazy loading.

Or skip the browser setup

If your goal is to capture a clean visual reference of a webpage alongside a Markdown extraction, ScreenshotNeo is a website screenshot API and MCP server for developers. It returns PNG, JPEG, WebP, or PDF from one GET request. It does not convert a page to Markdown; use a content extractor for Markdown and a screenshot when you also need the rendered visual. The API supports custom CSS and JavaScript, waiting for a selector, delay or network idle, blocking request types, and other capture controls. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie and consent banners are accepted like a visitor, and 60+ known consent platforms, newsletter popups, and chat widgets are removed before capture; each step can be turned off.
  • Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Response headers report the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
  • The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; all listed features are on every plan.

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

Performance, reliability, and cost

  • Local Pandoc: Avoids a third party extraction API for local files and makes the command easy to repeat. For URL input, results depend on what can be fetched as HTML; it does not make client side JavaScript execute.
  • Browser rendering: Can expose content generated by JavaScript, but rendering and waiting add work. A precise wait condition can avoid both premature extraction and unnecessarily long fixed delays.
  • APIs: Simplify integration and consistent output handling. For production, handle timeouts, non-success responses, empty Markdown, retries for transient failures, and provider limits. Keep API keys out of source control and logs.
  • Batch work: Bound concurrency, preserve a retry queue, and avoid refetching unchanged pages when your workflow permits caching. Record the source URL and retrieval time for each result.
  • Cost: A local conversion has no per-page vendor credit charge, while hosted extraction may have plan or usage limits. Compare current pricing against expected pages, retries, and rendering needs. Treat published free allowances as changeable.

Troubleshooting

Symptom Likely cause What to do
Markdown is empty or only contains a page shell Content is inserted after JavaScript executes, or the fetch was blocked. Use browser rendering; wait for the main content selector; verify the page is accessible with your authorized request.
Headings or paragraphs are missing Extraction selected the wrong region, or content is loaded on interaction. Inspect the source and extracted region; adjust selectors or use a rendered page.
Tables look broken Complex table structure is not represented cleanly by the target Markdown flavor. Try a richer supported output format or manually rewrite the table after comparing it with the source.
Images show as broken links Relative paths were copied without a base URL, or the image requires access. Resolve relative paths against the page URL and check that the assets are accessible to the reader.
Pandoc cannot read the URL Network access, redirects, or server behavior prevented a usable HTML input. Fetch the page through an authorized route, save its HTML, and run Pandoc on the local file.
API returns an error or times out Credentials, request shape, provider limits, or page load time may be the issue. Check the current API docs, validate the key and URL, increase the timeout only when justified, and retry transient failures with a limit.
Output includes cookie text, navigation, or footer boilerplate The extractor treated the full page as content. Use main-content extraction or selector removal where supported, then verify important content remains.

Frequently asked questions

Does converting a URL save the images in the Markdown file?

Usually the Markdown contains image references, not the image bytes. Download assets separately if you need an offline archive.

Can I convert a whole domain with one command?

The Pandoc examples convert one input page. A domain archive needs URL discovery and a workflow that processes each page.

Which output flavor should I choose?

Choose the flavor supported by the destination where the Markdown will be rendered. If you do not know the target yet, use a common Markdown output and review extensions such as tables and footnotes.

Can I convert a page I can view after logging in?

Only use a method and credentials that you are authorized to use, and confirm the provider supports that workflow. A public converter may not handle protected pages.