ScreenshotNeo

BlogHow-to

How to Convert a Web Page to Markdown: A Developer’s Guide

Learn how to fetch a web page, extract its main content, and convert HTML to Markdown with JavaScript, Python, or a hosted URL API.

By the ScreenshotNeo team4 October 202610 min read

To convert a web page to Markdown, handle three separate steps: fetch the page, isolate the content you want, then convert its HTML structure into Markdown. If you already have an HTML fragment, you can skip fetching. If the page renders its content with JavaScript, a basic HTTP request may not include that content; use a browser-rendered fetch or a service that supports rendering.

Choose the method based on your starting point and runtime: use Turndown for HTML in JavaScript, MarkItDown for a Python or CLI workflow, or a hosted URL conversion API when you want a service to fetch and optionally render a public URL. None of these steps guarantees perfect conversion. Inspect headings, links, tables, code, images, metadata, and relative URLs in the output.

1. Decide what you are converting

First identify your input. “Convert a web page” can mean several different jobs:

Input What you need to do Typical choice
HTML string or saved page Select the content and serialize its structure Turndown in JavaScript or MarkItDown in Python
Live, mostly server-rendered URL Fetch HTML, extract content, then convert HTTP client plus converter, or hosted URL API
Live page whose content appears after JavaScript runs Render the page in a browser, then extract and convert Browser automation or a hosted API with a render option
Article mixed with navigation, cookie notices, or widgets Identify and select the main content before conversion Site-specific selector or a suitable extraction step

HTML-to-Markdown conversion and main-content extraction are different tasks. A converter serializes the HTML it receives; do not assume it will identify the article body on every arbitrary site. Test representative pages from your target domain and decide which content belongs in the result.

2. Convert existing HTML in JavaScript with Turndown

Turndown is a JavaScript HTML-to-Markdown converter. It is a good fit when your application already has an HTML string or DOM node. This runnable example fetches a page with Node.js, selects a main-content element, and converts that element. The selector is site-specific: replace it with one that matches the page you are processing.

npm install turndown jsdom
// save as convert.mjs; run with: node convert.mjs https://example.com
import TurndownService from 'turndown';
import { JSDOM } from 'jsdom';
import { writeFile } from 'node:fs/promises';

const url = process.argv[2];
if (!url) throw new Error('Usage: node convert.mjs https://example.com');

const response = await fetch(url, {
  headers: { 'user-agent': 'MarkdownConverter/1.0' },
  signal: AbortSignal.timeout(30000),
});
if (!response.ok) throw new Error(`Fetch failed: HTTP ${response.status}`);

const html = await response.text();
const dom = new JSDOM(html, { url });
const { document } = dom.window;
const content = document.querySelector('article, main') ?? document.body;

const turndown = new TurndownService({
  headingStyle: 'atx',
  codeBlockStyle: 'fenced',
  bulletListMarker: '-',
});
const markdown = turndown.turndown(content);
await writeFile('page.md', markdown + '\n', 'utf8');
console.log('Wrote page.md');

This example uses ordinary HTTP fetching. It does not execute page JavaScript, and the fallback to body can include navigation or other unrelated elements. To improve results, inspect the page DOM and use a precise selector. If the desired content is inserted by client-side JavaScript, use a browser-rendered fetch before selecting and converting the content.

Turndown supports configurable rules for HTML elements. Use custom rules when a particular element needs a site-specific Markdown representation, and review the generated output rather than assuming a rule covers every page variation.

3. Convert HTML or a URL in Python with MarkItDown

Microsoft MarkItDown is a Python project for converting files and other input formats, including HTML, into Markdown for text analysis. Its README recommends a virtual environment and documents Python 3.10 through 3.14; confirm current requirements in the repository before installing because supported versions can change.

python -m venv .venv
source .venv/bin/activate
python -m pip install 'markitdown[all]'

To convert a local HTML file with the CLI:

markitdown page.html > page.md

To fetch a public page and convert the saved HTML in Python, use an HTTP client and MarkItDown. Install requests if needed. This path still does not execute JavaScript.

python -m pip install requests
from pathlib import Path
from urllib.parse import urlparse
import requests
from markitdown import MarkItDown

url = 'https://example.com'
parsed = urlparse(url)
if parsed.scheme not in {'http', 'https'} or not parsed.hostname:
    raise ValueError('Use an absolute http or https URL')

response = requests.get(url, timeout=(5, 30))
response.raise_for_status()
response.encoding = response.apparent_encoding
Path('page.html').write_text(response.text, encoding='utf-8')

converter = MarkItDown()
result = converter.convert('page.html')
Path('page.md').write_text(result.text_content.rstrip() + '\n', encoding='utf-8')
print('Wrote page.md')

MarkItDown aims to preserve document structure for text analysis; its project documentation says it may not be the best choice for high-fidelity, human-facing document conversion. Treat the output as a starting point and review it for your use case.

4. Convert a live URL with a hosted API

A hosted URL-to-Markdown API can combine fetching with conversion and may offer browser rendering for client-rendered pages. The reviewed markitdown.ai URL conversion documentation describes POST /v1/convert/url, API-key authentication, and render modes auto, force, and skip. These settings and commercial terms are vendor-specific; check the current documentation and plan before building against them.

The documentation describes auto as rendering when fetched HTML has no readable content, force as requesting rendering, and skip as bypassing it. Follow the provider’s current authentication and request schema; do not assume another service uses the same endpoint or options.

The provider overview describes requests that can complete synchronously or return an asynchronous conversion to poll or follow with a webhook. It also describes page-based credits, including standard and OCR pages at one credit per page and AI image understanding at five credits per image for paid-plan accounts. These figures are vendor-published and can change, so confirm current terms before estimating costs.

5. Inspect and improve the Markdown

Review output against the source page before using it downstream. Use this checklist:

  • Headings: Are heading levels in a sensible hierarchy, without navigation headings mixed into the article?
  • Lists: Are ordered and unordered lists intact, including nested items?
  • Links: Do link labels and destinations survive? Resolve relative URLs against the original page URL if your downstream reader needs absolute links.
  • Tables: Does a simple table remain readable? Complex layouts may need a different representation or manual handling.
  • Code: Are inline code and code blocks preserved, and are fenced blocks appropriate?
  • Images: Are image references useful? Markdown does not preserve every detail of an HTML image, such as layout or responsive behavior.
  • Metadata: If you need title, description, canonical URL, or publication date, extract and store those deliberately; do not assume they will appear in body Markdown.
  • Dynamic content: Compare the result with what a browser displays, especially for pages that load content after navigation.

There is no universal accuracy score established by the documentation reviewed here. Page structure, extraction choices, and conversion rules all affect the result.

6. Choose between a converter, browser, or hosted service

Approach Use it when Trade-offs to consider
Turndown Your JavaScript workflow already has HTML or a DOM You own fetching, content selection, and any browser rendering. You can configure conversion rules.
MarkItDown You want a Python or CLI workflow, especially across document types It runs with the current process’s permissions. Its stated focus is text analysis rather than high-fidelity human-facing conversion.
Hosted URL conversion API You want a managed URL-fetch and conversion step, potentially with rendering Review authentication, subscription, credit use, asynchronous behavior, data handling, and current provider terms.
Browser automation plus converter The target page needs JavaScript execution or browser-specific behavior You operate the browser runtime, wait conditions, resource controls, and extraction logic.

These are capability distinctions, not a universal quality ranking. Choose based on where your HTML comes from, what the page requires, your runtime, and the operational boundary you need.

7. Security for server-side conversion

Converting a user-supplied URL or file is a server-side I/O operation. Microsoft warns that MarkItDown accesses resources with the current process’s privileges. Apply controls appropriate to your service:

  • Allow only the URL schemes your application needs, typically http and https.
  • Validate destinations and block private, loopback, link-local, and cloud metadata addresses where relevant; account for redirects and DNS resolution.
  • Set connection and total timeouts, response-size limits, and concurrency limits.
  • Restrict file paths and avoid passing untrusted paths into a process with broad filesystem access.
  • Run converters with least privilege and constrain network access for untrusted inputs.
  • Do not forward internal credentials or privileged headers to a user-provided destination.

These are necessary defensive measures, not a complete security review. Consult the project’s current security guidance and your platform’s controls.

8. Troubleshooting

Symptom Likely cause What to try
Markdown is empty or nearly empty The server returned a JavaScript shell, or the content selector matched nothing Inspect the fetched HTML and selector result. Use browser rendering if content appears only after scripts run.
Navigation and footer overwhelm the result The converter received the whole body rather than the article content Select a site-specific article, main, or known content container before conversion; verify it exists on all target templates.
Request fails or stalls Network failure, blocked request, slow origin, or no timeout Check the HTTP status and response body, use explicit connect/read timeouts, and retry only transient failures with a bounded policy.
Output has broken or relative links HTML links were relative to the source URL Resolve relative destinations against the final response URL and preserve the canonical source URL as metadata.
Tables or layout look strange The source uses complex HTML structures that Markdown cannot represent directly Inspect the source structure and choose a simplified representation or keep the relevant data in another format.
CLI or import is unavailable Package installed into a different Python environment, or current package requirements differ Activate the intended virtual environment, check the installed package and interpreter, then consult the current README.
Hosted conversion takes longer than the request window The job exceeded the synchronous wait window Use the provider’s documented asynchronous polling or webhook flow and handle job failure and timeout states.
Unexpected access to internal resources Untrusted URL, redirect, or network destination reached from a privileged server Stop the job, validate every redirect and resolved destination, and restrict the converter’s network and process permissions.

9. Performance, reliability, and cost

  • Performance: Fetching HTML and converting it locally avoids a separate hosted conversion request, but browser rendering adds startup and page-load work. Measure your own target pages; no comparative speed figures are established here.
  • Reliability: Set timeouts, bound retries, record HTTP status and conversion failures, and distinguish an empty source from an extraction failure. For long hosted jobs, use the documented async mechanism rather than holding a request open indefinitely.
  • Cost: Local libraries have no per-page API credit charge described here, though you still operate compute and possibly browsers. Hosted APIs may charge by page, image processing, subscription, or another current plan rule. Verify provider pricing before estimating batch cost.
  • Batching: Limit concurrency to protect your process and target sites. Cache results where permitted, keyed by normalized URL and relevant options, and record fetch time so stale Markdown is identifiable.

10. Or skip the browser setup

If your goal is to capture a page visually as well as process its content, ScreenshotNeo is a website screenshot API and MCP server. It returns a PNG, JPEG, WebP, or PDF from one GET request. A screenshot is an image or PDF, not Markdown; use an HTML-to-Markdown workflow above when Markdown is the required output.

For a rendered page capture, use the API call below. See the ScreenshotNeo API documentation for options and setup.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie and consent banners are accepted like a visitor, then more than 60 known consent platforms, newsletter popups, and chat widgets are removed before the shot; each step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Response headers say the page verdict and whether the shot was billed.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
  • 1,000 screenshots per month are free with no card. Paid plans start at $5 for 3,000 screenshots; every feature is on every plan.

Sign up for ScreenshotNeo and get 1,000 free screenshots a month with no card.

Frequently asked questions

Does converting HTML to Markdown extract the article automatically?

Not necessarily. Conversion represents supplied HTML; select or extract the meaningful content separately and verify it for each page template.

Can Markdown preserve every detail of a web page?

No. Markdown represents common text structure well, but complex layout, dynamic behavior, and some image or table details need review or another output format.

Should I use a screenshot to get Markdown?

A screenshot is visual output, not structured Markdown. Use a converter on HTML for Markdown; use a screenshot API when you need a rendered image or PDF record of the page.

What is the safest way to convert URLs submitted by users?

Validate URLs and redirects, constrain network destinations and process privileges, set time and size limits, and avoid exposing internal credentials to fetched pages.