ScreenshotNeo

BlogHow-to

How to Extract Markdown, Links, and Emails from a URL

Learn how to fetch a URL, resolve links correctly, and extract Markdown links and email addresses with Python, cURL, and Node.js.

By the ScreenshotNeo team1 October 20268 min read

Direct answer: a URL is only an address; it does not contain Markdown links or email addresses by itself. Fetch the resource, identify its format, parse the document with a Markdown-aware parser, and resolve each relative destination against the response URL. Extract email autolinks as mailto: destinations, then normalize and deduplicate the results.

Python’s urllib.parse handles URL components and relative-reference resolution, while CommonMark defines inline links, reference links, URI autolinks, and email autolinks. Python’s URL parsing documentation warns that its behavior cannot be claimed fully compliant with either RFC 3986 or the WHATWG URL standard, so parsing is not the same as validating a URL. The CommonMark specification defines the Markdown syntax you need to handle.

1. Decide what you are extracting

There are two separate jobs:

  1. URL parsing: split a URL into scheme, authority (called netloc by Python), path, query, fragment, and optional path parameters; resolve relative references such as ../docs.
  2. Document parsing: parse Markdown syntax and collect link destinations and email autolinks. A plain substring search misses reference links, escaped characters, nested text, and links whose destination is defined elsewhere.

First check the response’s Content-Type. Parse text/markdown or a known Markdown extension as Markdown. Parse text/html with an HTML parser instead. Do not feed arbitrary binary responses or PDFs into a Markdown parser.

2. Python: complete Markdown extractor

Install the HTTP client and CommonMark parser:

python -m pip install requests markdown-it-py

The following script downloads a Markdown document, resolves relative links, records URI and email autolinks, and returns stable JSON. It intentionally does not claim that an extracted email address is deliverable.

from __future__ import annotations

import json
import sys
from urllib.parse import urljoin, urlparse

import requests
from markdown_it import MarkdownIt


def parse_url_parts(value: str) -> dict[str, str]:
    parsed = urlparse(value)
    return {
        "scheme": parsed.scheme,
        "netloc": parsed.netloc,
        "path": parsed.path,
        "params": parsed.params,
        "query": parsed.query,
        "fragment": parsed.fragment,
    }


def extract_markdown(url: str) -> dict:
    response = requests.get(
        url,
        timeout=(10, 30),
        headers={"User-Agent": "markdown-extractor/1.0"},
    )
    response.raise_for_status()
    content_type = response.headers.get("content-type", "").lower()
    if "text/markdown" not in content_type and not url.lower().endswith((".md", ".markdown")):
        raise ValueError(f"Expected Markdown, received {content_type or 'unknown content type'}")

    markdown = MarkdownIt("commonmark")
    tokens = markdown.parse(response.text)
    links = []
    emails = []

    for token in tokens:
        if token.type != "inline" or not token.children:
            continue
        children = token.children
        for index, child in enumerate(children):
            if child.type == "link_open":
                href = child.attrGet("href")
                if href:
                    links.append({
                        "raw": href,
                        "resolved": urljoin(response.url, href),
                    })
            elif child.type == "autolink_open":
                href = child.attrGet("href") or ""
                if href.lower().startswith("mailto:"):
                    emails.append(href[7:])
            elif child.type == "text" and index > 0:
                # markdown-it-py normally emits email autolinks as an autolink token.
                # This branch preserves visible text only; it does not guess emails with a regex.
                pass

    unique_links = list({item["resolved"]: item for item in links}.values())
    unique_emails = sorted({email for email in emails})
    return {
        "source": response.url,
        "url_parts": parse_url_parts(response.url),
        "links": unique_links,
        "emails": unique_emails,
    }


if __name__ == "__main__":
    target = sys.argv[1]
    print(json.dumps(extract_markdown(target), indent=2, ensure_ascii=False))

Run it with:

python extract_markdown.py https://example.com/guide.md

urljoin resolves a relative destination against the final response URL, which matters when redirects change the document’s base. Keep both the original destination and the resolved value when you need auditability.

3. cURL: fetch first, then inspect

cURL is useful for downloading the response and checking headers. It does not parse CommonMark by itself.

curl -L --fail --show-error --silent \
  -A 'markdown-extractor/1.0' \
  -D headers.txt \
  https://example.com/guide.md \
  -o document.md

cat headers.txt

Pass document.md to a CommonMark parser rather than extracting destinations with one regular expression. Check the final URL after redirects if relative links must be resolved.

4. Node.js: fetch and parse Markdown

For Node.js 18 or newer, install a CommonMark-compatible parser:

npm install markdown-it
import MarkdownIt from "markdown-it";

const input = process.argv[2];
if (!input) throw new Error("Usage: node extract.mjs https://example.com/guide.md");

const response = await fetch(input, {
  redirect: "follow",
  headers: { "user-agent": "markdown-extractor/1.0" },
});
if (!response.ok) throw new Error(`${response.status} ${response.statusText}`);

const type = (response.headers.get("content-type") || "").toLowerCase();
if (!type.includes("text/markdown") && !/\.(md|markdown)(?:$|\?)/i.test(response.url)) {
  throw new Error(`Expected Markdown, received ${type || "unknown content type"}`);
}

const source = await response.text();
const md = new MarkdownIt("commonmark");
const tokens = md.parse(source, {});
const links = [];
const emails = new Set();

for (const token of tokens) {
  if (token.type !== "inline" || !token.children) continue;
  for (const child of token.children) {
    if (child.type === "link_open" && child.attrGet("href")) {
      const raw = child.attrGet("href");
      links.push({ raw, resolved: new URL(raw, response.url).href });
    }
    if (child.type === "autolink_open") {
      const href = child.attrGet("href") || "";
      if (href.toLowerCase().startsWith("mailto:")) emails.add(href.slice(7));
    }
  }
}

console.log(JSON.stringify({
  source: response.url,
  links: [...new Map(links.map(link => [link.resolved, link])).values()],
  emails: [...emails].sort(),
}, null, 2));

5. Markdown forms your parser must support

Form Example Extraction rule
Inline link [Docs](/docs) Read the destination from the link node and resolve it against the final document URL.
Reference link [Docs][api] with [api]: /docs Use the parser’s resolved link definition; do not scan only the first line.
URI autolink <https://example.com> Collect the absolute URI destination.
Email autolink <person@example.com> Collect the mailto: destination as an address-like string.

CommonMark calls autolinks absolute URIs and email addresses inside angle brackets. Its email pattern is non-normative, so successful extraction does not prove that a mailbox exists, accepts mail, or is valid under every provider’s rules.

6. URL components and normalization

from urllib.parse import parse_qs, urlparse, urlunparse

value = "https://example.com:443/a/../guide;v=1?lang=en&lang=fr#intro"
parts = urlparse(value)
print(parts.scheme)      # https
print(parts.netloc)      # example.com:443
print(parts.path)        # /a/../guide
print(parts.params)      # v=1
print(parse_qs(parts.query))
print(parts.fragment)    # intro
print(urlunparse(parts))

Preserve fragments when they identify an in-page target, but do not send them to an HTTP server: browsers keep fragments client-side. Decide whether tracking parameters should be removed only if your application has an explicit policy; changing a URL can change the resource. Lowercasing a hostname is generally safe, while changing path case may not be.

7. Handling HTML pages

If the URL returns HTML, use an HTML parser and inspect <a href> elements, mailto: links, and visible text according to your requirements. Do not treat HTML source as Markdown: an HTML page can contain Markdown-looking text inside code blocks or scripts. A separate HTML implementation might use Python’s standard-library html.parser or a maintained HTML parser package.

8. Edge cases and defensive rules

  • Redirects: resolve links against the final response URL, not necessarily the URL typed by the user.
  • Scheme-relative links: //cdn.example.com/file inherit the response scheme.
  • Fragments: keep them for navigation, but exclude them when comparing network resources if that is your policy.
  • Empty destinations: []() and empty references may mean the current document; decide whether to retain or discard them.
  • Data and JavaScript URLs: classify data:, javascript:, and other non-HTTP schemes before storing or fetching them.
  • Internationalized domains: preserve the original text and use an IDNA-aware URL policy when making network requests.
  • Encoded characters: do not decode percent escapes indiscriminately; decoding can change path meaning.
  • Duplicate links: deduplicate after resolution if you want one record per target, or before resolution if source spelling matters.
  • Email privacy: avoid logging addresses unnecessarily and treat extracted addresses as unverified input.
  • Large or hostile documents: enforce byte, time, redirect, and parser limits. Never execute downloaded JavaScript just to extract Markdown.

9. Troubleshooting

Symptom Cause Fix
No links found The response is HTML, a PDF, or a Markdown extension your parser does not enable. Inspect Content-Type, save the response, and choose the parser for its actual format.
Relative links point to the wrong host Resolution used the requested URL instead of the final URL after redirects. Resolve against response.url (Python) or response.url after Node fetch follows redirects.
Reference links are missing A regex searched only for ](...). Use a CommonMark parser that builds reference definitions.
Email addresses are incomplete Only mailto: links were checked, or the document uses plain text. Define whether plain-text detection is required; if so, use a carefully scoped, documented detector and label results as unverified.
403, 429, or timeout Server policy, rate limiting, or a slow origin. Use a descriptive user agent, obey robots and rate limits, set connect/read timeouts, and retry only transient failures with backoff.
Parser crashes on huge input Unbounded download or deeply nested syntax. Limit response bytes, reject unexpected media types, and process trusted content in a constrained worker.

10. Performance, reliability, and cost

  • Network transfer and server latency usually dominate parsing time; reuse HTTP connections and avoid downloading the same URL repeatedly.
  • Cache by final URL and relevant request headers. Set an expiration policy because Markdown can change.
  • Bound concurrency so a batch job does not trigger rate limits. Retry 408, 429, and transient 5xx responses with exponential backoff and a maximum attempt count.
  • Record status code, final URL, content type, byte count, and parser version for reproducibility.
  • URL parsing is not URL validation. Apply an allowlist, block private network ranges where appropriate, and prevent server-side request forgery if users supply URLs.
  • Parsing locally has no API usage charge, but bandwidth, compute, storage, and any third-party fetch service still have costs.

11. Or skip the browser setup

If you need a clean image or PDF of the source page before processing it, ScreenshotNeo provides a website screenshot API. One GET request returns PNG, JPEG, WebP, or PDF; the API can capture a page or element and supports custom CSS and JavaScript, waits, headers, cookies, user agents, and other capture controls. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/guide.md -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/guide.md"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/guide.md' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed, and response headers identify the page verdict and billing status. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots.

Create your free ScreenshotNeo account.

12. FAQ

No. The URL identifies a resource; you need the resource representation or an API that exposes its links.

Does finding an email prove it is real?

No. It proves only that the document contained an address-like Markdown autolink or destination.

Should I use one regular expression for Markdown?

No. Inline links, reference links, escapes, nesting, and autolinks have different grammar. Use a CommonMark parser when destinations must be correct.

The raw value preserves what the author wrote; the resolved value is useful for crawling, deduplication, and downstream requests.