ScreenshotNeo

BlogHow-to

How to Extract URLs from Text

Extract URLs reliably with a candidate regex, URL parser, punctuation cleanup, validation rules, and runnable Python, JavaScript, and cURL examples.

By the ScreenshotNeo team1 October 20265 min read

Short answer: find URL-shaped candidates, trim wrappers and sentence punctuation, then parse and validate each candidate with a URL API and an explicit security policy. A regex is a locator, not a validator. For HTML or Markdown, parse link nodes instead of searching rendered text.

This workflow follows RFC 3986, which defines URI components and warns that prose punctuation and delimiters can be mistaken for URI content.

1. Choose an extraction strategy

Input Best method Reason
Plain text or logs Candidate regex, then parser No link structure exists.
HTML Parse a[href] Avoids scripts, comments, and visible text that is not a link.
Markdown Parse link and autolink nodes Handles escaping and destinations.
  1. Locate candidates beginning with allowed schemes such as https://, http://, or ftp://.
  2. Trim wrappers and punctuation only when outside the URL.
  3. Parse with a standards-aware URL API.
  4. Enforce scheme, host, port, and credential policy.
  5. Deduplicate a normalized key while retaining original text.

2. Python implementation

import re
from urllib.parse import urlsplit, urldefrag

CANDIDATE_RE = re.compile(r'''(?ix)\b(?:https?|ftp)://[^\s<>"']+''')
ALLOWED = {"http", "https", "ftp"}

def clean(raw):
    value = raw.strip()
    if len(value) >= 2 and value[0] == "<" and value[-1] == ">":
        value = value[1:-1]
    if len(value) >= 2 and value[0] in "'\"" and value[-1] == value[0]:
        value = value[1:-1]
    while value and value[-1] in ".,;:!?]}'}\"":
        value = value[:-1]
    while value.endswith(")") and value.count("(") < value.count(")"):
        value = value[:-1]
    return value

def extract_urls(text, drop_fragments=True):
    result, seen = [], set()
    for raw in CANDIDATE_RE.findall(text):
        candidate = clean(raw)
        try: parts = urlsplit(candidate)
        except ValueError: continue
        if parts.scheme.lower() not in ALLOWED or not parts.netloc: continue
        try: parts.port
        except ValueError: continue
        value = urldefrag(candidate)[0] if drop_fragments else candidate
        if value not in seen:
            seen.add(value); result.append(value)
    return result

print(*extract_urls('Read <https://example.com/docs?q=1>, then https://example.com/a_(demo).'), sep='\n')

urlsplit separates scheme, authority, path, query, and fragment; urldefrag removes a fragment. See the Python urllib.parse docs.

Resolve relatives deliberately

from urllib.parse import urljoin
print(urljoin("https://docs.example.com/guide/", "../api"))

Resolve /docs/page only against a trusted base; otherwise keep it relative.

3. JavaScript and Node.js

function extractUrls(text, baseUrl) {
  const rough = text.match(/\b(?:https?|ftp):\/\/[^\s<>"']+/gi) ?? [];
  const seen = new Set(), out = [];
  for (let value of rough) {
    if (value.startsWith("<") && value.endsWith(">")) value = value.slice(1,-1);
    value = value.replace(/[.,;:!?]}']+$/, "");
    while (value.endsWith(")") && (value.match(/\(/g)||[]).length < (value.match(/\)/g)||[]).length) value=value.slice(0,-1);
    try {
      const u = new URL(value, baseUrl);
      if (!["http:","https:","ftp:"].includes(u.protocol)) continue;
      if (!seen.has(u.href)) { seen.add(u.href); out.push(u.href); }
    } catch {}
  }
  return out;
}
console.log(extractUrls('See https://example.com/a_(demo), and <https://example.org>.'));

The MDN URL API also documents URL.canParse() for a non-throwing check.

4. cURL pipeline

curl -s https://example.com/page.txt | grep -Eo 'https?://[^[:space:]<>"'"'']+' | sed -E 's/[.,;:!?)}]+$//' | sort -u

This is useful for exploration. Use a parser for production validation.

5. Punctuation, wrappers, and line breaks

  • Trim sentence periods and commas at the end.
  • Trim ], }, or ) only when unmatched.
  • Keep balanced parentheses because paths may contain them.
  • Support <https://example.com> and quoted URLs when your input uses those wrappers.
  • Do not silently join lines unless your source format defines URL continuation.

Keep the original substring for display. Do not lowercase paths or decode percent escapes blindly; URI semantics can be scheme-specific.

6. Validation and security

  • Allowlist https (and http only when needed); allow ftp deliberately.
  • Require a host and validate ports.
  • Reject javascript:, data:, and unexpected schemes before navigation or fetching.
  • Treat userinfo, unusual IP forms, redirects, and credentials as sensitive. The rfc3986 validator supports required schemes/hosts and forbidding passwords in userinfo.
  • For server-side fetches, add SSRF controls, timeouts, redirect limits, and response-size limits.

Protocol-relative values such as //cdn.example.com/a.js need a trusted scheme. Never invent a base for an untrusted relative reference.

7. HTML and Markdown

Parse HTML and read href attributes from a elements, resolving relatives against the trusted document URL. Parse Markdown link and autolink nodes. Regex over raw HTML can return script text and comments.

8. Deduplication

  1. Store original_text and parsed URL separately.
  2. Decide whether fragments identify distinct resources.
  3. Deduplicate parsed values first; add scheme-specific canonicalization only when understood.
  4. Use the URL library for internationalized domains and percent encoding.

9. Troubleshooting

Symptom Cause Fix
Final period included Regex stopped only at whitespace Trim punctuation after candidate matching.
Closing parenthesis missing Unconditional trimming Count balanced parentheses.
/guide rejected Relative reference Resolve with a trusted base.
Dangerous value accepted Syntax mistaken for policy Allowlist schemes and hosts.
Unicode host differs Ad hoc normalization Use the platform URL API.
Duplicates remain Raw text used as key Deduplicate parsed values.

10. Performance, reliability, and cost

  • Scan once, parse only matches, and cap input and candidate lengths.
  • Record rejected candidates and reasons for debugging.
  • Extraction performs no network request; fetching requires timeouts, retries, SSRF defenses, and size limits.
  • Local parsing has no service charge; account separately for network fetches, storage, and rendering.

11. Or skip the browser setup

If extracted URLs need screenshots, ScreenshotNeo is a website screenshot API and MCP server. One GET returns PNG, JPEG, WebP, or PDF. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; X-Page-Verdict and X-Billed identify the result. MCP tools include take_screenshot, get_page_info, and capture_pdf.

See the ScreenshotNeo docs for options.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

12. FAQ

Can one regex validate every URL?

No. Combine candidate matching, parsing, and policy checks.

Should fragments be kept?

Keep them for section identity; remove them for resource-level deduplication.

Parse HTML and inspect href attributes, resolving relatives against the page URL.

Are email addresses URLs?

Not under this HTTP/HTTPS/FTP policy; handle mailto: separately.