How to Extract URLs from Text
Extract URLs reliably with a candidate regex, URL parser, punctuation cleanup, validation rules, and runnable Python, JavaScript, and cURL examples.
Short answer: find URL-shaped candidates, trim wrappers and sentence punctuation, then parse and validate each candidate with a URL API and an explicit security policy. A regex is a locator, not a validator. For HTML or Markdown, parse link nodes instead of searching rendered text.
This workflow follows RFC 3986, which defines URI components and warns that prose punctuation and delimiters can be mistaken for URI content.
1. Choose an extraction strategy
| Input | Best method | Reason |
|---|---|---|
| Plain text or logs | Candidate regex, then parser | No link structure exists. |
| HTML | Parse a[href] |
Avoids scripts, comments, and visible text that is not a link. |
| Markdown | Parse link and autolink nodes | Handles escaping and destinations. |
- Locate candidates beginning with allowed schemes such as
https://,http://, orftp://. - Trim wrappers and punctuation only when outside the URL.
- Parse with a standards-aware URL API.
- Enforce scheme, host, port, and credential policy.
- Deduplicate a normalized key while retaining original text.
2. Python implementation
import re
from urllib.parse import urlsplit, urldefrag
CANDIDATE_RE = re.compile(r'''(?ix)\b(?:https?|ftp)://[^\s<>"']+''')
ALLOWED = {"http", "https", "ftp"}
def clean(raw):
value = raw.strip()
if len(value) >= 2 and value[0] == "<" and value[-1] == ">":
value = value[1:-1]
if len(value) >= 2 and value[0] in "'\"" and value[-1] == value[0]:
value = value[1:-1]
while value and value[-1] in ".,;:!?]}'}\"":
value = value[:-1]
while value.endswith(")") and value.count("(") < value.count(")"):
value = value[:-1]
return value
def extract_urls(text, drop_fragments=True):
result, seen = [], set()
for raw in CANDIDATE_RE.findall(text):
candidate = clean(raw)
try: parts = urlsplit(candidate)
except ValueError: continue
if parts.scheme.lower() not in ALLOWED or not parts.netloc: continue
try: parts.port
except ValueError: continue
value = urldefrag(candidate)[0] if drop_fragments else candidate
if value not in seen:
seen.add(value); result.append(value)
return result
print(*extract_urls('Read <https://example.com/docs?q=1>, then https://example.com/a_(demo).'), sep='\n')
urlsplit separates scheme, authority, path, query, and fragment; urldefrag removes a fragment. See the Python urllib.parse docs.
Resolve relatives deliberately
from urllib.parse import urljoin
print(urljoin("https://docs.example.com/guide/", "../api"))
Resolve /docs/page only against a trusted base; otherwise keep it relative.
3. JavaScript and Node.js
function extractUrls(text, baseUrl) {
const rough = text.match(/\b(?:https?|ftp):\/\/[^\s<>"']+/gi) ?? [];
const seen = new Set(), out = [];
for (let value of rough) {
if (value.startsWith("<") && value.endsWith(">")) value = value.slice(1,-1);
value = value.replace(/[.,;:!?]}']+$/, "");
while (value.endsWith(")") && (value.match(/\(/g)||[]).length < (value.match(/\)/g)||[]).length) value=value.slice(0,-1);
try {
const u = new URL(value, baseUrl);
if (!["http:","https:","ftp:"].includes(u.protocol)) continue;
if (!seen.has(u.href)) { seen.add(u.href); out.push(u.href); }
} catch {}
}
return out;
}
console.log(extractUrls('See https://example.com/a_(demo), and <https://example.org>.'));
The MDN URL API also documents URL.canParse() for a non-throwing check.
4. cURL pipeline
curl -s https://example.com/page.txt | grep -Eo 'https?://[^[:space:]<>"'"'']+' | sed -E 's/[.,;:!?)}]+$//' | sort -u
This is useful for exploration. Use a parser for production validation.
5. Punctuation, wrappers, and line breaks
- Trim sentence periods and commas at the end.
- Trim
],}, or)only when unmatched. - Keep balanced parentheses because paths may contain them.
- Support
<https://example.com>and quoted URLs when your input uses those wrappers. - Do not silently join lines unless your source format defines URL continuation.
Keep the original substring for display. Do not lowercase paths or decode percent escapes blindly; URI semantics can be scheme-specific.
6. Validation and security
- Allowlist
https(andhttponly when needed); allowftpdeliberately. - Require a host and validate ports.
- Reject
javascript:,data:, and unexpected schemes before navigation or fetching. - Treat userinfo, unusual IP forms, redirects, and credentials as sensitive. The rfc3986 validator supports required schemes/hosts and forbidding passwords in userinfo.
- For server-side fetches, add SSRF controls, timeouts, redirect limits, and response-size limits.
Protocol-relative values such as //cdn.example.com/a.js need a trusted scheme. Never invent a base for an untrusted relative reference.
7. HTML and Markdown
Parse HTML and read href attributes from a elements, resolving relatives against the trusted document URL. Parse Markdown link and autolink nodes. Regex over raw HTML can return script text and comments.
8. Deduplication
- Store
original_textand parsed URL separately. - Decide whether fragments identify distinct resources.
- Deduplicate parsed values first; add scheme-specific canonicalization only when understood.
- Use the URL library for internationalized domains and percent encoding.
9. Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| Final period included | Regex stopped only at whitespace | Trim punctuation after candidate matching. |
| Closing parenthesis missing | Unconditional trimming | Count balanced parentheses. |
/guide rejected |
Relative reference | Resolve with a trusted base. |
| Dangerous value accepted | Syntax mistaken for policy | Allowlist schemes and hosts. |
| Unicode host differs | Ad hoc normalization | Use the platform URL API. |
| Duplicates remain | Raw text used as key | Deduplicate parsed values. |
10. Performance, reliability, and cost
- Scan once, parse only matches, and cap input and candidate lengths.
- Record rejected candidates and reasons for debugging.
- Extraction performs no network request; fetching requires timeouts, retries, SSRF defenses, and size limits.
- Local parsing has no service charge; account separately for network fetches, storage, and rendering.
11. Or skip the browser setup
If extracted URLs need screenshots, ScreenshotNeo is a website screenshot API and MCP server. One GET returns PNG, JPEG, WebP, or PDF. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; X-Page-Verdict and X-Billed identify the result. MCP tools include take_screenshot, get_page_info, and capture_pdf.
See the ScreenshotNeo docs for options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
12. FAQ
Can one regex validate every URL?
No. Combine candidate matching, parsing, and policy checks.
Should fragments be kept?
Keep them for section identity; remove them for resource-level deduplication.
How do I extract links from a web page?
Parse HTML and inspect href attributes, resolving relatives against the page URL.
Are email addresses URLs?
Not under this HTTP/HTTPS/FTP policy; handle mailto: separately.


