How to Scrape Emails from Any Website
Learn a permission-aware way to find public email addresses, validate results, and avoid privacy, site-policy, and marketing-law mistakes.
There is no universal permission to scrape email addresses from any website. An address being visible does not decide whether you may collect it, store it, combine it with other data, or use it for marketing. Before writing code, identify the purpose, the people represented by the addresses, the jurisdictions involved, and what the site permits.
This guide shows a narrow, auditable method for extracting addresses from a page you are allowed to access. It also explains classification, minimisation, validation, rate limits, failure modes, and the rules that apply to later outreach.
1. Decide whether collection is appropriate
Separate these decisions:
- Viewing: can your browser access the page?
- Collecting: may you copy addresses into a file or database?
- Using: may you combine them with other data or contact the people?
- Marketing: do the rules for commercial messages permit your planned email?
A named employee’s business address can be personal data under the European Commission’s examples, while a generic mailbox such as info@example.com may not be personal data in that context. The distinction depends on whether the address relates to an identifiable living person. See the European Commission’s company-data guidance.
Collection itself is processing. The Commission describes processing as including collection, recording, organisation, storage, retrieval, use, and disclosure. A public page therefore does not remove the need for a lawful basis, fairness, and transparency: GDPR processing principles and legal grounds for processing.
| Question | What to record before collecting |
|---|---|
| Purpose | Research, support, recruiting, account management, or a defined outreach campaign |
| Address type | Generic role mailbox, named employee, or ambiguous |
| Jurisdiction | Where you operate, where people are located, and where processing occurs |
| Site rules | Terms, access controls, robots.txt, rate limits, and copyright or database-rights notices |
| Retention | Deletion date, access controls, source URL, and collection timestamp |
2. Use a permission-aware collection workflow
- Read the site’s terms and relevant notices.
- Use a single page or a documented list of pages; do not crawl an entire domain by default.
- Identify yourself with a descriptive user agent and obey published rate limits.
- Fetch only the HTML needed for the stated purpose.
- Extract candidates, preserve the source URL and timestamp, and remove duplicates.
- Review candidates manually before any decision that affects a person.
- Delete records that are unnecessary, stale, or outside the stated purpose.
Do not bypass CAPTCHAs, login controls, paywalls, robots protections, or technical restrictions. If a page requires authentication, obtain the site’s permission and use an approved API or export instead.
3. Python: extract visible and mailto addresses from one page
The following script processes one URL supplied on the command line. It uses a timeout, a descriptive user agent, a maximum response size, and a simple email pattern. It does not crawl links or evade controls.
#!/usr/bin/env python3
import argparse
import re
import sys
from urllib.parse import urlparse
import requests
from bs4 import BeautifulSoup
EMAIL_RE = re.compile(
r"\\b[A-Za-z0-9.!#$%&'*+/=?^_`{|}~-]+@"
r"[A-Za-z0-9](?:[A-Za-z0-9-]{0,61}[A-Za-z0-9])?"
r"(?:\\.[A-Za-z0-9](?:[A-Za-z0-9-]{0,61}[A-Za-z0-9])?)+\\b"
)
def scrape(url: str) -> list[str]:
parsed = urlparse(url)
if parsed.scheme not in {"http", "https"} or not parsed.netloc:
raise ValueError("URL must start with http:// or https://")
headers = {"User-Agent": "EmailResearchBot/1.0 (contact: you@example.com)"}
response = requests.get(url, headers=headers, timeout=20, stream=True)
response.raise_for_status()
content_type = response.headers.get("content-type", "")
if "text/html" not in content_type:
raise ValueError(f"Expected HTML, received {content_type or 'unknown type'}")
content = response.raw.read(2_000_000 + 1)
if len(content) > 2_000_000:
raise ValueError("Response exceeds the 2 MB safety limit")
soup = BeautifulSoup(content, "html.parser")
candidates = set()
for link in soup.select('a[href^="mailto:"]'):
address = link.get("href", "")[7:].split("?", 1)[0].strip()
candidates.update(EMAIL_RE.findall(address))
visible_text = soup.get_text(" ", strip=True)
candidates.update(EMAIL_RE.findall(visible_text))
return sorted({address.lower() for address in candidates})
if __name__ == "__main__":
parser = argparse.ArgumentParser(description="Extract public email candidates from one HTML page")
parser.add_argument("url")
args = parser.parse_args()
try:
for address in scrape(args.url):
print(address)
except (requests.RequestException, ValueError) as exc:
print(f"error: {exc}", file=sys.stderr)
sys.exit(1)
Install dependencies with python -m pip install requests beautifulsoup4, then run python scrape_emails.py https://example.com/contact. Save the output together with the URL and date; an address without provenance is difficult to verify or delete later.
What this script deliberately does not do
- It does not discover or crawl every page.
- It does not decode obfuscations such as
name [at] example [dot] com; manually review those only when the site permits it. - It does not execute JavaScript. A client-rendered page may require an authorized browser workflow or an official API.
- It does not verify that a mailbox exists or that a person wants contact.
4. cURL and Node.js alternatives
For a quick inspection, cURL retrieves the HTML. It is not a parser, so pipe the response into a reviewed tool rather than treating every match as a valid address.
curl --fail --location --max-time 20 \\
-H 'User-Agent: EmailResearchBot/1.0 (contact: you@example.com)' \\
https://example.com/contact -o page.html
Node.js 18+ can fetch one page and extract candidates with a regular expression. For production parsing, use an HTML parser and enforce the same size, timeout, and rate limits as the Python example.
const url = process.argv[2];
if (!url) throw new Error('Usage: node scrape-emails.mjs https://example.com/contact');
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 20_000);
try {
const res = await fetch(url, {
signal: controller.signal,
headers: { 'user-agent': 'EmailResearchBot/1.0 (contact: you@example.com)' }
});
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const type = res.headers.get('content-type') || '';
if (!type.includes('text/html')) throw new Error(`Expected HTML, received ${type}`);
const html = await res.text();
if (Buffer.byteLength(html, 'utf8') > 2_000_000) throw new Error('Response exceeds 2 MB');
const pattern = /\\b[A-Z0-9.!#$%&'*+/=?^_`{|}~-]+@[A-Z0-9](?:[A-Z0-9-]{0,61}[A-Z0-9])?(?:\\.[A-Z0-9](?:[A-Z0-9-]{0,61}[A-Z0-9])?)+\\b/gi;
console.log([...new Set(html.match(pattern) || [])].map(x => x.toLowerCase()).sort().join('\\n'));
} finally {
clearTimeout(timer);
}
5. Improve accuracy without expanding scope
Extract context, not just strings
Store the page title, nearby heading, source URL, collection time, and whether the value came from a mailto: link or visible text. Context helps distinguish a support mailbox from an employee address and supports later correction requests.
Deduplicate safely
Lowercase domains, trim whitespace, and remove duplicates only after preserving each source. Do not silently merge aliases; sales@example.com and support@example.com may have different purposes.
Validate syntax and domain separately
Syntax checks catch malformed text. DNS or mailbox checks can create additional processing and network traffic, may be inaccurate, and should have a documented purpose. Never send a test message without permission.
Handle dynamic pages
If the initial HTML has no address, inspect whether the site provides a public API, embedded JSON, or a contact form. Browser automation should be the last step, rate-limited, and authorized. Do not use it to defeat bot checks or access controls.
6. Privacy, site-policy, and marketing checkpoints
CNIL explains that web scraping is not automatically incompatible with GDPR, but it requires a valid legal basis and measures that protect people’s rights. Its guidance also discusses site terms, database rights, and excluding sites that oppose scraping in the specific context addressed by the guidance: CNIL scraping guidance.
If data came from another source, transparency duties can apply. The European Commission describes information obligations for indirectly obtained data, including timing rules and exceptions: information for people whose data is collected.
Canada takes a materially different approach. The Office of the Privacy Commissioner says that, with very limited exceptions, PIPEDA prohibits address harvesting, including collecting electronic addresses with computer programs such as website scraping: Canadian e-marketing guidance (PDF).
Sending is a separate decision from collecting. In the United States, the FTC’s CAN-SPAM guide covers commercial messages, including business-to-business email. Requirements include accurate headers, non-deceptive subjects, ad identification, a valid postal address, an opt-out method, and prompt honoring of opt-outs. Other countries may impose consent or prior-relationship requirements.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| 403 or 429 response | Access denied or rate limit | Stop, read the site’s rules, slow down, and request permission or use an official API. |
| 200 response but no emails | JavaScript-rendered content or an image/PDF | Check content type and look for an authorized API or contact page; do not bypass controls. |
| Many false positives | Email-like text in scripts, examples, or documentation | Parse visible text and mailto links, exclude code blocks, and review context. |
| Obfuscated addresses | Human-readable anti-harvesting formatting | Respect the signal; do not automatically decode it without a clear permission basis. |
| Unicode or encoding errors | Missing or incorrect charset | Use the response’s declared encoding and retain the original source for review. |
| Repeated addresses across pages | Shared footer or navigation | Track source URLs and deduplicate only for the defined purpose. |
| Mailbox bounces | Stale, role-based, or invalid address | Mark the result unverified, suppress it from outreach, and delete it when no longer needed. |
8. Performance, reliability, and cost
- Performance: one-page requests with a 20-second timeout and response-size cap are easier to control than an unrestricted crawler.
- Reliability: record status code, content type, URL, timestamp, and parser version. Retry only transient failures, with exponential backoff, and never retry a refusal indefinitely.
- Cost: ordinary HTTP fetching uses your network and compute. Browser rendering, proxy services, and repeated validation add cost and operational risk; use them only when justified and authorized.
- Data quality: a syntactically valid address is not proof of ownership, consent, or deliverability.
- Security: treat downloaded HTML as untrusted input, avoid logging secrets, and protect any resulting dataset as sensitive business information.
9. Or skip the browser setup
If your task is to capture the contact page for review or documentation, ScreenshotNeo returns a screenshot or PDF from one GET request. Its consent step accepts cookie banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and responses include X-Page-Verdict and X-Billed headers.
See the ScreenshotNeo API documentation. Example:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. Every plan includes the features; 1,000 screenshots per month are free with no card, and paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
10. FAQ
Is scraping a public email address always legal?
No. Public visibility does not establish permission, a lawful basis, or permission for marketing. The answer depends on the address, purpose, jurisdiction, and site restrictions.
Does a company email avoid privacy rules?
No. A named employee’s business email can identify a person and may be personal data. A generic role mailbox is treated differently in the European Commission’s examples.
Can I email every address I find?
No. Collection and sending are separate activities. Apply the marketing rules where you and recipients are located, including opt-out and transparency requirements.
Should I use robots.txt as a complete legal test?
No. It is one access signal. Also review terms, technical restrictions, rights notices, and applicable law.
What is the safest way to build a contact list?
Use a permission-based form that states the purpose, records the applicable permission, and keeps a clear audit trail. This reduces uncertainty but does not replace local legal review.


