How to Scrape Email Addresses from a Website
Learn where websites expose email addresses and how to collect them responsibly, with a runnable Python example and practical privacy safeguards.

A website can expose an email address in visible text or in a mailto: link. A small Python script can find addresses on a page you are permitted to process, but first decide why you need them and whether collection and intended use are appropriate. A public address is not blanket permission to collect, retain, or message its owner.
This guide shows a limited, single-page example, explains how to inspect public mail links, and gives a checklist for site rules, privacy, validation, and retention. It does not cover evading bot checks, bypassing access controls, or building a bulk contact-harvesting system.
1. Check purpose, permission, and data handling first
Start by asking: Where does a website expose an email address, and what should I check before collecting it? The technical answer is often visible page text or a public mailto: URI. The governance answer depends on your purpose, the site, the people identified by the addresses, your jurisdiction, and what you do with the data afterward.
- Define a narrow purpose. For example, extracting a published support contact from your own site for a migration differs from gathering personal addresses for unsolicited outreach. Collect only the fields needed for that purpose.
- Review the site’s terms and crawler guidance. Read the site’s applicable terms and its
/robots.txtrules. RFC 9309 describes the Robots Exclusion Protocol as crawler guidance and states: “These rules are not a form of access authorization.” A permissive file does not establish permission under terms, privacy law, or other rules. RFC 9309. - Respect explicit objections and access controls. Do not work around a CAPTCHA, login requirement, rate limit, or other access restriction. CNIL notes in its legitimate-interest analysis that reasonable expectations may not be met where a site explicitly opposes scraping through technical measures such as robots.txt or CAPTCHA; that is guidance in its context, not a universal legal rule. CNIL guidance.
- Identify the relevant jurisdiction and role. In the EU, an email address can be personal data, and the EDPB says GDPR applies to scraping involving personal-data processing. Public availability alone does not settle whether processing is lawful or which lawful basis applies. European Commission: personal data; EDPB scraping guidance announcement.
- Set controls before collecting. Record the source and collection time, restrict access, validate only as needed, choose a retention period, and provide a deletion path. EDPB materials highlight purpose limitation, transparency, reliable sources, timestamps, validation, and data minimisation in the context discussed.
For US readers, the FTC’s CAN-SPAM compliance guide identifies email harvesting and dictionary attacks among aggravated violations that may lead to criminal penalties, and notes civil penalties for violations. That does not mean every act of viewing or collecting a publicly displayed address invariably violates CAN-SPAM; conduct and applicable law matter. FTC CAN-SPAM guide.
2. Know where email addresses appear
Common public exposure points include visible contact text and links such as <a href="mailto:help@example.com">Email support</a>. RFC 6068 specifically warns that “‘mailto’ URIs on public Web pages expose mail addresses for harvesting.” A mailto URI can also contain addresses in fields beyond the visible recipient, so inspect what you extract and avoid collecting unnecessary fields. RFC 6068.
Some pages render contact details after JavaScript runs, or publish them in structured data. There is no universal extraction recipe that reliably finds every address on every site. The example below intentionally handles a narrow, permitted case: it fetches one page, extracts addresses from visible text and mailto links, deduplicates them, and prints results. It does not crawl links or defeat obfuscation or access controls.
3. Extract addresses from one permitted page with Python
Use Python 3. Install the two dependencies with python -m pip install requests beautifulsoup4. Save the following as extract_emails.py. Replace the example URL with a page you are authorized to process.

import re
from urllib.parse import unquote, urlparse
import requests
from bs4 import BeautifulSoup
PAGE_URL = "https://example.com/contact"
TIMEOUT_SECONDS = 15
MAX_RESPONSE_BYTES = 2_000_000
# A practical matching pattern for ordinary email-shaped strings, not a
# complete implementation of every address form permitted by email standards.
EMAIL_RE = re.compile(
r"(?i)(?
Implementation notes: the expression is a practical filter, not an email-standard validator. It excludes many unusual but valid address forms and may still match strings that are not usable mailboxes. The script does not verify mailbox existence or permission to contact anyone. The response-size cap limits accidental processing of unexpectedly large pages; choose a limit appropriate to your permitted use.
What the script does and does not do
- It makes one ordinary HTTP GET request, checks for an HTML content type, and applies a timeout.
- It extracts address-shaped text and comma-separated recipients from public mailto links.
- It deduplicates results in memory and normalizes domain casing.
- It does not follow links, execute JavaScript, scrape hidden records, defeat a CAPTCHA, or send email.
- It prints addresses to standard output. For personal data, choose a protected destination and avoid retaining output longer than your purpose requires.
4. Adapt the request carefully
For a permitted internal workflow, you may want to save a provenance record alongside each result. Keep the original page URL, retrieval timestamp, and the narrow reason for collection. Do not add broad crawling simply because the one-page script works.
cURL: inspect the page response
curl --fail --location --max-time 15 \
--header 'User-Agent: ContactPageResearch/1.0' \
--output contact.html \
'https://example.com/contact'
This downloads the HTML response to a local file; it does not extract addresses. Review the response and page rules before processing it. Avoid saving data to shared or public locations.
Node.js: fetch one HTML page
const url = 'https://example.com/contact';
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 15_000);
try {
const response = await fetch(url, {
headers: { 'User-Agent': 'ContactPageResearch/1.0' },
signal: controller.signal,
});
if (!response.ok) {
throw new Error(`HTTP ${response.status}`);
}
const type = response.headers.get('content-type') || '';
if (!type.toLowerCase().includes('html')) {
throw new Error(`Expected HTML, received ${type || 'unknown content type'}`);
}
const html = await response.text();
console.log(`Fetched ${html.length} characters from ${url}`);
// Parse only for a defined, permitted purpose. For DOM parsing in Node,
// install and use a parser such as cheerio; do not treat regex as full HTML parsing.
} finally {
clearTimeout(timer);
}
Node's built-in fetch does not parse HTML. If you add a parser, pin and maintain the dependency, enforce a response-size limit, and extract only the fields required. The Python example shows a complete extraction path for a single page.
5. Handle edge cases without broadening the collection
| Case | What may happen | Careful response |
|---|---|---|
| Visible text differs from link target | A link label may say “Contact us,” while the href contains an address. | Inspect the public href and retain only the address needed for the stated purpose. |
| Mailto query fields | A URI can include cc, bcc, subject, or body fields in addition to recipients. | Do not ingest query fields unless your defined purpose requires them; treat them as potentially sensitive content. |
| JavaScript-rendered content | The initial HTML may omit what a browser later displays. | Prefer an official contact source or ask the site owner. Do not treat rendering differences as permission to evade controls. |
| Obfuscated address | A page may spell out “name [at] example [dot] com” or use a script. | Do not attempt to defeat an intentional protection. Use a published contact route or obtain authorization. |
| Internationalized or unusual address | A simple pattern can miss valid forms or misclassify text. | Do not assume a regex is authoritative. Validate only when necessary and avoid probing third-party mailboxes. |
| Duplicate address on a page | The same address may occur in text and several links. | Deduplicate for the narrow task, while preserving provenance if the source matters. |

6. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| HTTP 403 or CAPTCHA page | The site denied automated access or requires an interactive check. | Stop automated requests. Use the site's published contact method or ask for permission; do not bypass the challenge. |
| HTTP 404 | The path changed or the page is no longer public. | Check the URL manually and use a current official page if available. |
| Timeout | The host is slow, unreachable, or the request is stalled. | Retry sparingly after checking the site status and rules. Keep a finite timeout; do not use rapid retry loops. |
| “Expected HTML” error | The URL returned a PDF, image, redirect destination, or other content type. | Confirm that the selected public page is HTML. Do not parse unrelated content without a defined need. |
| No addresses found | The address may be absent from raw HTML, rendered later, encoded differently, or not present at all. | Inspect the page as a visitor and check official contact information. Treat a failed match as unknown, not proof of absence. |
| False positive or malformed result | A broad pattern matched text that resembles an address, or address syntax is unusual. | Review results before use; do not automatically contact them or enrich them against unrelated sources. |
| Unexpectedly large response | The server returned a large document or an unexpected resource. | Keep the size cap, verify the URL and content type, and stop if the response is outside the planned scope. |
7. Performance, reliability, and cost
A single-page request has little setup cost, but network latency and the site's response time dominate. Use a bounded timeout, check status and content type, cap response size, and avoid repeated requests. For a small authorized task, fetch only the needed page and cache your own result only as long as the purpose allows. For larger legitimate migrations, ask the site owner for an export or API; that is usually more reliable than inferring data from rendered pages.
HTML changes can break selectors and extraction assumptions. Log failures without logging unnecessary personal data, retain a timestamp and source for records you keep, and review extracted results before acting on them. Do not treat a syntactically plausible address as deliverable or as consent to receive messages.
Direct fetching has no software API charge in this example, though normal hosting, network, and engineering costs still apply. The material costs may instead be data protection, security, review, and remediation if you collect more than needed or use it in a way people do not reasonably expect. Requirements vary by jurisdiction and context; the sources here do not establish a worldwide legal answer.
8. Or skip the browser setup
If your task is to capture a page image for documentation or review, ScreenshotNeo can return a screenshot with one request. It is a website screenshot API and MCP server for developers, made by Yorker Media. See the ScreenshotNeo site and API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/contact -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/contact"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/contact' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Screenshot capture is for visual records: it does not extract email addresses or establish permission to collect or use them.
Sign up free for 1,000 screenshots a month, no card required.
9. A practical responsible-collection checklist
- The purpose is specific, documented, and limited.
- Site terms and robots.txt have been reviewed; no access control or explicit objection is being bypassed.
- The relevant jurisdiction and personal-data implications have been assessed.
- Only necessary addresses and provenance fields are retained.
- Results are protected, quality-reviewed, and deleted when no longer needed.
- Any downstream contact has a separately assessed lawful basis and applicable messaging compliance.
10. FAQ
Does a public email address mean I can email it?
No. Publication makes an address observable; it does not itself answer whether a message is lawful, expected, or wanted. Assess the intended use and applicable rules separately.
Can robots.txt give permission to scrape?
No. RFC 9309 explicitly says the protocol is not access authorization. It is one signal to review alongside terms, technical restrictions, privacy obligations, and your purpose.
Is a regex enough to validate an email address?
No. A regex can find common address-shaped strings but cannot establish that an address is valid, active, controlled by a person, or appropriate to use.
Does this script find every address on a site?
No. It processes one returned HTML page and common visible text and mailto patterns. It deliberately does not crawl an entire domain or bypass rendering and access restrictions.
This is practical technical information, not legal advice. The cited guidance is jurisdiction- and context-specific; seek qualified advice for a consequential collection or outreach program.


