ScreenshotNeo

BlogHow-to

How to Build an Email Database from Public Web Data

Build a useful, auditable email database from public web data while documenting provenance, consent, objections, and campaign limits.

By the ScreenshotNeo team29 September 202610 min read

How to Build an Email Database from Public Web Data

Direct answer: Build a narrow, purpose-defined directory of relevant contacts, not a list of every address a crawler can find. For each record, save the address, source URL, date found, publication context, role or business relevance, jurisdiction, restrictions, legal assessment, and suppression status. Public visibility does not automatically permit marketing. Collection and sending are separate decisions, and both must comply with the rules that apply to the recipient, sender, channel, and country.

This guide gives you an auditable workflow, a conservative Python collector for pages you are allowed to access, validation and maintenance steps, and a route that avoids browser automation when you need screenshots of source pages.

1. Define the database before collecting anything

Write a short specification before opening a crawler. Record:

  • Purpose: for example, contacting procurement leaders about a relevant software product.
  • Audience: job functions, organization types, company size, and countries.
  • Channel: one-to-one research, commercial email, or another channel.
  • Retention: how long a record remains useful and when it is rechecked.
  • Exclusions: personal addresses, unrelated roles, minors, restricted pages, and addresses with a no-contact instruction.

Purpose limitation and data minimisation are core GDPR principles. The European Commission explains that personal data should be collected for specified purposes, limited to what is necessary, accurate, and processed lawfully and transparently (European Commission GDPR guidance). A named employee’s business address can still be personal data, even when displayed on an employer’s site. The UK ICO likewise says public availability does not remove data-protection duties (ICO business-to-business marketing guidance).

2. Choose sources with useful context

Prefer official company contact pages, staff directories, investor-relations pages, professional associations, government registries, and press pages that publish contact details for a stated professional purpose. Context helps you decide whether an address is relevant and whether a later message would be expected.

Do not treat every visible string matching an email pattern as a qualified lead. A personal address copied into a public document, a support address intended only for existing customers, and a press address with an explicit “media inquiries only” notice have different purposes.

The ICO lists company websites, Companies House, social media, and press articles as examples of public sources, while warning that personal data in those sources remains regulated. A professional-network profile may be used in a person’s professional capacity, but outreach there can still be subject to privacy and electronic-marketing rules.

3. Collect only the fields you can explain

A practical record schema is:

Field Why keep it
organization Identifies the business context.
displayed_name Preserves the name exactly as published.
role Shows why the contact matches your audience.
email The published address; do not invent one.
source_url Lets you verify provenance.
captured_at Shows when the page was observed.
publication_context Stores nearby wording, such as “partnership inquiries.”
jurisdiction Supports the correct legal assessment.
restriction Records “no marketing,” “customers only,” or similar language.
collection_assessment Your documented reason for retaining the record.
campaign_relevance Explains why a proposed message relates to the role.
notice_status Tracks privacy notice, objection, unsubscribe, or unknown.
last_checked_at Supports accuracy and revalidation.

These fields are an operational recommendation based on purpose, minimisation, accuracy, and evidence principles. They are not a universal legal checklist.

4. A conservative Python collector

The following script reads a small, manually selected list of pages, respects a delay, extracts only published mailto: links and visible email-like text, and writes provenance to CSV. It does not guess addresses from names, crawl an entire site, bypass access controls, or send messages.

Preserve the page context and timestamp alongside each contact record.
Preserve the page context and timestamp alongside each contact record.
from __future__ import annotations

import csv
import re
import time
from datetime import datetime, timezone
from urllib.parse import unquote, urljoin, urlparse

import requests
from bs4 import BeautifulSoup

PAGES = [
    'https://example.com/contact',
    'https://example.org/team',
]
USER_AGENT = 'ResearchDirectoryBot/1.0 (replace-with-your-contact)'
EMAIL_RE = re.compile(r'[A-Z0-9._%+-]+@[A-Z0-9.-]+\\.[A-Z]{2,}', re.I)

session = requests.Session()
session.headers.update({'User-Agent': USER_AGENT})
rows = []

for page_url in PAGES:
    response = session.get(page_url, timeout=20)
    response.raise_for_status()
    soup = BeautifulSoup(response.text, 'html.parser')
    page_text = ' '.join(soup.stripped_strings)
    found = set()

    for link in soup.select('a[href^="mailto:"]'):
        address = unquote(link.get('href', '')[7:]).split('?', 1)[0].strip()
        if EMAIL_RE.fullmatch(address):
            found.add(address.lower())

    for address in EMAIL_RE.findall(page_text):
        found.add(address.lower())

    context = page_text[:500]
    for address in sorted(found):
        rows.append({
            'email': address,
            'source_url': page_url,
            'captured_at': datetime.now(timezone.utc).isoformat(),
            'publication_context': context,
            'collection_assessment': 'Review manually before retention or outreach',
            'notice_status': 'unknown',
        })
    time.sleep(2)

with open('public_contacts.csv', 'w', newline='', encoding='utf-8') as handle:
    writer = csv.DictWriter(handle, fieldnames=rows[0].keys() if rows else [
        'email', 'source_url', 'captured_at', 'publication_context',
        'collection_assessment', 'notice_status'
    ])
    writer.writeheader()
    writer.writerows(rows)

print(f'Wrote {len(rows)} records for manual review')

Before running it, read each site’s terms, robots guidance, and access restrictions. Keep the list small enough to review. A crawler’s ability to retrieve a page does not establish permission to collect or use the data.

5. Never generate addresses as a shortcut

Do not turn a name and domain into first.last@example.com and call it a public-data record. The Canadian Office of the Privacy Commissioner gives generated addresses as an example of why “not scraped” does not mean “consent exists.” Generated addresses can be wrong, personal, or unrelated to the intended role. Store only an address that was actually published or supplied through a compliant process.

6. Separate the right to collect from the right to send

For each record, make two decisions:

Retention and outreach require separate documented decisions.
Retention and outreach require separate documented decisions.
  1. Retention: Is there a documented, fair, purpose-limited reason to keep this personal data?
  2. Contact: Does the planned message satisfy the electronic-marketing rules for this recipient, channel, and jurisdiction?

In the EU, GDPR covers identifiable people acting professionally. Company data that identifies only a legal entity is treated differently, but a named employee remains a natural person. You still need a lawful, transparent purpose and a process for objections.

In the UK, the ICO says you cannot assume that an individual’s public data means agreement to direct marketing. UK GDPR and PECR can apply together. A publicly listed address is evidence of publication, not a universal opt-in.

In Canada, CASL generally requires express or qualifying implied consent. The CRTC says conspicuous publication can support implied consent only when the address is publicly displayed, there is no statement against receiving commercial electronic messages, and the message relates to the recipient’s business role, functions, or duties. The sender must prove those conditions (CRTC CASL guidance). ISED’s consent guidance explains the same general requirement (ISED consent guidance).

In the United States, CAN-SPAM applies to commercial email, including business-to-business email. FTC guidance requires accurate headers, non-deceptive subject lines, a physical postal address, an opt-out method, and honoring opt-outs within 10 business days (FTC CAN-SPAM compliance guide). It does not create a blanket permission to harvest addresses.

7. Preserve evidence and objections

Attach evidence to each record instead of relying on a vendor’s statement that a list is “compliant.” For a public-posting assessment, retain the URL, capture date, address, nearby wording, role, and whether a no-message instruction appeared. A screenshot can preserve context when the page may change. Keep a hash or immutable archive reference if your process requires tamper evidence.

Maintain a suppression table keyed by normalized address and, where appropriate, organization and person. Record the objection date, channel, source, and downstream systems updated. Propagate removals to exports, CRM segments, enrichment tools, and contractors. FTC guidance restricts transferring an opted-out address except to a compliance service provider. Canadian privacy guidance says the organization remains accountable when a supplier provides a list or runs a campaign (Office of the Privacy Commissioner of Canada).

8. Recheck before every campaign

  • Does the address still appear on the source page?
  • Is the person still in the relevant role?
  • Has the page added a no-contact or channel restriction?
  • Is the message genuinely related to the person’s business function?
  • Is an objection, unsubscribe, or suppression record present?
  • Does the message include the required sender identity, address, and opt-out controls?
  • Can you show why the selected legal and operational basis applies in this jurisdiction?

Recheck frequency should match risk and volatility. A frequently changing staff directory needs more frequent review than a stable statutory register. Update or delete records that are no longer accurate or necessary.

9. Source and contact comparison checklist

Question What to document
Who is identified? A legal entity, named employee, or ambiguous address.
Why was it published? Sales, support, press, recruiting, partnership, or unknown.
Is there a restriction? Any “no marketing,” “customers only,” or channel-specific notice.
Can you prove the observation? URL, timestamp, context, and retained evidence.
Is the message relevant? Specific connection to role, function, or stated inquiry type.
Which rules apply? Recipient country, sender country, channel, and purpose.
Can objections propagate? Suppression workflow across every export and supplier.

10. Or skip the browser setup

If you need a reliable record of what a public page looked like when collected, ScreenshotNeo can capture the page through one GET request. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Only clean shots are billed: bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status.

See the ScreenshotNeo API documentation for all options. This is a complete cURL example:

curl -G 'https://api.screenshotneo.com/v1/shot' \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Python:

import requests

r = requests.get(
    'https://api.screenshotneo.com/v1/shot',
    params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'},
    timeout=90,
)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', image);

For provenance workflows, use full-page capture when the context spans multiple sections, a CSS selector when only the contact block matters, custom CSS to hide irrelevant content, and a wait-for-selector or network-idle wait for JavaScript-rendered directories. You can also set a viewport or device preset, retina scale, timezone, geolocation, custom headers, cookies, user agent, or Authorization header. Caching with a chosen TTL helps avoid duplicate captures. Signed links are available when a public <img> reference is needed. Async jobs with signed webhooks, bulk capture for up to 100 URLs per call, PDF output, usage reporting, and an OpenAPI specification support larger evidence pipelines. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

Start with 1,000 screenshots a month free, with no card. Paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

11. Troubleshooting

The script finds no addresses

The page may render contacts with JavaScript, use an obfuscated address, expose only a form, or block the request. Inspect the HTML response, confirm the address is visibly published, and use a browser-rendered capture for evidence. Do not decode an address that the publisher intentionally restricted or bypass an access control.

The same address appears many times

Normalize case, trim whitespace, and deduplicate on the address plus source context. Keep multiple source URLs when they establish different purposes or restrictions.

A contact asks not to be contacted

Add the address to suppression immediately, stop queued messages, propagate the suppression to suppliers, and retain the objection evidence. Do not rely on deleting one CRM row.

A page changes after collection

Keep the original URL, timestamp, surrounding text, and an evidence capture. Mark the record stale and re-evaluate before use. A later page state does not erase what you observed, but it may change whether continued retention is justified.

The HTTP request returns a bot check or blank page

Do not retry aggressively. Slow down, check the site’s access terms, and use a permitted browser or screenshot workflow. With ScreenshotNeo, bot checks, blank pages, timeouts, failed loads, and cache hits are identified in response headers and are not billed.

12. Performance, reliability, and cost

Small, reviewable batches are safer than a high-volume crawl. Use connection reuse, bounded concurrency, exponential backoff for transient failures, and a clear stop condition. Store raw responses or evidence separately from the normalized database so a parser change does not destroy provenance. Treat HTTP success as transport success, not proof that the page contained a valid contact.

Budget for rechecks, evidence storage, suppression propagation, and manual review. A free or inexpensive collection script can create expensive compliance work if it produces untraceable records. Screenshot costs should be measured by clean captures rather than raw requests when choosing an evidence service; ScreenshotNeo bills only clean shots and offers a free monthly tier.

FAQ

Does a public business email mean I can send marketing?

No. Public visibility is not universal consent. Apply the recipient’s jurisdictional privacy and electronic-marketing rules and document the decision.

Can I buy a public-data email list?

A supplier’s claim does not transfer your responsibility. Ask for source URLs, dates, publication context, consent or legal-basis records, suppression handling, and update procedures before considering any list.

Should generic addresses such as info@ be treated differently?

They may identify a company rather than a person, but purpose and channel rules still matter. Respect stated restrictions and avoid assuming that a generic inbox welcomes unsolicited commercial messages.

How long should records be kept?

Keep them only while they serve the documented purpose and remain accurate. Set a review or deletion date based on source volatility, campaign risk, and applicable retention duties.

What is the minimum defensible record?

Address, source URL, capture date, publication context, relevance rationale, jurisdiction, restriction check, collection assessment, and suppression status.

This is a practical research guide, not a universal legal conclusion. Rules vary by recipient type, channel, jurisdiction, and purpose. Obtain advice for your specific campaign.