ScreenshotNeo

BlogGuides

How to Generate Marketing Leads with Web Scraping

Build a compliant web-scraping workflow for finding, validating, and routing qualified marketing prospects into your CRM.

By the ScreenshotNeo team1 October 20268 min read

Web scraping can help you identify organizations that match your ideal customer profile (ICP), collect a small set of relevant public business fields, validate those records, and route qualified prospects into a CRM. It is a research workflow, not permission to collect every visible contact detail or send indiscriminate outreach.

The practical sequence is:

  1. Define the organizations you want and the problem your offer solves.
  2. Choose sources whose rules permit your intended collection and use.
  3. Collect only the minimum fields needed for qualification.
  4. Validate, deduplicate, and record provenance.
  5. Apply privacy, suppression, and opt-out controls before outreach.

1. Define an ideal customer profile before scraping

Start with observable fit signals instead of a list of every field a page exposes. Write down:

  • Industry: for example, ecommerce, accounting, or logistics.
  • Geography: countries, regions, or service areas you can support.
  • Company size: employee range, locations, revenue band, or another public proxy.
  • Business trait: a technology used, a hiring pattern, a product category, or a newly opened location.
  • Problem solved: the operational issue that makes the company a plausible buyer.

Create a sample target list first. If you cannot explain why each sample organization fits, more scraping will only produce more noise.

2. Check permission, terms, and privacy scope

Technical accessibility does not establish permission to collect or use data. Review the source’s terms, robots guidance where relevant, platform rules, and the laws that apply to your organization and outreach channel. Do not bypass login controls, CAPTCHAs, bot checks, paywalls, or other access restrictions.

LinkedIn’s User Agreement prohibits third-party software that scrapes or automates activity on its site. Treat that as a source-specific restriction, not as a challenge to work around.

This article uses UK ICO guidance as its legal example. The ICO direct-marketing guidance says collection for direct marketing must be fair, lawful, and transparent. Public availability does not remove privacy obligations. UK B2B outreach can still involve UK GDPR when information identifies an individual, as explained in the ICO’s electronic-mail marketing guidance. Verify requirements for other jurisdictions separately.

3. Minimize the fields you collect

Separate organization-level research from personal data. A useful first schema might be:

Field Purpose Example validation
organization_name Identify the company Non-empty, normalized case
source_url Provenance and review Canonical URL stored
retrieved_at Freshness tracking UTC timestamp
industry_signal ICP qualification Matches allowed categories
location Geographic fit Normalized country or region
public_contact_route Choose a lawful channel Generic contact page or role inbox
fit_notes Human review Short evidence-based note

Do not collect names, personal email addresses, phone numbers, or social profiles unless they are necessary, lawful for your purpose, and covered by your privacy process. The ICO says organizations should consider whether using public-source personal information for marketing would be unexpected and whether the processing is fair and lawful.

4. Build a Scrapy crawler for permitted sources

Scrapy’s documentation covers structured extraction and export. The example below crawls a fictional directory that you are authorized to access. Replace selectors and the domain with a permitted source; do not use it to defeat access controls.

Project setup

python -m venv .venv
source .venv/bin/activate
pip install scrapy
scrapy startproject leadcrawler
cd leadcrawler

Spider

import scrapy
from datetime import datetime, timezone

class DirectorySpider(scrapy.Spider):
    name = 'directory'
    allowed_domains = ['example.com']
    start_urls = ['https://example.com/directory']

    custom_settings = {
        'DOWNLOAD_DELAY': 1.0,
        'CONCURRENT_REQUESTS_PER_DOMAIN': 2,
        'AUTOTHROTTLE_ENABLED': True,
        'AUTOTHROTTLE_START_DELAY': 1.0,
        'AUTOTHROTTLE_MAX_DELAY': 10.0,
        'FEEDS': {
            'leads.jsonl': {
                'format': 'jsonlines',
                'overwrite': True,
            }
        },
    }

    def parse(self, response):
        for card in response.css('article.company-card'):
            name = card.css('h2::text').get()
            detail_url = response.urljoin(card.css('a::attr(href)').get())
            yield scrapy.Request(
                detail_url,
                callback=self.parse_company,
                cb_kwargs={'name': name.strip() if name else None},
            )

        next_url = response.css('a[rel="next"]::attr(href)').get()
        if next_url:
            yield response.follow(next_url, callback=self.parse)

    def parse_company(self, response, name):
        yield {
            'organization_name': name,
            'source_url': response.url,
            'retrieved_at': datetime.now(timezone.utc).isoformat(),
            'industry_signal': response.css('[data-field="industry"]::text').get(),
            'location': response.css('[data-field="location"]::text').get(),
            'public_contact_route': response.css('a.contact::attr(href)').get(),
        }

Save this as leadcrawler/spiders/directory.py and run:

scrapy crawl directory

The crawler uses a delay, per-domain concurrency limit, and AutoThrottle. Those settings reduce load; they do not grant permission to crawl a site.

5. Configure crawling carefully

Control Why it matters
Allowed domains Prevents accidental requests to unrelated hosts.
Download delay Spaces requests and lowers server load.
Concurrency Limits simultaneous requests per domain.
AutoThrottle Adapts request rate to observed latency.
Pagination bounds Prevents unbounded crawls and duplicate loops.
Retry policy Retries transient failures without hammering a source.
Export format JSON Lines is convenient for streaming validation and replay.

Use a small test run before scaling. Log status codes, redirects, parser warnings, and the source URL for every record. Stop when the source signals that your traffic is unwelcome.

6. Validate and deduplicate records

Extraction produces candidates, not ready-to-contact leads. Add a validation step that:

  • Normalizes organization names, domains, countries, and URLs.
  • Rejects records missing the fields required by your ICP.
  • Checks that URLs resolve and that fields match the page evidence.
  • Deduplicates by canonical domain, registration identifier, or a reviewed composite key.
  • Flags uncertain values instead of silently guessing.
  • Stores source_url and retrieved_at so a reviewer can trace each value.
import json
from urllib.parse import urlparse

seen_domains = set()
with open('leads.jsonl', encoding='utf-8') as source, open('qualified.jsonl', 'w', encoding='utf-8') as output:
    for line in source:
        row = json.loads(line)
        url = row.get('public_contact_route') or row.get('source_url')
        domain = urlparse(url).netloc.lower().removeprefix('www.') if url else ''
        industry = (row.get('industry_signal') or '').lower()
        if not domain or domain in seen_domains:
            continue
        if industry not in {'saas', 'software', 'logistics'}:
            continue
        seen_domains.add(domain)
        output.write(json.dumps(row, ensure_ascii=False) + '\n')

7. Route only qualified records into your CRM

Keep the scraped dataset separate from your CRM until a qualification rule passes. Map fields explicitly, retain provenance, and record the lawful contact route. Add suppression-list checks before any campaign. The ICO says people have an absolute right to object to or opt out of direct marketing, and privacy information for data collected from other sources must be supplied within a reasonable period and no later than one month in the UK context. See the ICO marketing guidance for the applicable details.

Buying enrichment data does not transfer responsibility to the vendor. The ICO’s guidance on marketing data brokers says the organization using the data remains responsible and should establish its lawful basis before obtaining personal data.

8. Or skip the browser setup

If your workflow needs screenshots of company pages for human review, evidence, or an AI agent, ScreenshotNeo provides a single GET request. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result.

See the ScreenshotNeo API documentation for the available options.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', buffer);

ScreenshotNeo also supports full-page capture, CSS-element capture, device presets, custom viewport and retina scale, dark mode, custom CSS and JavaScript, click actions, selector waits, delays, network-idle waits, request blocking, headers, cookies, user agents, Authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

There are 1,000 free screenshots each month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is included on every plan. Create a free ScreenshotNeo account.

9. Performance, reliability, and cost

  • Performance: limit concurrency per domain, cache pages where permitted, and process exports as streams so memory does not grow with the crawl.
  • Reliability: make parsers tolerant of missing fields, persist checkpoints, retry only transient errors, and keep raw evidence for reprocessing.
  • Freshness: assign a refresh interval based on how quickly the source changes; store retrieval dates and do not imply current accuracy indefinitely.
  • Cost: Scrapy is open source, but hosting, storage, proxies, review time, and enrichment can still cost money. Do not judge a workflow by lead volume alone; measure qualified-record rate and review effort.
  • Compliance cost: privacy notices, suppression handling, legal review, and source monitoring are part of the operating cost.

10. Troubleshooting

Symptom Likely cause Fix
Empty fields Selector no longer matches or content is rendered by JavaScript. Inspect the permitted page structure, update selectors, or use an authorized rendered source. Do not bypass access controls.
Duplicate companies Multiple URLs represent one organization. Canonicalize domains and apply a reviewed deduplication key.
429 responses Request rate is too high. Stop, reduce concurrency, increase delay, and confirm the source permits collection.
403 or CAPTCHA The source is restricting automated access. Do not evade it. Use an allowed source or request permission.
Spider loops forever Unbounded pagination, calendars, or tracking parameters. Allow-list pagination links, strip known tracking parameters, and set page or item limits.
Stale CRM records No refresh policy or provenance. Store retrieval dates, schedule rechecks, and route uncertain records for review.
Outreach complaints Unexpected personal-data use or missing opt-out handling. Pause the campaign, honor objections, review lawful basis and notices, and consult the applicable regulator guidance.

11. Checklist before sending a campaign

  • ICP and qualification rules are written down.
  • Each source permits the intended collection and use.
  • No login, CAPTCHA, or access restriction was bypassed.
  • Only necessary fields were collected.
  • Every record has a source URL and retrieval date.
  • Duplicates, invalid values, and uncertainty were reviewed.
  • Privacy information and lawful-basis decisions are documented.
  • Suppression lists and objections are applied.
  • The CRM receives qualified records only.

FAQ

No. Public visibility does not settle permission, privacy, terms, or intended-use questions. Assess the source rules and applicable law.

Can I scrape LinkedIn for leads?

LinkedIn’s User Agreement prohibits third-party scraping and automation. Do not build a workflow that violates those terms.

What is the best field to deduplicate on?

A canonical company domain is a useful starting point, but confirm mergers, subsidiaries, franchises, and multi-domain organizations manually.

Does a data broker remove my compliance duties?

No. The organization using the data remains responsible for its lawful basis, transparency, accuracy, and opt-out handling.

When should I use screenshots in lead research?

Use them when a reviewer or agent needs visual evidence of a page, pricing display, product category, or qualification signal. Store the capture with the source URL and date.