ScreenshotNeo

BlogEngineering

A Guide to Matching Web-Scraped Data: Deduplicate and Reconcile Records

Learn how to deduplicate scraped records, link inconsistent fields, evaluate matches, and reconcile values with auditable Python workflows.

By the ScreenshotNeo team1 October 20266 min read

To deduplicate web-scraped data, preserve every source row, normalize comparison fields, match exact identifiers first, generate candidate pairs before fuzzy comparison, evaluate decisions with labeled examples, and reconcile values only after matching. Keep raw values, normalized values, source URLs, capture times, match scores, and decision reasons so every merge can be explained or reversed.

What matching, deduplication, and reconciliation mean

Deduplication usually removes repeated records inside one dataset. Record linkage connects records from different datasets that describe the same entity. Entity resolution is the broader task of deciding which records refer to the same real-world entity. Scraped-data projects often use all three terms, so define the decision you are making before writing rules.

A match group answers “which rows belong together?” Reconciliation answers “which value should the canonical record keep?” These are separate decisions. A product can be matched across shops while its price, title, stock status, and description are selected using different survivorship rules.

1. Preserve identity and provenance before cleaning

Assign a stable key to every scraped row. Never use a normalized name or URL as the only identity: two distinct products can normalize to the same text, and the same page can change between captures.

source_record_id = "shop_a:products/123:2026-09-30T12:45:00Z"
source_name = "shop_a"
source_url = "https://example.com/products/123"
captured_at = "2026-09-30T12:45:00Z"
raw_name = "Acme  Widget (Blue)"
raw_price = "$19.99"

Retain the original HTML-derived values, parser version, source, collection time, and any request metadata needed to reproduce the extraction. A unique input ID is also a requirement in AWS Entity Resolution matching workflows; the general design principle applies to custom pipelines as well.

2. Normalize comparison fields deliberately

Normalization makes equivalent representations comparable while preserving raw data for review. Typical operations include trimming whitespace, case folding, Unicode normalization, punctuation handling, and format-specific parsing.

Field Useful normalization Do not erase
Name Unicode NFKC, case folding, repeated-space removal Model numbers, edition markers, variant names
Address Standardize whitespace, postal-code format, common abbreviations Apartment, suite, unit, building identifiers
Phone Parse country code and digits Extension numbers and country context
Email Trim and case-fold the domain; apply local-part rules only when justified Provider-specific aliases unless documented
URL Lowercase host, remove tracking parameters, normalize trailing slash Path segments that identify products or editions
Price Parse decimal and currency separately Currency, tax state, unit size

AWS describes default normalization as removing special characters and extra spaces and formatting text to lowercase. Treat that as an example, not a universal recipe: field semantics determine what is safe to remove.

3. Match strong identifiers exactly first

Start with identifiers that are stable and specific: a merchant SKU plus source, ISBN, GTIN, canonical URL, account ID, or a verified email address. Use exact rules before fuzzy similarity because they are easier to audit.

def exact_key(row):
    sku = row.get("normalized_sku")
    source = row.get("source_name")
    return (source, sku) if sku else None

exact_groups = {}
for row in rows:
    key = exact_key(row)
    if key:
        exact_groups.setdefault(key, []).append(row["source_record_id"])

Do not force an exact rule when an identifier is reused, missing, or known to change. Record which rule produced each link.

4. Generate candidate pairs with blocking

Comparing every row with every other row grows quadratically. Blocking, also called indexing, creates plausible candidate pairs first. Examples include the same postal-code prefix, the same normalized domain, the same manufacturer plus model family, or a phonetic surname key.

from collections import defaultdict

blocks = defaultdict(list)
for row in rows:
    key = (row.get("country"), row.get("postal_prefix"), row.get("name_initial"))
    blocks[key].append(row)

candidate_pairs = []
for bucket in blocks.values():
    for i, left in enumerate(bucket):
        for right in bucket[i + 1:]:
            if left["source_name"] != right["source_name"]:
                candidate_pairs.append((left, right))

Blocking can miss true matches when its key is too restrictive. Measure candidate coverage against labeled examples and use multiple blocking keys when recall matters.

5. Compare candidates with explainable features

For names, addresses, and descriptions, calculate several similarities instead of trusting one score. Useful features include token overlap, edit distance, character n-gram similarity, numeric agreement, and whether a strong identifier agrees or conflicts.

from difflib import SequenceMatcher
import re

def norm(value):
    value = (value or "").casefold().strip()
    return re.sub(r"\\s+", " ", value)

def ratio(a, b):
    return SequenceMatcher(None, norm(a), norm(b)).ratio()

def compare(a, b):
    return {
        "name_similarity": ratio(a.get("name"), b.get("name")),
        "address_similarity": ratio(a.get("address"), b.get("address")),
        "same_postal_code": bool(a.get("postal_code")) and a.get("postal_code") == b.get("postal_code"),
        "same_sku": bool(a.get("sku")) and a.get("sku") == b.get("sku"),
    }

def decision(features):
    if features["same_sku"]:
        return "match", "exact_sku"
    score = 0.55 * features["name_similarity"] + 0.35 * features["address_similarity"] + 0.10 * features["same_postal_code"]
    if score >= 0.90:
        return "match", f"weighted_score:{score:.3f}"
    if score <= 0.55:
        return "nonmatch", f"weighted_score:{score:.3f}"
    return "review", f"weighted_score:{score:.3f}"

Thresholds are domain decisions, not constants supplied by a source. A false positive can merge two people or products; a false negative can leave fragmented histories. Keep a review band for uncertain cases.

6. Evaluate with labeled examples

Create a small, representative set of candidate pairs labeled match or nonmatch. Include difficult cases: missing fields, transliteration, product variants, shared addresses, reused SKUs, and changed names. Report precision (how many accepted matches are correct) and recall (how many true matches were found). Inspect false positives and false negatives, then adjust normalization, blocking, features, or thresholds.

The U.S. Census quality standard treats automated record linkage as a process requiring documentation and evaluation. The Record Linkage Toolkit describes cleaning, indexing, comparing, classifying, and evaluation as connected workflow steps. Neither source establishes one universal threshold for scraped data.

7. Reconcile matched groups into canonical records

After forming a match group, choose field values with explicit survivorship rules. Examples:

  • Prefer a source ranked as authoritative for that field.
  • Prefer the most recent capture when freshness matters.
  • Prefer values with complete address components.
  • Retain all conflicting prices with currency and capture time instead of silently overwriting.
  • Store contributing source IDs and the rule that selected each canonical value.
def choose_value(records, field, source_rank):
    available = [r for r in records if r.get(field) not in (None, "")]
    if not available:
        return None, []
    available.sort(key=lambda r: (source_rank.get(r["source_name"], 999), -r["captured_at_epoch"]))
    winner = available[0]
    return winner[field], [r["source_record_id"] for r in available]

Keep a merge log containing group ID, member IDs, rule version, selected values, discarded alternatives, and review status. This makes merges reversible.

Exact rules, fuzzy rules, or machine learning?

Approach Strength Risk or cost Best use
Exact rules Fast and highly explainable Misses formatting and spelling variation Reliable IDs and high-precision links
Fuzzy rules Transparent handling of small variations Threshold tuning and false positives Names, addresses, descriptions
ML matching Combines fields and can handle missing values Needs training or labeled data and careful evaluation Large, variable datasets with review capacity

AWS Entity Resolution documents rule-based exact and fuzzy matching plus machine-learning workflows. Its ML workflow considers fields together and accounts for missing fields, but a confidence value is not proof that two rows are identical. Choose based on precision, recall, explainability, scale, reviewability, and reversal requirements.

Complete Python workflow

import csv, re, unicodedata
from difflib import SequenceMatcher
from collections import defaultdict

def normalize(value):
    value = unicodedata.normalize("NFKC", value or "").casefold().strip()
    value = re.sub(r"\\s+", " ", value)
    return value

def sim(a, b):
    return SequenceMatcher(None, normalize(a), normalize(b)).ratio()

with open("scraped.csv", newline="", encoding="utf-8") as f:
    rows = list(csv.DictReader(f))

for r in rows:
    r["norm_name"] = normalize(r.get("name"))
    r["norm_address"] = normalize(r.get("address"))
    r["norm_sku"] = normalize(r.get("sku"))

blocks = defaultdict(list)
for r in rows:
    key = (r.get("country", ""), normalize(r.get("postal_code", ""))[:3])
    blocks[key].append(r)

links = []
for bucket in blocks.values():
    for i, left in enumerate(bucket):
        for right in bucket[i + 1:]:
            if left["source_name"] == right["source_name"]:
                continue
            exact_sku = left["norm_sku"] and left["norm_sku"] == right["norm_sku"]
            score = 1.0 if exact_sku else 0.6 * sim(left["name"], right["name"]) + 0.4 * sim(left["address"], right["address"])
            status = "match" if score >= 0.90 else "review" if score >= 0.60 else "nonmatch"
            links.append({"left_id": left["source_record_id"], "right_id": right["source_record_id"], "score": round(score, 4), "status": status})

with open("match_decisions.csv", "w", newline="", encoding="utf-8") as f:
    writer = csv.DictWriter(f, fieldnames=["left_id", "right_id", "score", "status"])
    writer.writeheader(); writer.writerows(links)

Common errors and fixes

Symptom Likely cause Fix
Many false matches Over-aggressive normalization or low threshold Preserve discriminating tokens, add blocking fields, raise threshold, expand review band
True matches never become candidates Blocking key is too restrictive Add alternate keys and measure candidate coverage
Variants collapse together Model, edition, unit, or apartment information was removed Keep semantic tokens and compare them separately
Scores change after reruns Unversioned parser, normalization, or source data Version rules and retain capture timestamps
Canonical values cannot be explained Matching and survivorship were combined Store field-level winner IDs and rule versions
Memory or runtime grows sharply All-pairs comparison Use blocking, chunking, indexes, and batch writes

Performance, reliability, and cost

  • Performance: blocking and exact indexes reduce comparisons; cache normalized fields; process candidates in chunks; avoid repeatedly parsing the same HTML.
  • Reliability: make matching jobs idempotent, checkpoint batches, retain failed rows, and version rules so a rerun is reproducible.
  • Cost: fuzzy comparisons and ML inference cost more than exact joins. Sample difficult pairs for review instead of sending every row through an expensive model.
  • Operations: monitor candidate counts, match and review rates, field missingness, and drift by source. Sudden changes often indicate a scraper or schema change.

Or skip the browser setup

If your pipeline needs screenshots of source pages for provenance or visual verification, ScreenshotNeo provides a website screenshot API at screenshotneo.com. It removes cookie banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; and response headers identify the page verdict and billing status. Its MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for capture options. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Should I deduplicate before or after scraping?

Preserve and store every scrape first. Deduplicate in a repeatable downstream step so parser changes and source corrections can be replayed.

Is a high fuzzy score enough to merge?

No. Scores are evidence. Validate them on labeled examples and keep uncertain pairs for review.

How do I handle missing fields?

Use available-field features, record which fields were missing, and avoid treating missing values as agreement.

Can I merge records from different countries?

Only with country-aware normalization. Phone, address, postal-code, currency, and language rules vary by locale.

How can I undo a bad merge?

Keep immutable source IDs, match decisions, group history, and field-level survivorship records. Rebuild canonical records from that log.