A Guide to Matching Web-Scraped Data: Deduplicate and Reconcile Records
Learn how to deduplicate scraped records, link inconsistent fields, evaluate matches, and reconcile values with auditable Python workflows.
To deduplicate web-scraped data, preserve every source row, normalize comparison fields, match exact identifiers first, generate candidate pairs before fuzzy comparison, evaluate decisions with labeled examples, and reconcile values only after matching. Keep raw values, normalized values, source URLs, capture times, match scores, and decision reasons so every merge can be explained or reversed.
What matching, deduplication, and reconciliation mean
Deduplication usually removes repeated records inside one dataset. Record linkage connects records from different datasets that describe the same entity. Entity resolution is the broader task of deciding which records refer to the same real-world entity. Scraped-data projects often use all three terms, so define the decision you are making before writing rules.
A match group answers “which rows belong together?” Reconciliation answers “which value should the canonical record keep?” These are separate decisions. A product can be matched across shops while its price, title, stock status, and description are selected using different survivorship rules.
1. Preserve identity and provenance before cleaning
Assign a stable key to every scraped row. Never use a normalized name or URL as the only identity: two distinct products can normalize to the same text, and the same page can change between captures.
source_record_id = "shop_a:products/123:2026-09-30T12:45:00Z"
source_name = "shop_a"
source_url = "https://example.com/products/123"
captured_at = "2026-09-30T12:45:00Z"
raw_name = "Acme Widget (Blue)"
raw_price = "$19.99"
Retain the original HTML-derived values, parser version, source, collection time, and any request metadata needed to reproduce the extraction. A unique input ID is also a requirement in AWS Entity Resolution matching workflows; the general design principle applies to custom pipelines as well.
2. Normalize comparison fields deliberately
Normalization makes equivalent representations comparable while preserving raw data for review. Typical operations include trimming whitespace, case folding, Unicode normalization, punctuation handling, and format-specific parsing.
| Field | Useful normalization | Do not erase |
|---|---|---|
| Name | Unicode NFKC, case folding, repeated-space removal | Model numbers, edition markers, variant names |
| Address | Standardize whitespace, postal-code format, common abbreviations | Apartment, suite, unit, building identifiers |
| Phone | Parse country code and digits | Extension numbers and country context |
| Trim and case-fold the domain; apply local-part rules only when justified | Provider-specific aliases unless documented | |
| URL | Lowercase host, remove tracking parameters, normalize trailing slash | Path segments that identify products or editions |
| Price | Parse decimal and currency separately | Currency, tax state, unit size |
AWS describes default normalization as removing special characters and extra spaces and formatting text to lowercase. Treat that as an example, not a universal recipe: field semantics determine what is safe to remove.
3. Match strong identifiers exactly first
Start with identifiers that are stable and specific: a merchant SKU plus source, ISBN, GTIN, canonical URL, account ID, or a verified email address. Use exact rules before fuzzy similarity because they are easier to audit.
def exact_key(row):
sku = row.get("normalized_sku")
source = row.get("source_name")
return (source, sku) if sku else None
exact_groups = {}
for row in rows:
key = exact_key(row)
if key:
exact_groups.setdefault(key, []).append(row["source_record_id"])
Do not force an exact rule when an identifier is reused, missing, or known to change. Record which rule produced each link.
4. Generate candidate pairs with blocking
Comparing every row with every other row grows quadratically. Blocking, also called indexing, creates plausible candidate pairs first. Examples include the same postal-code prefix, the same normalized domain, the same manufacturer plus model family, or a phonetic surname key.
from collections import defaultdict
blocks = defaultdict(list)
for row in rows:
key = (row.get("country"), row.get("postal_prefix"), row.get("name_initial"))
blocks[key].append(row)
candidate_pairs = []
for bucket in blocks.values():
for i, left in enumerate(bucket):
for right in bucket[i + 1:]:
if left["source_name"] != right["source_name"]:
candidate_pairs.append((left, right))
Blocking can miss true matches when its key is too restrictive. Measure candidate coverage against labeled examples and use multiple blocking keys when recall matters.
5. Compare candidates with explainable features
For names, addresses, and descriptions, calculate several similarities instead of trusting one score. Useful features include token overlap, edit distance, character n-gram similarity, numeric agreement, and whether a strong identifier agrees or conflicts.
from difflib import SequenceMatcher
import re
def norm(value):
value = (value or "").casefold().strip()
return re.sub(r"\\s+", " ", value)
def ratio(a, b):
return SequenceMatcher(None, norm(a), norm(b)).ratio()
def compare(a, b):
return {
"name_similarity": ratio(a.get("name"), b.get("name")),
"address_similarity": ratio(a.get("address"), b.get("address")),
"same_postal_code": bool(a.get("postal_code")) and a.get("postal_code") == b.get("postal_code"),
"same_sku": bool(a.get("sku")) and a.get("sku") == b.get("sku"),
}
def decision(features):
if features["same_sku"]:
return "match", "exact_sku"
score = 0.55 * features["name_similarity"] + 0.35 * features["address_similarity"] + 0.10 * features["same_postal_code"]
if score >= 0.90:
return "match", f"weighted_score:{score:.3f}"
if score <= 0.55:
return "nonmatch", f"weighted_score:{score:.3f}"
return "review", f"weighted_score:{score:.3f}"
Thresholds are domain decisions, not constants supplied by a source. A false positive can merge two people or products; a false negative can leave fragmented histories. Keep a review band for uncertain cases.
6. Evaluate with labeled examples
Create a small, representative set of candidate pairs labeled match or nonmatch. Include difficult cases: missing fields, transliteration, product variants, shared addresses, reused SKUs, and changed names. Report precision (how many accepted matches are correct) and recall (how many true matches were found). Inspect false positives and false negatives, then adjust normalization, blocking, features, or thresholds.
The U.S. Census quality standard treats automated record linkage as a process requiring documentation and evaluation. The Record Linkage Toolkit describes cleaning, indexing, comparing, classifying, and evaluation as connected workflow steps. Neither source establishes one universal threshold for scraped data.
7. Reconcile matched groups into canonical records
After forming a match group, choose field values with explicit survivorship rules. Examples:
- Prefer a source ranked as authoritative for that field.
- Prefer the most recent capture when freshness matters.
- Prefer values with complete address components.
- Retain all conflicting prices with currency and capture time instead of silently overwriting.
- Store contributing source IDs and the rule that selected each canonical value.
def choose_value(records, field, source_rank):
available = [r for r in records if r.get(field) not in (None, "")]
if not available:
return None, []
available.sort(key=lambda r: (source_rank.get(r["source_name"], 999), -r["captured_at_epoch"]))
winner = available[0]
return winner[field], [r["source_record_id"] for r in available]
Keep a merge log containing group ID, member IDs, rule version, selected values, discarded alternatives, and review status. This makes merges reversible.
Exact rules, fuzzy rules, or machine learning?
| Approach | Strength | Risk or cost | Best use |
|---|---|---|---|
| Exact rules | Fast and highly explainable | Misses formatting and spelling variation | Reliable IDs and high-precision links |
| Fuzzy rules | Transparent handling of small variations | Threshold tuning and false positives | Names, addresses, descriptions |
| ML matching | Combines fields and can handle missing values | Needs training or labeled data and careful evaluation | Large, variable datasets with review capacity |
AWS Entity Resolution documents rule-based exact and fuzzy matching plus machine-learning workflows. Its ML workflow considers fields together and accounts for missing fields, but a confidence value is not proof that two rows are identical. Choose based on precision, recall, explainability, scale, reviewability, and reversal requirements.
Complete Python workflow
import csv, re, unicodedata
from difflib import SequenceMatcher
from collections import defaultdict
def normalize(value):
value = unicodedata.normalize("NFKC", value or "").casefold().strip()
value = re.sub(r"\\s+", " ", value)
return value
def sim(a, b):
return SequenceMatcher(None, normalize(a), normalize(b)).ratio()
with open("scraped.csv", newline="", encoding="utf-8") as f:
rows = list(csv.DictReader(f))
for r in rows:
r["norm_name"] = normalize(r.get("name"))
r["norm_address"] = normalize(r.get("address"))
r["norm_sku"] = normalize(r.get("sku"))
blocks = defaultdict(list)
for r in rows:
key = (r.get("country", ""), normalize(r.get("postal_code", ""))[:3])
blocks[key].append(r)
links = []
for bucket in blocks.values():
for i, left in enumerate(bucket):
for right in bucket[i + 1:]:
if left["source_name"] == right["source_name"]:
continue
exact_sku = left["norm_sku"] and left["norm_sku"] == right["norm_sku"]
score = 1.0 if exact_sku else 0.6 * sim(left["name"], right["name"]) + 0.4 * sim(left["address"], right["address"])
status = "match" if score >= 0.90 else "review" if score >= 0.60 else "nonmatch"
links.append({"left_id": left["source_record_id"], "right_id": right["source_record_id"], "score": round(score, 4), "status": status})
with open("match_decisions.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=["left_id", "right_id", "score", "status"])
writer.writeheader(); writer.writerows(links)
Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Many false matches | Over-aggressive normalization or low threshold | Preserve discriminating tokens, add blocking fields, raise threshold, expand review band |
| True matches never become candidates | Blocking key is too restrictive | Add alternate keys and measure candidate coverage |
| Variants collapse together | Model, edition, unit, or apartment information was removed | Keep semantic tokens and compare them separately |
| Scores change after reruns | Unversioned parser, normalization, or source data | Version rules and retain capture timestamps |
| Canonical values cannot be explained | Matching and survivorship were combined | Store field-level winner IDs and rule versions |
| Memory or runtime grows sharply | All-pairs comparison | Use blocking, chunking, indexes, and batch writes |
Performance, reliability, and cost
- Performance: blocking and exact indexes reduce comparisons; cache normalized fields; process candidates in chunks; avoid repeatedly parsing the same HTML.
- Reliability: make matching jobs idempotent, checkpoint batches, retain failed rows, and version rules so a rerun is reproducible.
- Cost: fuzzy comparisons and ML inference cost more than exact joins. Sample difficult pairs for review instead of sending every row through an expensive model.
- Operations: monitor candidate counts, match and review rates, field missingness, and drift by source. Sudden changes often indicate a scraper or schema change.
Or skip the browser setup
If your pipeline needs screenshots of source pages for provenance or visual verification, ScreenshotNeo provides a website screenshot API at screenshotneo.com. It removes cookie banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; and response headers identify the page verdict and billing status. Its MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for capture options. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
FAQ
Should I deduplicate before or after scraping?
Preserve and store every scrape first. Deduplicate in a repeatable downstream step so parser changes and source corrections can be replayed.
Is a high fuzzy score enough to merge?
No. Scores are evidence. Validate them on labeled examples and keep uncertain pairs for review.
How do I handle missing fields?
Use available-field features, record which fields were missing, and avoid treating missing values as agreement.
Can I merge records from different countries?
Only with country-aware normalization. Phone, address, postal-code, currency, and language rules vary by locale.
How can I undo a bad merge?
Keep immutable source IDs, match decisions, group history, and field-level survivorship records. Rebuild canonical records from that log.


