ScreenshotNeo

BlogGuides

How to Collect Data for Machine Learning

A practical guide to sourcing, labeling, validating, protecting, and monitoring machine-learning data without bias or leakage.

By the ScreenshotNeo team29 September 20269 min read

How to Collect Data for Machine Learning

Collecting data for machine learning starts with the decision your model must make. Define the target, the people and conditions involved, and the cost of errors before you download a dataset or build a scraper. Then choose sources, document permissions and provenance, collect representative examples, label them with written rules, test quality, split and version the data, and monitor it after deployment.

There is no universal minimum number of rows. A dataset is adequate when it covers the cases the system will meet, labels are accurate enough for the intended error budget, evaluation is reliable, and known gaps are understood and monitored.

1. Define the data need before collecting

Write a one-page data specification. It should answer:

  • Decision: What will the model predict or recommend?
  • Unit of observation: A customer, transaction, image, document, session, device reading, or time window?
  • Target label: The answer to predict. AWS describes a supervised example as a target plus variables or features (AWS definition).
  • Features: Information available when the prediction is made, excluding future information.
  • Operating range: Languages, geographies, devices, seasons, lighting, network conditions, and user groups.
  • Error budget: Which mistakes are acceptable, and which require human review?
  • Intended use: Where will outputs be shown, and what actions can they trigger?

List positive and negative cases, rare but consequential cases, and examples that should be rejected. This prevents a convenient sample from silently becoming the specification.

2. Choose a collection approach

Most projects combine several sources. Compare them on coverage, label cost, expected error, consent and legal basis, provenance, privacy risk, update frequency, and operational cost.

A reproducible collection pipeline connects sources, labels, quality checks, and dataset versions.
A reproducible collection pipeline connects sources, labels, quality checks, and dataset versions.
Source Good for Risks and checks
Existing labeled dataset Fast baseline and model prototyping License, stale distributions, hidden label policy, duplicate records
Operational records Real production behavior and outcomes Selection bias, missing outcomes, historical decisions encoded as labels
Directly contributed data Surveys, uploads, interviews, user feedback Informed consent, incentives, self-selection, retention terms
Observed or acquired data Web, public records, sensors, logs, vendors Terms of use, personal data, provenance, regional restrictions, drift
New human or synthetic collection Rare cases and controlled coverage Sampling realism, annotation burden, privacy, distribution mismatch

The OECD’s Mapping relevant data collection mechanisms for AI training explains that each mechanism has different implications for developers, data subjects, and other rights holders (OECD, 2025). Google’s People + AI Guidebook recommends evaluating predictive power, relevance, fairness, privacy, and security when deciding whether to reuse or build a dataset (Google People + AI Guidebook).

3. Collect representative examples

Design a sampling plan instead of taking whatever is easiest to access. Stratify by variables that affect performance, such as language, region, device, age group, lighting, product version, or income band. Include both common and edge cases. If deployment traffic is seasonal, collect across the relevant period.

For web or document sources, save the source URL or record identifier, retrieval time, response status, parser version, and a checksum of the raw artifact. Keep raw data immutable; create transformed copies for cleaning and labeling.

A screenshot can be a useful training artifact for visual regression, document understanding, or page classification. Capture the same viewport and state deliberately, and record consent or access conditions. For dynamic pages, wait for the relevant selector or network activity rather than assuming the first HTML response is complete.

4. Label with a written scheme

Labels are the answers your model learns. Features are the observations used to infer those answers. Write a labeling guide before assigning work:

  1. Define every class or numeric target in plain language.
  2. Show positive, negative, borderline, and “cannot determine” examples.
  3. Specify precedence when multiple labels apply.
  4. Define an escalation path for ambiguous cases.
  5. Train labelers on a calibration set and revise confusing instructions.
  6. Measure agreement, adjudicate disagreements, and retain the original label plus the final decision.

Google notes that accurate labels are crucial for supervised learning and that instructions and interface design affect quality (Google People + AI Guidebook). Microsoft advises obtaining voluntary informed consent and using data only for purposes covered by documented consent (Microsoft Azure responsible AI guidance). Pay labelers fairly, provide support for sensitive content, and avoid making them infer private attributes that are not needed.

5. Automate collection with a reproducible manifest

Keep a manifest that maps each item to its source, timestamp, license or consent record, transformation code version, label status, and split. A simple JSON Lines record might look like this:

{"id":"img-00042","source":"https://example.org/item/42","collected_at":"2026-09-29T12:00:00Z","sha256":"…","label":null,"consent_ref":"project-terms-v3","split":null}

Example Python collector with retries, a rate limit, and a manifest:

import hashlib, json, time
from pathlib import Path
import requests

URLS = ["https://example.org/a", "https://example.org/b"]
out = Path("raw"); out.mkdir(exist_ok=True)
with requests.Session() as session, open("manifest.jsonl", "w") as mf:
    for i, url in enumerate(URLS):
        for attempt in range(3):
            try:
                r = session.get(url, timeout=30, headers={"User-Agent": "dataset-collector/1.0"})
                r.raise_for_status()
                body = r.content
                path = out / f"{i:06d}.bin"
                path.write_bytes(body)
                rec = {"id": path.stem, "source": url,
                       "status": r.status_code,
                       "sha256": hashlib.sha256(body).hexdigest()}
                mf.write(json.dumps(rec) + "\n")
                break
            except requests.RequestException:
                if attempt == 2: raise
                time.sleep(2 ** attempt)
        time.sleep(1)

Respect robots directives, published terms, access controls, rate limits, and applicable privacy law. Do not bypass authentication, bot protections, or consent choices.

6. Validate quality before training

Run checks on raw, labeled, and feature data. The UK Data and AI Ethics Framework names these dimensions: completeness, accuracy, validity, consistency, uniqueness, timeliness, duplicates, missingness, outliers, class balance, leakage, and subgroup coverage (UK Data and AI Ethics Framework).

  • Completeness: Required fields exist at expected rates.
  • Validity: Values match type, range, and format constraints.
  • Consistency: Related fields and labels do not contradict each other.
  • Uniqueness: Near-duplicates are identified before splitting.
  • Leakage: Features do not contain the target or information from after prediction time.
  • Coverage: Compare subgroup and edge-case counts with expected deployment traffic.

Generate a data-quality report for every version. Fail the pipeline when thresholds are exceeded, but retain quarantined records for review instead of silently deleting them.

7. Protect people, rights, and provenance

For personal or sensitive data, document the lawful basis, purpose, retention period, access roles, and deletion process. Apply minimisation: collect only fields needed for the stated decision. Use encryption in transit and at rest, separate identifiers from content, and prefer aggregation, masking, pseudonymisation, or de-identification where they preserve utility. The UK guidance recommends metadata, stewardship, transformation documentation, catalogs, access controls, audit logs, and continuous monitoring (UK data and AI guidance). NCSC lists filtering, sanitisation, differential privacy, masking, aggregation, swapping, and pseudonymisation as possible controls (NCSC guidance).

Keep provenance with the dataset: collector, supplier, location, time range, permissions, transformations, label policy, and known limitations. The EU AI Act says training, validation, and test sets should be relevant, sufficiently representative, as error-free and complete as possible for their intended purpose (EU AI Act, Recital 67).

8. Split and version without leakage

Create training, validation, and test sets according to the evaluation you need. For time-dependent predictions, use a chronological split. For users, patients, devices, or households, group by entity so related records cannot cross splits. Deduplicate before splitting. Freeze the test set, version raw and transformed data, and record code and schema versions. A data version should be reproducible from its manifest and transformation steps.

9. Monitor after deployment

Collection does not end at model release. Track missingness, new categories, label delays, distribution shift, subgroup metrics, and data drift. Sample predictions for human review and feed confirmed outcomes into the next labeled version. Set alerts for source failures, sudden volume changes, and quality-threshold breaches. Review whether the original purpose, consent, and retention terms still apply when data is reused.

10. Or skip the browser setup

When web pages are one of your sources, ScreenshotNeo provides a single GET request that returns a PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners like a visitor, removes more than 60 known consent platforms plus newsletter popups and chat widgets, and lets you turn each step off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and whether it was billed.

Cleaning page overlays before capture produces more consistent visual training examples.
Cleaning page overlays before capture produces more consistent visual training examples.

See the ScreenshotNeo API docs for all options. Basic calls:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

For dataset collection, relevant options include full-page capture with lazy images loaded, CSS-selector element capture, custom CSS and JavaScript, click actions, selector or network-idle waits, ad and tracker blocking, custom headers, cookies, user agents, authorization, timezone and geolocation, transparent backgrounds, resizing, a chosen cache TTL, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, and a usage API. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

There are 1,000 free shots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is included on every plan. Create a free ScreenshotNeo account to start collecting visual data.

11. Troubleshooting checklist

Symptom Likely cause Fix
Labels disagree frequently Ambiguous classes or weak examples Rewrite definitions, add borderline examples, calibrate, and adjudicate.
Excellent test score, poor production results Leakage or non-representative split Split by time or entity, deduplicate, and evaluate on deployment-like data.
Many missing fields Source or parser changed Validate schemas, quarantine failures, and alert on missingness thresholds.
Dataset is biased toward easy cases Convenience sampling Stratify collection and deliberately add underrepresented conditions.
HTTP 403 or CAPTCHA during capture Access policy or bot protection Obtain permission, authenticate legitimately, reduce rate, or use an approved source. Do not bypass controls.
Screenshot is blank or incomplete JavaScript not finished, blocked resource, or wrong viewport Wait for a selector or network idle, allow required resources, and capture the correct state.

12. Performance, reliability, and cost

Batch network requests within source limits, reuse connections, cache immutable artifacts, and process labeling asynchronously. Store compressed derivatives while retaining hashes of raw files. For high-volume collection, checkpoint manifests so a failed run resumes without duplicates. Measure cost per accepted, labeled example rather than cost per downloaded item; rejected, duplicate, or unusable records are part of the real collection cost. Review vendor and storage costs, annotation time, legal review, and maintenance when comparing sources.

For screenshot collection, caching with a deliberate TTL, bulk capture, and asynchronous jobs reduce repeated work. Keep the page-verdict and billed headers beside each artifact so accounting and quality audits can distinguish clean shots from failed loads or cache hits.

FAQ

How much data do I need?

No fixed threshold applies to every task. Justify adequacy with coverage, label quality, confidence intervals on evaluation, and performance across important subgroups and edge cases.

Should I buy a dataset or collect my own?

Reuse is faster, but inspect license, provenance, relevance, fairness, privacy, and freshness. Collect new data when existing sources cannot cover the intended operating range.

Can unlabeled data help?

Yes, for exploration or semi-supervised and self-supervised methods, but you still need a carefully labeled evaluation set to measure the decision you care about.

Store consent or license references in the manifest, including version, purpose, date, geography, retention terms, and withdrawal or deletion handling.

What should I do when labels change?

Version the labeling policy, preserve historical labels, relabel an evaluation sample, and report metrics under the policy used for each model.