ScreenshotNeo

BlogGuides

How to Collect Big Data from Online Sources

Learn how to plan, collect, validate, and maintain large datasets from online sources while respecting privacy, source rules, and site reliability.

By the ScreenshotNeo team30 September 202611 min read

How to Collect Big Data from Online Sources

To collect big data from online sources, begin with a defined purpose, a target schema, and an inventory of candidate sources. Prefer an official API, bulk download, or licensed feed when one meets the need. If you must scrape pages, keep the collection narrow, identify your crawler, limit requests, check site rules and applicable law, and record enough provenance to reproduce each result.

Large collection is not just a crawling problem. A technically successful crawl can still produce biased, stale, duplicated, legally restricted, or personally sensitive data. Plan permissions, quality checks, storage, retention, and deletion before scaling up.

1. Define the dataset before collecting it

Write down what decision, analysis, or product the data will support. Then define the smallest collection that can answer the question. “Collect every page about companies” is not a workable specification. “Collect public product prices for these categories from these named sources once per day, retaining the observed price and retrieval time” is closer.

Create a collection brief with these fields:

  • Purpose: the intended use, users, and output.
  • Scope: domains, page types, geographic or temporal coverage, exclusions, and collection frequency.
  • Schema: field names, types, units, allowed nulls, and stable identifiers.
  • Source authority: who publishes the information and whether it is primary, derivative, or user submitted.
  • Permission and constraints: API terms, license, robots.txt directions, contractual limits, privacy, copyright, and database rights questions.
  • Retention: how long raw and processed data are needed, who may access them, and how deletion requests are handled.
  • Quality gates: checks that must pass before data are used.

Keep the schema explicit. For example, a price observation might contain source_url, item_id, currency, amount, observed_at_utc, parser_version, and collection_run_id. Store the source and observation time with the value, rather than expecting someone to recover them later.

2. Choose an access method

Compare APIs, licensed feeds, bulk files, and page scraping before writing a crawler. An API usually provides structured fields and documented behavior; a bulk file can be more efficient for a complete historical snapshot; a licensed feed can make coverage and permitted reuse clearer. Scraping is useful when the information is only exposed as web content and no suitable authorized channel exists.

Method Good fit Check before choosing
Official API Repeated queries, structured records, incremental updates Quota, pagination, fields, versioning, terms, freshness
Bulk download Large snapshots or historical backfills Update cadence, file format, license, completeness
Licensed feed or agreement Production use where rights, coverage, and support matter Cost, allowed use, redistribution, corrections, retention
Web scraping Necessary content without a suitable data channel Terms, robots.txt, legal basis, server impact, fragility, privacy

Eurostat’s European Statistical System guidance describes APIs and scraping as web-content retrieval methods and encourages agreements and alternative channels such as APIs and file transfer. It also recommends minimizing server impact, identifying the bot, and following applicable scraping policies. See the ESS web-content retrieval guidelines.

Do not interpret public accessibility as blanket permission to reuse information. Robots.txt communicates crawler preferences, but does not settle every question about contractual terms, copyright, database rights, privacy law, or whether your intended use is permitted. Review the actual source terms and seek qualified advice for consequential or uncertain collection plans.

3. Handle personal data as regulated processing

Determine whether the collected records include information about identifiable people. Names, contact details, profiles, opinions, and combinations of attributes may be personal data depending on context. A page being viewable without login does not remove privacy obligations. The Canadian privacy commissioners state that publicly accessible personal information remains subject to privacy laws in most jurisdictions; their joint statement on data scraping and privacy explains the risks.

The EDPB states that GDPR applies to web scraping when it includes personal-data processing such as collection, storage, organization, and retrieval. Its guidance recommends reliable sources, recording timestamps, and validating data before use. See the EDPB announcement on web-scraping guidance. The exact requirements depend on jurisdiction, purpose, data type, and processing design.

Before collection, document the purpose and applicable lawful basis, assess transparency duties, minimize fields, protect credentials and raw data, restrict access, and set a retention period. Separate direct identifiers from analytical records where practical. Define a process for access, correction, deletion, and objections where the laws that apply require it. The European Commission’s GDPR principles guidance describes data protection by design and default, including collecting only what is necessary, limiting retention, and limiting access.

4. Build a restrained and identifiable collector

If scraping is justified and permitted, start with a small pilot. Check the site’s robots.txt and terms, identify yourself in a meaningful user-agent string with a contact route, and request only the pages and fields needed. Respect documented rate limits. Use delays, bounded concurrency, and backoff; do not try to evade access controls or bot checks. Coordinate with the site owner when volume or frequency could materially affect its service.

Here is a small Python example for a site you are authorized to collect from. It reads a fixed list of URLs, fetches sequentially with a delay, extracts a page title and description, and writes JSON Lines with source and retrieval provenance. Adapt the selectors and schema to the source’s documented structure. This is a teaching example, not a high-volume distributed crawler.

import json
import time
from datetime import datetime, timezone
from urllib.parse import urlparse

import requests
from bs4 import BeautifulSoup

URLS = [
    "https://example.org/authorized-page-1",
    "https://example.org/authorized-page-2",
]
USER_AGENT = "ExampleResearchBot/1.0 (+https://example.org/data-collection)"
OUTPUT = "records.jsonl"
DELAY_SECONDS = 2
PARSER_VERSION = "title-description-v1"

session = requests.Session()
session.headers.update({"User-Agent": USER_AGENT})

with open(OUTPUT, "w", encoding="utf-8") as out:
    for url in URLS:
        host = urlparse(url).hostname
        try:
            response = session.get(url, timeout=(5, 25))
            response.raise_for_status()
            content_type = response.headers.get("Content-Type", "")
            if "text/html" not in content_type.lower():
                raise ValueError(f"Unexpected content type: {content_type}")

            soup = BeautifulSoup(response.text, "html.parser")
            description = soup.find("meta", attrs={"name": "description"})
            record = {
                "source_url": response.url,
                "source_host": host,
                "retrieved_at_utc": datetime.now(timezone.utc).isoformat(),
                "http_status": response.status_code,
                "title": soup.title.get_text(" ", strip=True) if soup.title else None,
                "description": description.get("content", "").strip() if description else None,
                "parser_version": PARSER_VERSION,
            }
            out.write(json.dumps(record, ensure_ascii=False) + "\n")
        except (requests.RequestException, ValueError) as exc:
            error = {
                "source_url": url,
                "retrieved_at_utc": datetime.now(timezone.utc).isoformat(),
                "parser_version": PARSER_VERSION,
                "error_type": type(exc).__name__,
                "error": str(exc),
            }
            out.write(json.dumps(error, ensure_ascii=False) + "\n")
        time.sleep(DELAY_SECONDS)

Install the dependencies with python -m pip install requests beautifulsoup4, replace the example URLs only with authorized targets, and run python collect.py. A production system should distinguish retriable errors from permanent ones, enforce a per-host request budget, and avoid storing response bodies unless there is a documented need and an appropriate retention policy.

5. Make collection scalable without losing control

For large datasets, split discovery, fetching, parsing, validation, and storage into separate stages. Use a durable queue so a worker restart does not lose the collection plan. Partition work by source or date, and ensure a single host’s policy is enforced centrally across workers. A list of URLs can be deduplicated before enqueueing; canonicalization should be conservative because query parameters may change a page’s meaning.

A controlled pipeline separates source selection, collection, validation, and storage.
A controlled pipeline separates source selection, collection, validation, and storage.
  1. Discover: create a finite, auditable set of permitted URLs or API queries.
  2. Schedule: enforce per-domain concurrency, interval, retry budget, and collection window.
  3. Fetch: record status, headers needed for diagnostics, final URL, and retrieval time.
  4. Parse: transform into versioned, typed records; quarantine unexpected layouts.
  5. Validate: run schema, freshness, duplicate, range, and completeness checks.
  6. Publish: expose only records that pass quality and privacy gates.

Use idempotent writes keyed by a stable record identity plus observation time or source version. This makes retries safer. For APIs, persist pagination cursors and use incremental update markers where supported. For pages, use conditional requests such as ETag or Last-Modified only when the source supports them and the terms permit it. Cache responses when appropriate, but honor expiration and revalidation instructions.

Scale by measuring requests, bytes, parse failures, queue age, and per-source error rates. More workers do not automatically mean a better dataset: they can increase server load, amplify blocks, and make records arrive in an inconsistent order. Apply exponential backoff with jitter for transient failures, cap retries, and pause a source after repeated errors or a change in access policy.

6. Validate data and preserve provenance

Run quality checks before analysis and on every recurring batch. At minimum:

Validation and provenance checks help make collected records auditable and reproducible.
Validation and provenance checks help make collected records auditable and reproducible.
  • Validate required fields, data types, formats, units, and allowed value ranges.
  • Measure missing values and flag sudden changes in missingness.
  • Deduplicate exact repeats and investigate near-duplicates rather than blindly deleting them.
  • Detect outliers against domain rules; an unusual value may be a source error or a real event.
  • Check freshness against the promised collection schedule.
  • Compare samples with the source page or API response to catch parser drift.
  • Track coverage by source, date, category, and geography to expose selection bias.

CNIL’s guidance identifies correcting empty values and errors, detecting outliers, removing duplicates, and eliminating unnecessary fields as data-cleaning tasks. See CNIL guidance for developing AI systems. Cleaning rules should be versioned and documented: preserve raw observations where appropriate, and make transformations reproducible rather than silently overwriting source facts.

For each record or batch, retain source URL or query, retrieval timestamp in UTC, source response identifier where available, collector version, parser version, transformation version, and run ID. Maintain a manifest of input sources, scope, schema version, and quality results. Keep raw data access restricted and apply retention rules to both raw and derived copies.

7. Common failures and fixes

Symptom Likely cause Response
HTTP 403 or 429 Access is restricted, rate limit reached, or policy changed Stop retrying rapidly. Recheck terms and robots guidance, lower request rate, contact the owner, or use an authorized API/feed.
Timeouts or intermittent 5xx Temporary source issue, overloaded server, or aggressive concurrency Use bounded retries with backoff, lower concurrency, and record failed attempts separately.
Empty or malformed records Markup changed, content loads through scripts, or an unexpected response was parsed Check status and content type, compare a source sample, update parser tests and version, and quarantine suspect rows.
Duplicate rows Repeated discovery, pagination overlap, redirects, or retries Define a stable identity and idempotent upsert; inspect canonical URLs and page boundaries.
Stale or inconsistent values Sources update at different times, caching is misunderstood, or snapshots overlap Store observation time and source version, define a freshness window, and avoid implying simultaneity across sources.
Unexpected personal data Page content includes names, contact details, or user-generated material Pause the affected collection, reassess purpose and legal basis, minimize or remove the fields, and follow the documented privacy process.

8. Performance, reliability, and cost

Estimate volume before launch: number of records, pages per record, refresh frequency, average response size, retention duration, and expected change rate. The dominant cost may be engineering and maintenance rather than raw storage. Scrapers need ongoing attention when pages, access rules, or source behavior change. APIs and licensed feeds can reduce parser maintenance, but may have quotas or subscription fees.

Minimize work by selecting only necessary sources and fields, skipping irrelevant assets, using incremental updates, and deduplicating before fetching. Keep request rates low enough to avoid burdening the source. Treat a source’s availability and schema as external dependencies: monitor them, alert on quality regressions, preserve checkpoints, and keep a documented pause procedure.

Reliability is a data property as much as an infrastructure property. A pipeline that runs continuously but silently drops a category is less useful than a slower pipeline that detects the gap. Set explicit freshness and completeness thresholds, report them alongside downstream data, and do not mark a batch successful merely because its jobs completed.

9. Capture visual evidence when the source is a rendered page

Some collection workflows need a visual record of what a user-facing page displayed at a particular time, alongside structured fields. A screenshot is evidence of rendered appearance, not a substitute for permission, provenance, or a structured data record. Store its URL and timestamp with the extracted record, and apply the same privacy and retention controls to the image.

For a local browser-based workflow, capture only the page you are authorized to access, save the screenshot with the run identifier, and preserve the capture settings with the record. For pages whose content appears after rendering, wait for the relevant element or a reasonable load condition instead of assuming the initial HTML represents the visible page.

Or skip the browser setup

If your collection needs rendered screenshots as source evidence, ScreenshotNeo is a website screenshot API and MCP server. A single GET request returns an image or PDF. Its capture workflow removes cookie banners, newsletter popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents use the take_screenshot, get_page_info, and capture_pdf tools. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.org -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.org"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.org' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));

For authenticated collection or personal data, review permissions and handling before sending a URL to any external service. Sign up for ScreenshotNeo and get 1,000 free screenshots a month, with no card.

FAQ

Public access alone does not answer whether a particular collection and reuse are permitted. Rules depend on location, data type, source terms, intellectual property rights, privacy law, and purpose. Assess those issues for the specific project.

How often should I collect the same source?

Choose a cadence that matches the use case and source update behavior. Collecting more often than the data can meaningfully change wastes resources and adds server load. Document the freshness requirement and revisit it when the source or use changes.

Can I use scraped data for machine learning?

That depends on the collection’s permissions, rights, personal-data implications, and intended model use. Keep dataset provenance and transformations, and validate accuracy and representation before training.

What should I do if a website changes its layout?

Fail visibly: quarantine records that violate schema or quality checks, record the parser version, inspect representative source pages, then update and document the parser before resuming normal publication.

Practical launch checklist

  • Purpose, scope, schema, and retention are written down.
  • API, bulk, feed, and agreement options were considered first.
  • Source terms, robots directions, privacy, and rights questions were reviewed.
  • The collector identifies itself, limits requests, and has a pause mechanism.
  • Retries are bounded and data writes are idempotent.
  • Validation covers schema, missingness, duplicates, outliers, freshness, and source coverage.
  • Records retain timestamps, source references, parser versions, and transformation history.
  • Access, security, retention, and deletion processes cover raw and derived data.

A useful large dataset comes from a controlled process: select reliable sources, collect only what the purpose requires, validate the result, and preserve the history needed to explain how each record got there.