Web Scraping Data Protection and Privacy Best Practices
Publicly visible information can still be personal data. Use this practical workflow to assess a scraping project, limit collection, and protect data.

Publicly accessible does not mean free of privacy obligations. If a page contains information about an identifiable person, collecting, storing, organizing, or retrieving it may be regulated personal-data processing. Before scraping, define the purpose, assess the applicable law and source-site rules, collect only what is needed, and protect the data through its full lifecycle. No checklist can determine whether a particular project is lawful without its facts and jurisdictions.
This guide is for developers, product teams, and reviewers planning or operating web-scraping projects. It covers privacy and data protection; separate legal review may also be needed for copyright, database rights, contract, computer-misuse rules, sector-specific requirements, and international transfers.
1. Is scraping public data legal?
There is no universal yes or no. Public visibility alone does not remove personal information from data-protection laws. A joint statement by privacy regulators says publicly accessible personal information is subject to privacy and data-protection laws in most jurisdictions. Whether a particular collection is permitted depends on the data, purpose, collection method, jurisdiction, your role, and how the information will be used.
For the EU GDPR, the European Data Protection Board states that the GDPR applies to web scraping when it includes personal-data processing such as collection, storage, organization, and retrieval. Its July 2026 statement focuses on scraping for generative-AI development; it is useful guidance, not a complete rulebook for every purpose or country.
Permission from a website can be relevant, but it does not settle the privacy analysis. Regulators caution that contractual authorization can be a safeguard but cannot, by itself, make processing lawful. You may still need a lawful basis, transparency, consent where required, and oversight of contractual limits.
Questions to answer before deciding
- What specific purpose will the scraped information serve, and who will use it?
- Which fields are necessary, and can the purpose be achieved with fewer fields or aggregated information?
- Could fields identify people directly or indirectly when combined? Could the project infer sensitive traits?
- Which countries’ laws apply to your organization, the people, and the processing?
- What do the source site’s terms, access policies, and technical signals say about collection and reuse?
- Will data be shared with vendors, customers, or AI systems, and how long will it remain available?
Do not treat robots.txt compliance, a public page, an API key, or a site owner’s permission as a complete legal conclusion. They answer different questions and may be relevant safeguards.
2. Does GDPR apply to web scraping?
It may apply when the project processes personal data within the GDPR’s scope. Personal data is information relating to an identified or identifiable person. Names and email addresses are obvious examples, but combinations of location, job title, profile details, identifiers, or other attributes can also identify someone. Consider what can be learned by linking fields to other available information.
Under GDPR, identify and document an applicable Article 6 lawful basis and apply the principles of purpose limitation, transparency, data minimization, and accuracy. If special-category data is processed, the EDPB says an Article 6 basis and an Article 9(2) exception are both needed. Design the collection to avoid capturing such information incidentally where feasible; assess safeguards for anything that remains in scope.
Transparency and rights handling depend on the facts and applicable law. Plan how people can raise concerns and how you will evaluate requests to correct, suppress, delete, or otherwise address data. Do not promise a universal response without checking the relevant jurisdiction and any applicable exceptions.
For AI training, the EDPB recommends reliable sources, recording timestamps, and validating data before use to support accuracy. AI use also makes it especially important to define the purpose, screen for sensitive information, and avoid collecting a broad corpus first and deciding on uses later.
3. A practical privacy workflow for scraping
Step 1: Write down the purpose and limits
Describe the concrete outcome, intended users, and downstream uses. State what is out of scope. “Build a dataset” or “find useful data later” is not a useful limit. If the purpose changes, review whether the new use is compatible with the original one and whether additional notice or legal review is needed.

Step 2: Map fields and identify personal or sensitive data
Make a field inventory before implementation. For each field, record why it is needed, whether it can identify a person alone or in combination, how sensitive it may be, and whether a less detailed alternative works. Include free-text content and derived inferences in the assessment. Avoid collecting sensitive personally identifying information without a legitimate need; the FTC’s business guidance puts the minimization principle plainly: “If you don’t have a legitimate business need for sensitive personally identifying information, don’t keep it. In fact, don’t even collect it.”
Step 3: Review jurisdictions, roles, and source access rules
Identify the organization responsible for the project, any parties acting on its behalf, the jurisdictions connected to the people and processing, and restrictions on collection or reuse. Review the site’s terms and access policy. Eurostat’s guidance suggests contacting site operators in advance about access and issues such as privacy and database protection. That is practical operational guidance, not a universal legal rule.
For EU/EEA personal-data processing, document the lawful-basis analysis and GDPR principles. If special-category information may be encountered, assess both Article 6 and Article 9(2), and add feasible filters, exclusions, or review controls. Escalate uncertainty before collection rather than assuming that public visibility or a site’s authorization resolves it.
Step 4: Choose a collection route
| Route | Scope and permission | Auditability and impact | Questions to resolve |
|---|---|---|---|
| Direct scraping under site rules | Depends on the site’s policies, applicable law, and the project’s purpose. | Requires you to manage pacing, identify the crawler where appropriate, and keep your own records. | Are the fields necessary? Are access rules clear? Can the source handle the request pattern? |
| Site-provided API or authorized feed | May define permitted fields, uses, and access controls. | Can give the platform more control and facilitate logging and monitoring; still needs responsible downstream use. | What do the scope, retention, and reuse terms allow? What is logged and for how long? |
| Licensed or otherwise lawfully sourced dataset | Check the license, provenance, and permitted purposes. | May provide documented provenance, but your handling and later use still require review. | Is it current and accurate? Does the license cover your intended use and affected data? |
No route is automatically lawful or best. Compare documented permission and scope, ability to limit fields and purposes, freshness and accuracy, auditability, burden on the source’s infrastructure, and ongoing cost. An API is not impenetrable and does not automatically make downstream processing lawful.
Step 5: Make collection restrained and identifiable
Collect only the fields needed for the stated purpose. Use a clear crawler identity in the user-agent where appropriate. Follow current site directions, use controlled request rates, and pause between requests to avoid overloading the service. Eurostat gives a one-second pause as an example, not a universal rate limit. Follow the source’s instructions and tune the request pattern to the operational context.
Respect robots exclusion directives and terms as part of responsible access. A robots.txt file is an operational signal; it does not by itself decide privacy, copyright, contract, or database-rights questions. Avoid bypassing access controls or collecting from reserved areas without appropriate authorization.
Step 6: Protect, review, and dispose of data
- Inventory storage and flows. Record what was collected, where it is stored, who can access it, and which vendors or services process it.
- Restrict access. Give access only to people and services that need it. Protect retained data in proportion to its sensitivity.
- Set vendor expectations. Document security expectations for service providers and check that they meet them. The FTC recommends written expectations and verification.
- Set retention periods. Tie retention to the stated need and any applicable legal duties. Define deletion or secure disposal when the need ends.
- Maintain review paths. Decide how to assess correction, suppression, deletion, and source concerns under the applicable rules. Record the decision and its rationale.
4. A minimal implementation pattern
The following Python sketch demonstrates a restrained collection loop: explicit fields, a descriptive user-agent, a delay, and a place to validate before storing. It is illustrative, not a legal-compliance mechanism. Replace the example domain and parser with a source you are authorized to access, and check the site’s current terms and robots directives before running it.
import time
import requests
from urllib.parse import urljoin
BASE = "https://example.org"
HEADERS = {"User-Agent": "ExampleResearchBot/1.0 (contact: privacy@example.org)"}
PAUSE_SECONDS = 1.0
# Keep only fields justified by the documented project purpose.
def parse_needed_fields(html):
# Replace with a parser for the authorized source.
return []
def collect(urls):
rows = []
for url in urls:
response = requests.get(url, headers=HEADERS, timeout=20)
response.raise_for_status()
for row in parse_needed_fields(response.text):
# Validate, minimize, and screen before persistence.
if set(row) & {"unneeded_field", "sensitive_field"}:
continue
rows.append(row)
time.sleep(PAUSE_SECONDS)
return rows
if __name__ == "__main__":
pages = [urljoin(BASE, "/public-page")]
records = collect(pages)
print(f"Collected {len(records)} reviewed records")
Before production, add an approved source list, robots and policy review, retry limits with backoff, request logging that avoids unnecessary personal data, schema validation, secure storage, retention enforcement, and an incident process. A pause alone does not make access appropriate.
5. How can a website prevent data scraping?
No single control prevents all scraping. Privacy regulators recommend a regularly reviewed mix suited to the data, threat, technical context, proportionality, and cost. Options include rate limits, monitoring unusual activity, bot detection, access controls, terms, reserved areas, APIs, and incident response. The Italian authority described these as measures controllers should assess; it explicitly did not make each one mandatory in itself.

- Reduce exposure: review whether personal information needs to be publicly visible, and avoid exposing unnecessary fields.
- Control access: use authentication or reserved areas where appropriate, and provide a scoped API for legitimate uses.
- Watch traffic: monitor unusual request volumes and patterns, with proportionate bot detection and blocking.
- Set and enforce terms: define permitted information and purposes when authorizing collection, monitor compliance, and respond to misuse.
- Prepare response: document escalation, investigation, and mitigation steps for suspected scraping.
APIs can facilitate controls and logging, but they are not impenetrable. A contractual clause telling users to obey the law is not sufficient by itself; authorization still needs a lawful basis and oversight.
6. Or skip the browser setup
If the task is to capture a page for review or documentation, a screenshot is a different workflow from scraping structured personal data. ScreenshotNeo is a website screenshot API and MCP server by Yorker Media. One GET request returns a PNG, JPEG, WebP, or PDF; see the API documentation for options and authentication.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.org \
-o shot.webp
Cookie banners are accepted and removed before the shot, along with known newsletter popups and chat widgets. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. A screenshot does not replace privacy review if the captured page contains personal data.
Sign up free for 1,000 screenshots a month, with no card.
7. Troubleshooting common project failures
| Problem | Likely cause | Practical response |
|---|---|---|
| The team cannot explain why a field is collected. | Collection began before purpose and field mapping were defined. | Pause that field, document necessity and downstream use, then reassess whether it should be collected at all. |
| A page is public, so reviewers assume no privacy rules apply. | Visibility is being confused with an exemption. | Check whether information relates to an identifiable person and assess the applicable jurisdiction and purpose. |
| robots.txt permits a path, so the project assumes reuse is allowed. | An operational crawler directive is treated as full legal permission. | Review terms, privacy obligations, and other applicable rights separately; robots.txt does not answer all of them. |
| The site authorized access, so the project skips legal-basis review. | Contractual permission is treated as sufficient on its own. | Assess lawful basis, transparency, consent where required, and contractual limits independently. |
| Collection overloads or disrupts the source. | Request rate is too high or ignores site-specific directions. | Stop or slow the crawler, identify it where appropriate, respect current guidance, and use bounded retries. |
| Data remains in vendor systems after the project ends. | Retention and data flows were not mapped. | Inventory copies and processors, enforce the retention schedule, and dispose of data when no longer needed, subject to legal duties. |
| AI training data is stale or difficult to validate. | Sources and collection times were not recorded. | Prefer reliable sources, timestamp records, and validate data before training. |
8. Performance, reliability, and cost considerations
Request pacing is a reliability and courtesy control: aggressive concurrency can burden a source and trigger blocking, while a fixed delay may be too slow or too fast for a particular service. Follow site-specific limits, cap concurrency, bound retries, and use backoff for transient failures. Log enough to investigate errors and demonstrate the collection process, while avoiding logs that create an unnecessary second store of personal data.
Data quality affects downstream cost. Validate records near ingestion, track provenance and timestamps when freshness matters, and reject malformed or out-of-scope fields before persistence. Scope creep raises storage, review, security, and deletion work; minimization reduces those operational burdens as well as privacy exposure.
Compare route costs across engineering and review time, source fees or licensing, storage, vendor controls, and ongoing data maintenance. A free or technically accessible source can still have meaningful infrastructure and compliance costs. There are no universal request-rate, retention, or cost figures: set them from the source’s rules, project purpose, sensitivity, and applicable obligations.
9. FAQ
Can I scrape personal data from public websites?
Sometimes, but public availability alone does not answer whether collection and reuse are permitted. Identify the applicable rules, purpose, fields, source constraints, and safeguards before collecting.
Does respecting robots.txt make scraping legal?
No single robots directive resolves privacy, contract, copyright, database-rights, or computer-misuse questions. Treat it as one operational signal and review the rest separately.
Does a site-provided API make the data safe to use?
An API can define scope and improve logging and control. It does not automatically authorize every downstream use or remove privacy obligations.
What should I do if a source complains?
Pause the affected collection while you review the source’s concern, your authorization, the data involved, and applicable obligations. Preserve only records needed for a proportionate review, then document corrective action.
Sources and scope
This is general operational guidance, not legal advice or a determination that a particular scraping project is lawful. The cited regulator and public-sector material provides the factual basis; requirements differ by jurisdiction and project. Review the primary sources and obtain qualified advice for decisions with legal consequences.
- European Data Protection Board: web scraping for generative AI.
- Privacy regulator joint statement on data scraping and privacy.
- Eurostat: web scraping guidance.
- Federal Trade Commission: Protecting Personal Information, a Guide for Business.
- Italian data protection authority: guidance and measures concerning web scraping.


