ScreenshotNeo

BlogGuides

Is Web Scraping Legal? Key Insights and Guidelines

Web scraping can be lawful, but public access is not a complete defense. Learn the CFAA, GDPR, robots.txt, contracts, and practical risk controls.

By the ScreenshotNeo team30 September 202610 min read

Is Web Scraping Legal? Key Insights and Guidelines

Short answer: Web scraping has no universal legal status. Collecting genuinely public, unauthenticated pages may fall outside the U.S. Computer Fraud and Abuse Act (CFAA) in some circumstances, including the Ninth Circuit’s reasoning in hiQ Labs v. LinkedIn. That does not create a general right to copy data. Contracts, copyright, database rights, trespass, privacy laws, state laws, technical barriers, server impact, and your downstream use can still create liability.

In the European Union, scraping personal data is generally processing under the GDPR. You need a lawful basis and must apply principles such as transparency, purpose limitation, minimization, accuracy, security, and storage limitation. Robots.txt is a crawler instruction, not an access authorization.

This guide gives you a practical way to assess a project before collecting data, explains the main U.S. and EU rules, and provides an operational checklist. It is general information, not jurisdiction-specific legal advice.

1. The facts that determine whether scraping is risky

Two projects can both copy HTML and have very different legal outcomes. Evaluate these facts together:

Question Lower-risk indication Higher-risk indication
How is the page accessed? Anyone can view it without an account. Login, paid subscription, invitation, or another restricted area is required.
Are technical controls present? You request ordinary public pages at a conservative rate. You bypass CAPTCHAs, IP blocks, paywalls, authentication, or access controls.
What data is collected? Non-personal product or publication information. Personal data, sensitive data, profiles, or inferred attributes.
What do the site rules say? An API or license expressly permits the intended use. Terms, registration conditions, or a cease-and-desist prohibit or limit collection.
How much load do you create? Low request volume, caching, and clear identification. High-volume parallel requests that degrade the service.
What happens afterward? Limited internal research with retention controls. Resale, public republication, profiling, or AI training at scale.

Document each answer. A written record of your purpose, sources, fields, rate limits, retention period, and decisions helps you make consistent choices and respond to questions later.

The CFAA distinction

In HIQ LABS, INC. V. LINKEDIN CORPORATION, No. 17-16783 (9th Cir. 2022), the Ninth Circuit held that hiQ had raised a serious question that the CFAA’s “without authorization” language does not apply when information is generally available to the public without authentication. The court affirmed a preliminary injunction. It did not finally resolve every claim or create a nationwide safe harbor. Read the Ninth Circuit opinion.

The practical distinction is between entering a protected computer area without permission and requesting pages that anyone may view. Facts still matter: using a login or fake account, evading a technical block, circumventing a control, creating excessive load, copying protected expression, and exploiting the results can change the analysis. Courts outside the Ninth Circuit may also apply different reasoning.

The U.S. Department of Justice’s CFAA policy says prosecutors will not charge “exceeding authorized access” solely because someone violated a contractual terms-of-service restriction on a generally available public website. The policy describes a narrower test involving separated areas, authorization to some areas but not others, knowledge, and enforcement goals. See the DOJ Justice Manual.

Other U.S. claims can remain

A CFAA argument can fail while another claim remains viable. Review:

  • Contract: Terms of service, API agreements, registration terms, and paid access conditions may create contractual duties.
  • Copyright: Facts are treated differently from original expression. Copying an entire article, image, database, or other protected work creates separate questions.
  • Database and state-law claims: Some jurisdictions recognize database rights or state-law theories that do not depend on the CFAA.
  • Trespass or interference: Excessive requests or deliberate service disruption can create risk even where pages are public.
  • Privacy and publicity laws: Personal information can trigger state and sector-specific obligations.

If an owner sends a cease-and-desist, stop the relevant collection while counsel reviews the request. Preserve your source list, timestamps, request logs, terms, robots.txt copy, and deletion actions. Do not respond by increasing traffic or attempting to evade a block.

3. Does robots.txt make scraping forbidden?

No. RFC 9309 standardizes the Robots Exclusion Protocol and expressly says its rules are “not a form of access authorization.” Robots.txt is a request to crawlers, not a statute or a permission grant. Read RFC 9309.

A lawful scraping workflow combines access checks, privacy review, and controlled retention.
A lawful scraping workflow combines access checks, privacy review, and controlled retention.

That does not make it irrelevant. Treat directives as an operational and evidentiary signal. Fetch and record the file, honor applicable Disallow rules, identify your crawler, rate-limit requests, cache responses, and stop when the owner blocks you or asks you to stop. Never use robots.txt as a reason to bypass authentication, a paywall, a CAPTCHA, an IP block, or another technical control.

A conservative robots.txt preflight in Python

This example checks a site’s crawler instructions before a planned request. It does not prove that a project is lawful or that a particular URL is permitted.

from urllib.parse import urlparse
from urllib.robotparser import RobotFileParser

TARGET = "https://example.com/articles/one"
USER_AGENT = "ResearchBot/1.0 (+https://example.org/contact)"

parsed = urlparse(TARGET)
robots_url = f"{parsed.scheme}://{parsed.netloc}/robots.txt"
parser = RobotFileParser(robots_url)
parser.read()

if not parser.can_fetch(USER_AGENT, TARGET):
    raise SystemExit(f"Robots policy disallows {TARGET}")

crawl_delay = parser.crawl_delay(USER_AGENT) or parser.crawl_delay("*")
print({"allowed_by_robots": True, "crawl_delay_seconds": crawl_delay})

Use a real contact URL and a rate limiter in production. Cache the result and retain the retrieval time so your decision can be reproduced.

4. Can you scrape personal data under the GDPR?

Usually, scraping personal data is GDPR processing. The European Commission defines personal data as information relating to an identified or identifiable living person, and processing includes collection, recording, organization, storage, retrieval, consultation, use, and disclosure. See the European Commission’s GDPR overview.

The European Data Protection Board stated on 8 July 2026 that GDPR applies when web scraping involves operations such as collecting, storing, organizing, and retrieving personal data. It highlights purpose limitation, transparency, reliable sources, timestamps, accuracy validation, and data minimization. If special-category data is involved, you need both an Article 6 lawful basis and an Article 9(2) condition; there is no blanket public-data exemption. Read the EDPB statement.

Apply the principles throughout the pipeline:

  • Lawfulness, fairness, and transparency: Identify and document a lawful basis. Consider how people will be informed, including whether an Article 14 exception genuinely applies.
  • Purpose limitation: Collect for a defined purpose and do not silently repurpose the dataset.
  • Data minimization: Exclude fields that are unnecessary. Do not collect sensitive attributes merely because they are visible.
  • Accuracy: Record source timestamps and provide correction processes where appropriate.
  • Storage limitation: Set deletion dates and remove stale records.
  • Integrity and confidentiality: Restrict access, encrypt where appropriate, and monitor exports.
  • Accountability: Keep a decision log, data map, retention schedule, and records of objections or deletion requests.

Assess controller and processor roles, international transfers, data-subject rights, security, and downstream recipients. For high-risk or large-scale projects, document necessity and proportionality and consider a data protection impact assessment.

5. Is scraping for AI training allowed?

There is no single global answer. Analyze the source’s access conditions, the data categories, your lawful basis, copyright and database rights, purpose, model-training method, retention, and whether outputs expose or reproduce personal information.

The EDPB adopted Guidelines 03/2026 on web scraping in the context of generative AI in July 2026. The consultation page stated that comments were open through 30 October 2026. Treat this as regulator guidance subject to consultation at the time of publication, not as a new statute. Check the EDPB consultation page.

For an AI dataset, record the exact sources, collection dates, filtering rules, exclusions, licenses, personal-data assessment, and deletion process. Reassess whether model training, retrieval, resale, or public release is compatible with the original purpose.

6. A practical pre-scrape checklist

  1. Define the purpose, jurisdictions, sources, fields, volume, and retention period.
  2. Classify every field as non-personal, personal, or special-category data.
  3. Read terms, API rules, registration requirements, copyright notices, database notices, and robots.txt.
  4. Confirm that collection does not require bypassing authentication or technical controls.
  5. Choose and document a privacy lawful basis; record necessity, proportionality, and transparency decisions.
  6. Use conservative rate limits, caching, source timestamps, and provenance.
  7. Filter sensitive fields and exclude sources that object to collection.
  8. Provide a contact and complaint process and honor cease-and-desist or opt-out signals.
  9. Control access to raw data, exports, and derived profiles.
  10. Reassess downstream uses such as resale, profiling, publication, and AI training.
  11. Obtain jurisdiction-specific advice for high-volume, sensitive, or cross-border work.

7. Safer collection engineering

Legal analysis and engineering controls reinforce each other. Use one identifiable user agent, a queue with bounded concurrency, exponential backoff for transient failures, response caching, and a maximum request budget per host. Keep raw HTML separate from extracted fields, encrypt sensitive stores, and attach source URL and retrieval time to every record.

Build stop controls: a robots change, owner request, elevated error rate, CAPTCHA, authentication wall, or unexpected personal-data field should pause the job for review. Do not rotate identities to defeat a block. Do not use fake accounts to reach content that is not public.

When a source provides an official API or export, compare its license, fields, limits, and retention terms with a scraper. An API may be the clearest permission path even when public HTML is technically reachable.

8. Or skip the browser setup

If your project needs screenshots of public pages for documentation, monitoring, or review, ScreenshotNeo provides a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. Before capture, it can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off.

Consent banners and overlays can be handled before a clean screenshot is captured.
Consent banners and overlays can be handled before a clean screenshot is captured.

Only clean shots are billed. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses identify the result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo documentation for authentication and options. A minimal call is:

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Relevant capture controls include full-page shots with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or a custom viewport, retina scale, PDF paper size, margins, landscape and page ranges, custom CSS and JavaScript, clicks, selector or network-idle waits, blocked ads or resource types, custom headers, cookies, user agents and Authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTL, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, and a usage API. Parameter names used by other screenshot APIs also work, which can simplify migration.

Plans include 1,000 free shots per month with no card; paid plans start at $5 for 3,000 shots. Yearly billing gives two months free, and every feature is included on every plan. A screenshot service does not decide whether your underlying collection is lawful: you still need permission, a lawful purpose, and controls appropriate to the page and data.

Sign up for 1,000 free screenshots a month with no card.

9. Troubleshooting common problems

Problem Likely cause Fix
The page is blocked by robots.txt. Your crawler is disallowed. Stop, document the directive, seek permission, or use an authorized API.
You receive a CAPTCHA or login page. The content is protected or the site detected automation. Do not bypass it. Request access or use a licensed source.
A cease-and-desist arrives. The owner objects or asserts contractual or other rights. Pause collection, preserve records, and obtain legal review.
The dataset contains unexpected names or profiles. Your fields or selectors are too broad. Stop the job, classify the data, delete unnecessary records, and reassess your lawful basis.
Requests cause errors or slow the site. Concurrency or rate is too high. Reduce concurrency, add backoff and caching, and set a host request budget.
A screenshot is blank or shows a popup. The page needs a wait, consent handling, or a selector. Use an appropriate wait or selector; with ScreenshotNeo, configure consent, popup, wait, and hide options and inspect verdict headers.

10. FAQ

Can I scrape LinkedIn profiles because they are public?

Public visibility can matter to a CFAA analysis, and the Ninth Circuit’s hiQ decision involved LinkedIn, but it was a preliminary, regional ruling. Terms, privacy law, technical barriers, copyright, and your use of the profiles still matter.

Does a public page require permission?

Not every public request requires individual permission, but public access is not a complete defense. Check contracts, privacy obligations, intellectual-property rights, robots directives, and server impact.

Is robots.txt legally binding?

RFC 9309 says robots.txt is not access authorization. Honor it as a responsible operational rule and never treat it as permission to access restricted areas.

Can I publish scraped data?

Publication adds downstream copyright, privacy, defamation, database, and consumer-protection questions. Reassess the project before release.

What should I retain for an audit?

Keep your purpose statement, source and field inventory, terms and robots snapshots, lawful-basis analysis, rate limits, timestamps, provenance, deletion records, objections, and incident decisions.

Use counsel for authentication or circumvention issues, personal or special-category data, large-scale or cross-border collection, profiling, AI training, resale, a cease-and-desist, or a high-impact publication.