ScreenshotNeo

BlogGuides

Compliance and Regulatory Web Scraping APIs

No scraping API makes collection compliant by itself. Use this workflow to assess data, source restrictions, vendor terms and operational controls.

By the ScreenshotNeo team29 September 202610 min read

Compliance and Regulatory Web Scraping APIs

Direct answer: No general-purpose web scraping API makes a project lawful or compliant by itself. An API can provide technical access and contractual commitments, but your team still needs to assess its purpose, target sources, data, rights, jurisdiction, downstream use and the provider’s terms. Publicly accessible pages are not blanket permission to collect and reuse everything on them.

This is a practical engineering and procurement framework, not legal advice for a particular collection. The strongest recent regulator material in the research is EU-focused and, in parts, concerns data collection for AI development. Rules vary with jurisdiction and facts. Start with the relevant primary guidance: CNIL’s web-scraping focus sheet and the EDPB’s July 2026 announcement on scraping for generative AI.

1. What does a compliance-oriented scraping API do?

A scraping API may fetch pages, render JavaScript, manage retries, or provide structured extraction. Separately, a vendor may publish a data processing agreement (DPA), acceptable-use policy, subprocessor information and data-location terms. These documents help you assess a provider; they do not answer whether your specific collection is permitted.

CNIL says scraping is not prohibited per se, but it must be assessed case by case and can conflict with other rules, including copyright, database rights and website terms. Where personal data is collected, a controller needs a lawful basis and must consider data-protection obligations. The right question is not simply “Which API is compliant?” but “Can we justify this collection, and does this provider’s contract and operation fit our obligations?”

A vendor’s DPA is a procurement input, not a permission slip. For example, ScrapingBee’s DPA places responsibility for the customer’s lawful basis and notices with the customer. Read the current ScrapingBee DPA for its scope and terms. Also review Oxylabs’ DPA, its acceptable-use policy, and Apify’s GDPR documentation as examples of provider materials to compare. This is not an endorsement; documentation and service scope can change.

2. A practical compliance workflow

Step 1: Write down purpose and scope

Before choosing an endpoint, record the business or research purpose, target domains and page types, fields to collect, collection frequency, retention period, downstream users, and any model-training or commercial reuse. Define what is out of scope. CNIL recommends setting collection criteria in advance and excluding data or sites that are not necessary.

Assess purpose, source signals, data and vendor terms before collection begins.
Assess purpose, source signals, data and vendor terms before collection begins.

Turn that scope into a collection specification developers can enforce. For example, allow only named domains and paths, limit fields to an explicit schema, cap requests per source, and set a retention deadline. A broad instruction such as “collect useful public web data” is difficult to audit or minimise.

Step 2: Identify personal and sensitive data

Determine whether pages may contain information relating to identifiable people. Consider whether the collection could encounter special-category data, information about minors, or sensitive information about people or groups. “Publicly available” does not mean “not personal data.” For EU processing, identify the lawful basis and safeguards that apply before collection, and document who is controller or processor for each activity.

The EDPB’s 2026 announcement says that where scraping involves personal-data processing, GDPR applies; for special-category data, both an Article 6 lawful basis and an Article 9(2) exception are needed. CNIL’s AI-development guidance calls for minimising collection, automatically excluding irrelevant sensitive data where appropriate, and deleting irrelevant data when identified. These points are grounded in the described EU context and should not be generalized into a single worldwide rule.

Step 3: Check source signals and rights

Review each source’s terms, robots.txt, CAPTCHA or other access barriers, copyright notices, and applicable database or text-and-data-mining reservations. Keep a record of the date and the signals you reviewed. CNIL says sites that clearly oppose AI-training scraping through exclusion protocols or CAPTCHA should be excluded in the context covered by its guidance.

Do not treat robots.txt as a complete legal answer in either direction. It is an important signal, but the OECD notes that the protocol may not be legally enforceable or technically binding in every circumstance, and a site’s terms and robots.txt may not match. See the OECD’s 2025 review of intellectual-property issues in AI trained on scraped data. A technical ability to fetch a page does not resolve copyright, contract, privacy or computer-access questions.

Step 4: Review provider terms for the exact service

For every candidate provider, verify the current DPA and acceptable-use rules for the endpoint and account you plan to use. Check:

  • Role and scope: Does the DPA cover the exact service, data and processing purpose?
  • Customer obligations: Who determines the lawful basis, gives notices, handles data-subject requests and performs impact assessments?
  • Data locations and transfers: Where are data processed and stored? Which subprocessors are involved, and what transfer terms apply?
  • Acceptable use: Are your target sources, personal-data categories, minors-related data or AI use restricted?
  • Retention and security: What deletion controls, access controls, breach support and audit evidence are described?
  • Reuse rights: Can results be cached, retained, displayed, used to train a model or passed downstream? Are attribution or other conditions required?

Do not infer that a DPA covers every product a vendor sells. Confirm the listed service, processing purpose, account type and geographic terms in the actual documents. Treat security statements as items to verify, not as substitutes for your own assessment.

Step 5: Check API-specific terms too

An API can impose use conditions independent of the underlying website’s rules. Microsoft’s cited Bing Search API terms, for example, restrict use of results, require attribution for LLM grounding, and prohibit using results for a site where a crawler is restricted, including through robots.txt. Read the current Microsoft Bing Search APIs with your LLM legal terms and confirm the exact product and scope before relying on an interpretation.

Step 6: Preserve evidence and revisit the decision

Keep a versioned record containing the purpose, source list, date checked, site signals, fields collected, lawful-basis assessment, provider contract version, safeguards, retention and deletion decisions. Reassess when the purpose changes, a new source is added, the provider changes terms, or the legal environment changes. The EDPB announcement recommends reliable sources, timestamps and validation in its AI-training context.

3. How to evaluate and integrate an API

Technical controls should make the approved scope enforceable. Use a source allowlist, rate limits, a strict output schema, bounded retries, and an automated retention or deletion policy. Keep access credentials out of source control. Log the source URL, timestamp, request outcome and policy version without unnecessarily duplicating personal data in logs.

A visual capture is a technical output, not a decision about access or reuse rights.
A visual capture is a technical output, not a decision about access or reuse rights.
  1. Choose one permitted source and a minimal set of fields.
  2. Confirm the provider’s terms and DPA cover the intended collection.
  3. Implement source restrictions and data minimisation before scaling.
  4. Test failure handling and deletion paths using non-sensitive sample data.
  5. Review a sample of collected records and confirm that the output matches the approved schema.
  6. Record the assessment and assign an owner to revisit it.

Do not confuse a screenshot with a structured scrape. A screenshot is a visual capture and does not itself establish permission to access a site, use its content, or collect personal data. If the task is visual evidence or page rendering, ScreenshotNeo is a website screenshot API and MCP server; its capture features do not replace the source and purpose review described above.

4. Common mistakes and fixes

Problem Why it fails Better approach
“The page is public, so collection is allowed.” Public access does not settle privacy, copyright, database rights, contract or API restrictions. Assess the source, purpose, data and jurisdiction together; document the decision.
“The vendor has a DPA, so our project is compliant.” A DPA describes a provider relationship and allocation of obligations; it does not create your lawful basis or rights to reuse content. Check the DPA scope and separately assess your controller duties and source rights.
“robots.txt is either irrelevant or the whole answer.” It is a meaningful site signal, but its legal effect depends on context and it may differ from site terms. Record it alongside terms, access barriers, rights notices and the applicable legal analysis.
“We will filter sensitive data after the whole corpus is collected.” Unnecessary collection can already create risk and undermine minimisation. Narrow sources and fields in advance; exclude irrelevant data early and delete it when identified.
“The provider’s general policy covers this endpoint.” Terms may apply only to listed products, regions or processing. Verify current documents against the exact service and account you will use.
“The data can be reused because the API returned it.” API results can have separate caching, attribution and downstream-use restrictions. Review API terms and source rights before storing, displaying, training on or redistributing results.

5. Performance, reliability and cost

Compliance controls affect system design. Allowlisting sources and minimising fields reduce unnecessary requests and storage. Rate limits and bounded retries make load predictable; retries should not turn access failures or blocks into an attempt to evade a source’s restrictions. Cache only when the provider’s terms, source conditions and retention policy allow it. Log enough to audit decisions while avoiding unnecessary personal-data copies.

Estimate total cost across API calls, storage, review, deletion, monitoring and engineering time. A low per-request price does not make an unsuitable provider a good fit if the service’s DPA, locations, subprocessors or acceptable-use policy do not match your requirements. Likewise, a provider contract that looks suitable on paper does not make a poorly scoped collection appropriate. No benchmark or universal cost comparison can replace checking current vendor pricing and terms for your workload.

6. Troubleshooting

The vendor will not confirm that my use is allowed

Pause procurement and provide a precise use description: sources, data categories, purpose, geography, retention and downstream use. Ask whether the named product and endpoint are covered by its terms. If the answer remains unclear, do not assume coverage.

Our crawler is blocked or sees a CAPTCHA

Treat this as a source signal requiring review. Check whether the site objects to automated collection and whether your project is allowed to proceed under its rules and applicable law. For the AI-training context in CNIL’s guidance, sites clearly opposing scraping through exclusion protocols or CAPTCHA should be excluded. Do not design retries to evade a restriction.

The DPA and product page describe different services

Compare the exact product names, processing purposes, data locations and account terms. Ask the provider to identify the applicable contractual documents. Do not assume a DPA for one service extends to another.

We discovered personal or sensitive data after collection

Limit access, follow your incident and data-protection procedures, assess whether collection was necessary and permitted, and apply the documented deletion or exclusion process. Update source and field filters so the same issue does not recur. Escalate to the appropriate privacy or legal owner where required.

Terms changed after launch

Keep dated copies or references to the terms used in your assessment, assign an owner to monitor relevant changes, and trigger a review when a vendor or source changes its conditions. Reassess whether continued collection and downstream use remain within scope.

7. Or skip the browser setup

For a screenshot task, ScreenshotNeo provides a one-request capture. It is not a web-scraping compliance service, and using it does not grant permission to access or reuse a target page. Review the ScreenshotNeo API documentation and the relevant site and provider terms.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

ScreenshotNeo removes cookie banners, newsletter popups and chat widgets before the shot; bot checks, blank pages and failed loads are never billed; its MCP server lets AI agents take screenshots; and 1,000 screenshots a month are free with no card, with paid plans starting at $5 for 3,000. These are product features, not a compliance determination. Sign up for 1,000 free screenshots a month, with no card required.

8. Short FAQ

There is no single answer for every source and purpose. CNIL says scraping is not prohibited per se, but legality depends on the circumstances and other applicable rules can restrict collection or reuse.

Does robots.txt make scraping illegal?

Not as a universal rule. It is an important signal to evaluate alongside terms, access barriers, rights and jurisdiction. Its legal effect can depend on the facts.

Can I scrape publicly available personal data under GDPR?

Public availability does not remove data-protection duties. Identify a lawful basis and safeguards; special-category data raises additional Article 9 requirements under the EU guidance described by the EDPB.

Does a scraping API provider make my collection compliant?

No. A provider’s DPA and policies are part of vendor diligence. Your organization still needs to assess its purpose, sources, data, legal basis, rights and downstream use.

What should I keep for an audit?

Keep the source and purpose assessment, fields and minimisation rules, dated source signals, provider terms and DPA version, safeguards, retention decisions and records of reassessment.

Conclusion

Choose a scraping API only after you have defined what you need to collect and checked the sources, data, reuse conditions and vendor contract. Put that decision into technical controls and preserve enough evidence to review it later. An API can make collection easier to operate; it cannot make the underlying collection compliant by itself.