ScreenshotNeo

BlogEngineering

Top Web Scraping Trends for E-Commerce in 2026

Retail scraping attacks remain high as AI crawlers and shopping agents focus on product discovery. Here’s what the 2025 data means for retailers in 2026.

By the ScreenshotNeo team29 September 202611 min read

Top Web Scraping Trends for E-Commerce in 2026

Retailers in 2026 face two related but distinct developments: persistent automated scraping pressure and a growing concentration of AI crawler and browser-agent activity around product discovery. The practical response is to improve visibility into what automated clients do, understand the sensitivity of the systems they touch, and apply proportionate controls instead of treating every bot as the same threat.

The strongest available figures describe activity observed during 2025 and published in 2026. They come from security vendors’ telemetry and a practitioner survey, not a universal census of all websites or all legitimate competitive research. They measure different populations and should not be combined into one estimate of “how much scraping” occurs.

1. Retail scraping pressure remains high, especially on product pages

HUMAN Security’s 2026 benchmark reports more than 150 billion attempted scraping attacks against retail and e-commerce businesses during 2025. It reports a 3.17% median scraping attack rate for the retail/e-commerce cohort. For heavily targeted businesses, the reported rate reached 57.01% of product-page traffic. These numbers describe different measures and cohorts: the high-target figure is not a typical store’s rate.

That distinction matters operationally. A high aggregate count says retailers see sustained automated pressure; it does not mean that every request is abusive, nor that every store receives the same volume. A store’s exposure depends on its traffic, product economics, public endpoints, defenses, and position in a competitive market. The vendor’s classifications and customer coverage also shape the measurement.

Product pages are a natural focus because they expose structured commercial information: descriptions, variants, availability, prices, and other details that can be monitored at scale. Search and category pages can reveal similar information while allowing automated clients to enumerate a catalog efficiently. Retailers should therefore examine which endpoints are being requested, at what rate, and in what sequence, rather than relying on a single site-wide bot percentage.

2. AI crawlers and browser agents are concentrating on commerce

HUMAN’s 2026 retail bulletin says 62.5% of AI crawler requests in its dataset went to retail and e-commerce in 2025. It also reports that 77% of AI agent/browser traffic to e-commerce websites visited product and search pages. A separate HUMAN measure says 46.6% of AI agent/browser traffic went to retail and e-commerce organizations.

AI crawlers and browser agents increasingly interact with commerce through product and search pages, while traffic purpose remains a separate classification question.
AI crawlers and browser agents increasingly interact with commerce through product and search pages, while traffic purpose remains a separate classification question.

Akamai’s 2026 research release reports that commerce represented 47.9% of AI bot traffic observed across its global network from July through December 2025. That is Akamai’s network view. Its denominator, traffic categories, time window, and classification approach differ from HUMAN’s, so the 47.9% and HUMAN percentages are not directly interchangeable.

Together, these findings point to a practical shift: product discovery is becoming a major interaction surface for automated clients. Some clients may index pages, answer shopping questions, compare products, or help a person navigate a catalog. Others may collect data at scale for purposes that create commercial, security, or contractual concerns. The traffic label alone does not tell a retailer which purpose applies.

3. Retailers have an intent-classification problem

“Bot” is a description of how a request is made, not a complete policy decision. Automated traffic can include search crawlers, AI crawlers, shopping agents, competitive monitors, accessibility or quality tools, and harmful activity. The same public product page may be useful to a customer-facing agent and valuable to a scraper trying to build a competing catalog.

That makes blanket allow/block decisions costly. Blocking too broadly can interfere with discovery or customer journeys; allowing too broadly can expose sensitive endpoints, increase infrastructure load, or support unwanted extraction. HUMAN and Akamai document substantial automated activity, while NRF/PwC’s retail work is framed around governance and security foundations for agentic commerce. The sources support the need to distinguish use cases, but they do not establish a universal classification system that every retailer should adopt.

A practical policy should consider several factors together:

  • Purpose and observed behavior: Does the client identify itself, respect access rules, and make requests consistent with its declared purpose? Identity claims alone are not proof.
  • Data sensitivity and business impact: Is the endpoint public product content, account information, inventory operations, pricing logic, or a workflow that could expose customer data?
  • API and product-page exposure: Which interfaces provide the data, and are they intended for public discovery, partner access, or authenticated use?
  • Classification accuracy: What is the likely false-positive cost of challenging or blocking a client? Can a policy be applied to a specific endpoint or behavior?
  • Operational cost: What are the costs of traffic, infrastructure, proxy use, fraud investigation, and manual review?
  • Applicable rules: What do site access policies, contracts, and jurisdictional requirements permit?

Policies should be documented with an owner, a reason, and a review path. When the classification is uncertain, a retailer can consider graduated responses such as rate limits, additional verification, narrower access, monitoring, or escalation rather than treating every case as an immediate permanent block.

4. API visibility and proportionate controls are rising priorities

Akamai reports that API attacks against commerce rose 9% year over year. In the API Security Impact Study summarized in its release, 85% of commerce respondents said they had experienced at least one API-related incident in the prior year, while 22% knew which APIs exposed sensitive data. These are Akamai-attributed findings from its study, not universal rates for every retailer.

Endpoint inventory and proportionate controls help teams respond according to data sensitivity and observed behavior.
Endpoint inventory and proportionate controls help teams respond according to data sensitivity and observed behavior.

The gap between incidents and visibility is a reason to inventory APIs before trying to classify all automated traffic. Teams need to know which APIs exist, who owns them, what data they return, how they are authenticated, and whether they were designed for direct public access. An undocumented endpoint can be a risk even if a retailer has a sophisticated bot product at the front door.

Akamai recommends moving beyond binary allow/block models toward risk-based governance that categorizes bots by intent and business value. In practice, that calls for coordination between security, fraud, API owners, e-commerce operations, and legal or privacy teams. A bot decision can affect product discovery, fraud exposure, customer experience, and data handling at once.

5. Scraping operations are getting more expensive, while AI adoption remains mixed

The 2026 State of Web Scraping summary from Apify and The Web Scraping Club surveyed hundreds of scraping professionals. In that community-recruited respondent pool, 65.8% reported increased proxy usage, 58.3% said proxy spending increased year over year, and more than 62% reported increased infrastructure spending. The summary attributes some cost pressure to stronger anti-bot protections. Treat these figures as a practitioner pulse, not a representative forecast for every scraping organization.

The same survey suggests that AI adoption is unsettled. Some 54.2% of respondents said they did not use AI in scraping workflows, while 66.2% planned to try AI-assisted scraping. Among respondents already using AI, 72.7% reported productivity advantages. These are separate survey responses, and “planned to try” does not mean adoption or proven gains across the industry.

For retailers, the cost trend matters because defenses can shift expenses to both sides: more verification and infrastructure for the site, and more proxies or automation work for clients attempting access. Controls should be evaluated against the business risk they reduce. A measure that creates substantial friction for customers or legitimate discovery needs a clear rationale and monitoring for unintended effects.

The European Data Protection Board published Guidelines 03/2026 on web scraping in the context of generative AI for feedback, with comments due 30 October 2026 according to the consultation page. At the time researched, these were draft consultation guidelines, not a final rule. The research available for this article establishes the draft status and consultation date, but does not analyze the linked draft’s detailed legal tests.

Retailers should not infer a specific legal conclusion from the title of a consultation document. They should track the consultation and relevant guidance in the jurisdictions where they operate, and get advice for decisions involving personal data, access restrictions, or commercial agreements. Technical bot classification and legal permission are related governance questions, but one does not answer the other.

7. A practical retailer checklist for 2026

  1. Inventory public and internal interfaces. List product, search, pricing, inventory, account, and partner APIs. Record owners, authentication, data returned, and intended users.
  2. Measure behavior by endpoint. Track request volume, rates, sequences, failures, and unusual enumeration patterns. Keep vendor classifications and your own observations distinguishable.
  3. Separate intent categories. Define how the organization handles search crawlers, AI crawlers, shopping agents, partner integrations, competitive monitoring, and suspected abuse. Make uncertain classifications reviewable.
  4. Set graduated responses. Choose actions appropriate to endpoint sensitivity and confidence: observe, rate-limit, challenge, restrict, or block. Define who can authorize exceptions and how quickly they are reviewed.
  5. Coordinate security and commerce. Include fraud, API, product, customer experience, and privacy stakeholders in policy decisions that affect catalog discovery or checkout-adjacent services.
  6. Track false positives and business impact. Monitor blocked or challenged legitimate traffic, customer friction, API incidents, operational load, and the cost of controls.
  7. Review access and legal rules. Align technical policy with published site rules, contracts, and applicable law. Revisit the policy as agent capabilities and guidance change.

8. Capture public pages for visual review

Automated monitoring does not replace checking what a human visitor sees. Teams may need screenshots of their own product pages or public pages for QA, competitive research conducted under applicable rules, and incident records. A browser automation setup gives control over viewport, timing, and page state, but it also requires browser installation, rendering configuration, and careful handling of failures.

DIY capture with Playwright in Node.js

This runnable example captures a full-page screenshot. Install Playwright, save the code as capture.mjs, and run it with Node.js. The browser binary may need to be installed with Playwright’s browser installation command for your environment.

import { chromium } from 'playwright';

const target = 'https://example.com';
const browser = await chromium.launch({ headless: true });
try {
  const page = await browser.newPage({
    viewport: { width: 1440, height: 1000 },
    deviceScaleFactor: 1
  });
  await page.goto(target, { waitUntil: 'networkidle', timeout: 60000 });
  await page.screenshot({ path: 'page.png', fullPage: true });
} finally {
  await browser.close();
}

For pages that keep long-lived network connections open, networkidle can wait too long or time out. Use domcontentloaded or load, then wait for a known selector or a short, bounded delay when the page needs client-side rendering. Avoid capturing authenticated or personal information into files without an appropriate access and retention policy.

For Python, install the Playwright package and its browser, then use this minimal equivalent:

from pathlib import Path
from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    try:
        page = browser.new_page(
            viewport={"width": 1440, "height": 1000},
            device_scale_factor=1
        )
        page.goto("https://example.com", wait_until="networkidle", timeout=60000)
        page.screenshot(path="page.png", full_page=True)
    finally:
        browser.close()

Choose full-page capture for a whole document, or a locator screenshot when only one component matters. Larger pages and higher device scale factors use more memory and produce larger files. If you need repeatable results, pin browser and dependency versions, use a consistent viewport, and set explicit timeouts.

Troubleshooting browser captures

  • Browser executable missing: Install the browser build matching the Playwright package, or use the browser cache expected by the runtime.
  • Navigation timeout: The page may have slow assets or never-idle connections. Use a less strict readiness condition and wait for the specific content you need.
  • Blank or partial output: Confirm the target URL is reachable from the runtime, wait for client-rendered content, and inspect console or navigation errors.
  • Different output between runs: Fix viewport, locale, timezone, fonts, and browser version where possible; dynamic content and experiments can still change a page.
  • Memory pressure: Avoid concurrent full-page captures of very long documents, lower device scale, and capture only the required element if appropriate.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. A GET request returns a PNG, JPEG, WebP, or PDF. Its clean-shot flow accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

For the documented request options, see the ScreenshotNeo docs. Here is a one-call example:

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));

The API also supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF settings, HTML/CSS rendering, custom CSS and JavaScript, click-before-capture, selector hiding and waiting, request blocking, headers, cookies, user agent, authorization, timezone, geolocation, transparency, resizing, configurable cache TTL, signed image links, async jobs with signed webhooks, bulk capture for up to 100 URLs per call, usage API, and OpenAPI spec. Parameter names used by other screenshot APIs also work to make switching easier.

Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed; an MCP server lets AI agents take screenshots; and 1,000 screenshots a month are free with no card, with paid plans starting at $5 for 3,000. Plans are Starter $5/3,000, Growth $15/15,000, Pro $39/60,000, Scale $99/250,000, and Business $249/1,000,000 per month; yearly billing gives two months free, and every feature is on every plan. Sign up for 1,000 free screenshots a month, with no card.

10. Frequently asked questions

Are these statistics estimates of all scraping on the web?

No. HUMAN and Akamai report activity from their own telemetry and classifications. They quantify observed automated traffic and attacks in their coverage, not a universal count of every benign and malicious scrape.

Does high AI bot traffic mean every AI agent is harmful?

No. Traffic volume and destination do not establish intent. Retailers need to assess behavior, endpoint sensitivity, and business impact before choosing controls.

Is the EDPB’s 2026 scraping guidance final?

No. The cited EDPB page described Guidelines 03/2026 as open for feedback through 30 October 2026. Treat it as draft consultation guidance at the time researched.

Do Apify survey results predict what every scraping team will do?

No. The findings are from hundreds of community-recruited scraping professionals and are best read as a practitioner pulse.