ScreenshotNeo

BlogHow-to

How to Detect and Handle Anti-Bot Measures in Browser Automation

Learn to tell anti-bot challenges from application and network failures, capture useful evidence, and handle authorized browser tests responsibly.

By the ScreenshotNeo team4 October 20269 min read

A browser automation run may reach a challenge page, an access-denied response, an unexpected redirect, or content that differs from the application you expected. Those symptoms can indicate that a site is challenging automation, but none proves which detector acted—or that automation was the cause. Diagnose the rendered page and network activity together, then use a permitted test path.

This guide is for developers and QA engineers testing systems they own or are explicitly authorized to test. If a third-party site blocks a run, stop automated retries and seek its published API, access policy, or support channel. Do not try to defeat the challenge or disguise or rotate automation identities.

1. What anti-bot signals can look like

Bot-management systems may combine request, session, browser, and behavioral signals. Cloudflare describes multiple detection engines because different bot types require different strategies; available engines depend on the plan. Its documentation describes heuristic checks, JavaScript detections, machine-learning scoring on applicable plans, and anomaly detection. These are examples of one vendor’s system, not a universal detector checklist. Cloudflare: Bot detection engines

Possible clues include:

  • A challenge, verification, or access-denied page instead of the expected application.
  • A redirect to a security or verification route, or a redirect loop.
  • A successful navigation but missing expected content or a changed page structure.
  • Requests for challenge scripts or other resources that fail, or a challenge that remains unresolved.
  • A difference between controlled runs that tracks a session, route, network condition, or browser configuration.

These clues are ambiguous. A challenge may fail for legitimate client-side reasons such as network problems, blocked scripts, disabled JavaScript, or browser configuration. A challenge is not proof of malicious intent or proof that a specific automation framework was identified. Cloudflare: JavaScript detections

2. Diagnose the page and the network

  1. Check what rendered. Assert a meaningful heading, landmark, or record that identifies the intended application state. Do not treat a completed navigation as proof that the right page loaded.
  2. Record the main document response and redirects. Note the final URL and status. A 403 or 503 is useful evidence, but status alone cannot identify the cause.
  3. Capture failed requests and console errors. Distinguish a transport failure from an HTTP error response. In Playwright, a 404 or 503 is still an HTTP response; requestfailed means the browser did not receive an HTTP response for that request. Playwright Request API
  4. Save the first failure’s evidence. Keep the timestamp, screenshot, final URL, response status, relevant request failures, and the browser and test configuration. Avoid blind retries that overwrite the initial evidence or add needless traffic.
  5. Compare controlled runs. Hold the browser version, test account, environment, and session setup steady. Check whether the symptom follows a route, account, network condition, or browser configuration. This narrows diagnosis; it does not reveal a vendor’s exact decision rule.

Runnable Playwright diagnostic (Node.js)

This example visits a system you control, logs the main document response, redirects and failed requests, saves a screenshot, and checks for an expected heading. Set TARGET_URL and EXPECTED_HEADING to values for your authorized test environment.

// diagnose.mjs
import { chromium } from 'playwright';

const target = process.env.TARGET_URL;
const expectedHeading = process.env.EXPECTED_HEADING;
if (!target || !expectedHeading) {
  throw new Error('Set TARGET_URL and EXPECTED_HEADING for an authorized test target.');
}

const browser = await chromium.launch({ headless: true });
const page = await browser.newPage();
page.on('response', response => {
  if (response.request().isNavigationRequest()) {
    console.log('navigation response', response.status(), response.url());
  }
});
page.on('requestfailed', request => {
  console.log('request failed', request.url(), request.failure()?.errorText);
});
page.on('console', message => {
  if (message.type() === 'error') console.log('console error', message.text());
});

try {
  const response = await page.goto(target, { waitUntil: 'domcontentloaded', timeout: 30000 });
  console.log('final URL:', page.url());
  console.log('main document status:', response?.status() ?? 'no response');
  await page.screenshot({ path: 'diagnostic.png', fullPage: true });

  const heading = page.getByRole('heading', { name: expectedHeading, exact: true });
  const expectedPageLoaded = await heading.isVisible().catch(() => false);
  console.log('expected heading visible:', expectedPageLoaded);
  if (!expectedPageLoaded) process.exitCode = 2;
} catch (error) {
  console.error('navigation or inspection error:', error);
  process.exitCode = 1;
} finally {
  await browser.close();
}

Run it with Playwright installed in your project: npm install playwright, then TARGET_URL='https://your-staging.example/' EXPECTED_HEADING='Account overview' node diagnose.mjs. Use a heading that identifies the expected page, not merely a generic heading such as “Error.” This script observes symptoms; it does not attempt to solve or bypass challenges.

Python equivalent with Playwright

# diagnose.py
import asyncio
import os
from playwright.async_api import async_playwright

async def main():
    target = os.environ.get("TARGET_URL")
    expected_heading = os.environ.get("EXPECTED_HEADING")
    if not target or not expected_heading:
        raise RuntimeError("Set TARGET_URL and EXPECTED_HEADING for an authorized test target.")

    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        page = await browser.new_page()
        page.on("response", lambda response: print(
            "navigation response", response.status, response.url
        ) if response.request.is_navigation_request() else None)
        page.on("requestfailed", lambda request: print(
            "request failed", request.url, request.failure
        ))
        try:
            response = await page.goto(target, wait_until="domcontentloaded", timeout=30000)
            print("final URL:", page.url)
            print("main document status:", response.status if response else "no response")
            await page.screenshot(path="diagnostic.png", full_page=True)
            heading = page.get_by_role("heading", name=expected_heading, exact=True)
            loaded = await heading.is_visible()
            print("expected heading visible:", loaded)
            if not loaded:
                raise RuntimeError("Expected page landmark is missing; inspect the saved evidence.")
        finally:
            await browser.close()

asyncio.run(main())

Install with pip install playwright and playwright install chromium. Run using TARGET_URL='https://your-staging.example/' EXPECTED_HEADING='Account overview' python diagnose.py.

cURL as a transport-level check

cURL can record response headers and redirects, but it does not execute the page’s JavaScript or validate browser-rendered content. Use it as one diagnostic signal, not a substitute for the browser checks above. Run only against a system you may access.

curl --silent --show-error --location --max-time 30 \
  --dump-header response-headers.txt \
  --output response-body.html \
  --write-out 'final_status=%{http_code}\nfinal_url=%{url_effective}\nredirects=%{num_redirects}\n' \
  'https://your-staging.example/'

3. Choose a permitted handling path

Your application or authorized test environment

  • Prefer a staging environment or test configuration designed for automation. Keep test accounts and browser storage isolated so a session from one test does not contaminate another.
  • Verify user-visible application behavior with stable page landmarks and records, not only navigation completion.
  • Mock third-party services that are not part of the system under test. Playwright supports routing and fulfilling requests so a test can use a known response instead of depending on an external service. Playwright: Best practices and Playwright: Mock APIs
  • If your organization operates the bot-management layer, coordinate with its security or site reliability owner. Agree on a documented, narrowly scoped test path or policy and review the relevant logs. Do not assume a particular configuration is available on every plan.
  • Keep a repeatable run configuration: browser version, account, storage state, environment, and test data. When a failure occurs, preserve the first response and screenshot before investigating.

A third-party site

If an external site challenges or blocks the run, stop automated retries. Check its published API and access policy, or contact its support channel for an approved route. Do not disguise automation, defeat the challenge, or rotate identities to get around the site’s controls. For a test that depends on external data, mock that dependency or use an authorized API instead. Playwright likewise recommends testing only dependencies you control. Playwright: Avoid testing third-party dependencies

4. Common problems and fixes

Symptom What it can mean Responsible next step
Navigation succeeds but the expected heading is absent The browser may have reached an intermediate, challenge, error, or unexpected application state. Inspect the final URL, main document status, screenshot, and rendered content. Make the assertion identify the real expected page.
requestfailed appears No HTTP response was received for that request; possible causes include network, DNS, TLS, or client cancellation issues. Log the failure text and timestamp, then check the environment and whether the failure is repeatable. Do not classify it as a bot block based on this event alone.
A request returns 403, 429, or 503 The server returned an HTTP response. It may reflect access policy, rate limiting, service trouble, or another application-specific condition. Record the response and context. If you control the system, ask its operator to inspect logs; if not, stop and seek an approved access route.
A challenge loops or remains incomplete Challenge scripts may not run or validate. Network problems, blocked scripts, disabled JavaScript, and browser configuration are among possible causes. For your own test setup, check browser console and failed challenge-resource requests, then consult the detector owner. Do not automate challenge solving on a production third-party site. Cloudflare: Challenge solve issues
Results differ between runs Session state, account, route, network, browser version, or test data may have changed. Stabilize those inputs, isolate storage state, and compare logs from controlled runs.
Network routing misses requests A service worker or another interception layer may affect which events Playwright exposes. Check the test’s service-worker and routing setup; Playwright documents blocking service workers as one option when native route events are missing. Playwright: Network events and service workers

5. Reliability, performance, and cost

Browser checks are more reliable when they test a controlled environment, use stable landmarks, and isolate accounts, storage, and test data. Third-party dependencies can change independently and introduce overlays, downtime, or access challenges, so mock them when they are outside the test’s purpose. Keep diagnostic output small and useful: the main response, final URL, relevant failures, console errors, timestamp, and a screenshot are usually more actionable than repeatedly dumping every request.

Retries add runtime and request volume, and can obscure the first failure. Retry only when your own system’s documented policy says the failure is transient, and use bounded retries with a delay; do not use retries to press through a challenge or access denial. No universal retry count or timing applies across services.

6. Or skip the browser setup

For authorized screenshot capture, ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF. It is useful when the task is to capture a page without maintaining browser setup; it is not a way to bypass a site’s access controls. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners are accepted like a visitor and removed before the shot, along with 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. An MCP server lets AI agents using Claude, Cursor, or any MCP client take screenshots. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots.

Sign up for 1,000 free screenshots a month, with no card required.

7. Frequently asked questions

Does a challenge prove the site detected Playwright?

No. A challenge is a clue that access is being evaluated, but it does not identify the exact signal or prove that a particular framework caused it.

Can a 200 response still be the wrong page?

Yes. A successful HTTP response can contain an intermediate page or unexpected content. Check rendered landmarks and final URL as well as the status.

Should I use a real third-party site as a test fixture?

Usually not when the test does not control that site. Use its approved API or mock the dependency so the test has a predictable response.

Is there one universal anti-bot detector?

No. Vendors can combine different request and browser signals, and their systems and configuration vary. Diagnose observable behavior rather than guessing at an unseen rule.