ScreenshotNeo

BlogGuides

What Is a Web Bot? Types, Uses, and Examples

A web bot automates tasks over the internet. Learn what bots do, how to tell useful automation from abuse, and what website owners can control.

By the ScreenshotNeo team4 October 202611 min read

A web bot is software that performs tasks automatically over the internet. Bots can crawl pages for search, answer customer questions, monitor a service, compare prices, or collect data. Automation alone does not make a bot harmful: its purpose, authorization, behavior, and effect on a site determine whether it is useful, unwanted, or malicious.

For developers, the practical question is what the automation does and how it interacts with your site. A search crawler fetching pages at a reasonable rate is different from a scraper ignoring access rules, a bot buying tickets for resale, or a botnet generating traffic to disrupt a service.

What is a web bot?

A bot is software programmed to perform tasks. On the web, it may make requests to websites, interact with visitors through chat, or scan and process page content. Bots may run on a schedule, respond to an event, or act on instructions from a person or another system. [Cloudflare’s bot documentation]

“Web bot” is a broad term, not a single technology or a verdict about intent. The same general techniques—automated requests, browser interaction, and data processing—can support useful services or harmful activity. Assess the behavior and its impact rather than assuming every automated visitor is either legitimate or abusive.

What are bots used for?

Common bots automate tasks that would otherwise require people to repeatedly browse, answer, compare, or check information.

  • Search and discovery: Crawlers fetch pages and help build search indexes.
  • Customer support: Chatbots answer common questions, guide visitors through a menu, or route requests.
  • Commerce: Shopping bots compare products or prices; transaction bots can automate permitted steps in a purchase flow.
  • Monitoring: Bots check availability, performance, traffic, or other system conditions and can alert operators to changes.
  • Data collection: Scrapers gather specified information. Whether this is appropriate depends on authorization, site rules, the data involved, and request behavior.
  • Security operations: Automated systems can inspect activity or perform defensive tasks. Automation can also be misused for credential stuffing, spam, or denial-of-service activity.

These categories describe tasks, not an automatic permission status. A crawler may help a site appear in search, while a high-volume scraper may strain its infrastructure. A monitoring bot operated by the site owner is different from an unknown client probing the same endpoints.

Types of web bots

Type Typical task What to consider
Web crawler or spider Discovers, fetches, and indexes pages Which pages it requests, how often, and whether the crawler is genuine
Chatbot Responds to visitor questions or guides a conversation Task complexity, accuracy limits, escalation path, and disclosure that it is automated
Scraper Collects selected page data Authorization, data sensitivity, request rate, and site impact
Shopping bot Compares products, availability, or prices Whether its access is permitted and whether it creates excessive load
Monitoring bot Checks a service or traffic condition Expected request pattern, alerting, and access scope
Transaction bot Automates a transaction or workflow User authorization, safeguards, and effects on other users
Abusive bot or botnet Spams, hoards inventory, steals credentials, or generates harmful traffic Misuse, scale, evasion, and damage to users or infrastructure

A botnet is a network of compromised bots controlled remotely. It is not a synonym for a crawler or for ordinary automation. Botnets can be used for activities including spam, identity theft, denial-of-service attacks, and click fraud. [IETF RFC 6561]

What are examples of bots?

Search crawlers

Search engines use crawlers to discover and fetch web pages. Google uses “Googlebot” as the generic name for its smartphone and desktop crawlers. Google says it primarily discovers URLs through links on previously crawled pages. These details describe Google’s crawlers, not every crawler on the web. [Google Search Central: Googlebot]

Customer support chatbots

A support bot might offer predefined menu choices, recognize keywords, or use more advanced natural-language processing. Some systems combine approaches. The right design depends on the job: a short menu may work for routing, while open-ended questions need more flexible handling and a clear way to reach a person. GOV.UK recommends making a chatbot’s automated nature and limits clear, and considering whether better site content or search would serve users more effectively. [GOV.UK chatbot and webchat guidance]

Shopping and transaction bots

A shopping bot can compare listed prices or availability. A transaction bot may automate a user-authorized workflow. The same automation can become harmful when it bypasses limits, hoards scarce inventory, or buys tickets for resale. The task and its effect on other users matter.

Monitoring and scraping bots

A monitoring bot can periodically check whether a site or service responds. A scraper extracts selected content for a defined purpose. Operators should consider the source’s access rules, the sensitivity of the data, request volume, and whether the activity affects site performance.

Spam, credential abuse, and denial-of-service bots

Harmful bots may create fake accounts, harvest email addresses, attempt credential stuffing, or generate enough requests to impair a service. A coordinated botnet can amplify such activity. Defenses should focus on observed behavior and impact; a user-agent string by itself is not proof of identity or intent.

Are all bots bad?

No. Many bots provide useful services, including search crawling, support, and monitoring. Other bots are unwanted because they disregard site rules, collect data without authorization, burden infrastructure, or harm users. Whether a bot is acceptable depends on factors such as:

  • Purpose: What task is it carrying out?
  • Authorization: Is the activity allowed by the site owner and relevant users?
  • Behavior: Does it respect access controls, request limits, and the intended workflow?
  • Impact: Does it degrade service, take opportunities from people, expose data, or cause other harm?
  • Transparency: Can operators identify the activity and understand who or what initiated it?

For site operators, blocking all automation can also block useful crawlers. Bot management is generally a classification and control problem: distinguish activity by its behavior and apply suitable rules, instead of treating every automated request alike. This is a general operational principle; vendor documentation describes each vendor’s own tools and classifications.

How do web crawlers work, and what can site owners control?

A crawler requests pages, follows discovered links, and may process the returned content for a purpose such as search indexing. Google says Googlebot discovers URLs primarily through links found on pages it has already crawled. Its documented crawl behavior and limits should not be treated as universal rules for other bots.

Use robots.txt for crawl guidance

A robots.txt file communicates crawl rules to cooperating crawlers. It is not access control, and blocking a URL from crawling does not by itself guarantee that the URL will be absent from search results. A URL may still be known from links or other information.

Google distinguishes crawl blocking from indexing controls. If the goal is to prevent a page from appearing in Google Search, use an applicable noindex directive and ensure Googlebot can crawl the page to see it. Blocking the page in robots.txt can prevent the crawler from seeing a page-level directive.

Use authentication to restrict access

If content must be accessible only to authorized users, protect it with authentication such as password protection. Crawl directives do not protect private content from clients that ignore them.

Verify claimed crawler identity

Request headers, including a user-agent string, can be spoofed. Google recommends verifying purported Googlebot requests with reverse DNS or by checking Google’s published IP ranges. Apply the same general caution to other claimed bot identities: a name supplied by the requester is not authentication.

Account for crawler limits accurately

Google documents that, when crawling for Google Search, it reads up to 2 MB of supported file types and up to 64 MB of PDF files before stopping that fetch. Google also says most sites should not see Googlebot access more than once every few seconds on average, though short bursts can appear faster because of delays. These are Google-specific statements, not universal crawler guarantees. [Googlebot documentation]

How to assess automated traffic on your site

  1. Describe the observed action. Record the requested paths, methods, timing, response patterns, and resources accessed.
  2. Check the claimed identity. Treat user-agent labels as clues, not proof. Verify known crawler claims using the operator’s published verification method.
  3. Compare activity with site rules. Check authentication, crawl directives, documented APIs, terms, and rate limits that apply.
  4. Measure impact. Look at request volume, error rates, resource use, user experience, and whether the activity interferes with other users.
  5. Choose a proportionate control. Depending on the evidence, allow, rate-limit, challenge, authenticate, or block the traffic. Revisit the rule if the behavior changes.
  6. Keep useful automation working. Make sure a rule does not accidentally prevent legitimate search crawling, monitoring, or integrations that your service depends on.

For AI-related traffic, behavior can be a more useful distinction than a broad label. Cloudflare describes search activity that gathers content for later answers, real-time agent activity acting for a person, and training crawls that collect material for model training or fine-tuning. Those represent different purposes and may call for different site policies. [Cloudflare: Bots]

Performance, reliability, and cost considerations

Automated traffic has real infrastructure costs even when a bot’s intent is benign. A crawler or scraper that requests pages too frequently can consume server capacity, trigger expensive rendering, or interfere with visitors. Set and monitor sensible request limits for integrations you operate, cache repeatable responses where appropriate, and watch traffic and error patterns.

Reliability depends on the task. A monitoring bot needs a meaningful check and an alert path; a chatbot needs a fallback when it cannot answer; a crawler depends on reachable pages and stable responses. Avoid treating a single failed request as proof of abuse, and avoid depending on a bot for a critical workflow without handling timeouts and partial results.

Bot-management services and defensive controls have their own costs and configuration tradeoffs. This dossier does not establish prices, performance benchmarks, or a universally best provider. Choose controls based on the traffic you need to classify, the impact of false positives, and the operational effort required to maintain rules.

Capture a web page without managing a browser

If your bot-related workflow needs a screenshot of a page—for example, to document what a visitor sees—you can run a browser yourself or request a capture from an API. ScreenshotNeo is a website screenshot API and MCP server for developers, made by Yorker Media. It returns a PNG, JPEG, WebP, or PDF from a GET request, and supports full-page and element captures, custom waits, headers, cookies, and other capture options. See the ScreenshotNeo site and API documentation.

DIY: capture a page with Playwright

This runnable Node.js example launches Chromium, opens a page, waits for it to load, captures a full-page PNG, and closes the browser. Install Playwright and its Chromium browser first with npm install playwright and npx playwright install chromium.

const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch({ headless: true });
  try {
    const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
    await page.goto('https://example.com', {
      waitUntil: 'networkidle',
      timeout: 30000
    });
    await page.screenshot({ path: 'shot.png', fullPage: true });
  } finally {
    await browser.close();
  }
})();

For sites with long-lived network connections, networkidle may never occur. In that case, wait for a specific selector or use a bounded delay instead. For content below the fold, a full-page screenshot may require the page to load lazy images as it scrolls; inspect the output and use a deliberate scroll-and-wait strategy if needed. Keep browser processes bounded and close them even after navigation or capture errors.

Or skip the browser setup

Make one GET request to capture a page. Replace the example URL with the page you want to capture and supply your API key. The complete parameter and response options are in the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo accepts cookie and consent banners as a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up free for 1,000 screenshots a month, with no card required.

Troubleshooting bot and crawler behavior

Symptom Likely cause What to do
A page is crawled despite a robots.txt disallow rule The rule is crawl guidance for cooperating crawlers, not an access restriction; the URL may also be known without its content being crawled. Use authentication for private content. If the goal is to keep a public page out of Google Search, use an appropriate noindex directive and allow Googlebot to fetch it.
A page appears in search even though crawling is blocked Crawl blocking alone does not guarantee removal from search results. Use Google’s indexing controls for the intended outcome; do not treat robots.txt as a removal or privacy mechanism.
A request claims to be Googlebot but behaves unexpectedly The user-agent header may be spoofed. Verify through reverse DNS or Google’s published crawler IP ranges before applying an identity-based exception.
Legitimate crawlers are blocked by bot protection A broad rule may be classifying all automation as harmful. Review the triggering signals and scope the rule to observed abusive behavior while preserving necessary crawler access.
Automated traffic causes latency or errors Request volume, expensive page rendering, or repeated uncached work may be straining capacity. Measure request patterns and resource use, then apply proportionate rate limits, caching, or access controls.
A chatbot gives poor answers or traps visitors The task may exceed the bot’s limits, or the conversation may lack a human fallback. State the bot’s limits, narrow its job, improve the content it relies on, and provide a route to human support.

FAQ

Is a web bot the same as a crawler?

No. A crawler is one kind of web bot, usually designed to discover and fetch pages. Chatbots, monitors, scrapers, and transaction automation are other examples.

Does a bot have to use AI?

No. A bot can follow straightforward programmed rules. Chatbots may use menus, keyword matching, natural-language processing, or a combination.

Does robots.txt stop every bot?

No. It communicates crawl rules to cooperating crawlers; it does not authenticate clients or prevent a bot from requesting a URL.

How can I tell whether an automated request is legitimate?

Consider its purpose, verified identity where possible, authorization, request behavior, and impact. A user-agent label alone does not establish legitimacy.

Can AI agents be considered web bots?

Yes. An agent that makes automated requests or interacts with pages on a person’s behalf is a form of web automation. Its permissions and effects still matter.