7 Web Scraping Tips for Reliable Scraping
Make web scraping more reliable with seven practical habits for crawler rules, request pacing, batching, monitoring, and handling robots.txt.

Reliable web scraping starts before the first page request: check the target site’s crawler rules, identify your crawler, make requests at a considerate pace, and watch how the site responds. Then keep collection focused with sitemaps and manageable batches. These practices help make a run easier to observe and resume while reducing avoidable load on the site.
This guide covers seven practical habits, including what robots.txt does and does not mean, how to interpret common response signals, and how to check that the rules you read apply to the host you plan to fetch. The specific rates cited below are examples from AWS guidance, not universal limits. A target’s own instructions and your legal and contractual obligations still matter.
1. Check crawler rules before fetching
Before collecting pages, inspect the target host’s robots.txt file and follow its parseable rules. The Robots Exclusion Protocol (REP), standardized in RFC 9309, lets site operators communicate which crawler paths they request bots to avoid. AWS also recommends checking and respecting robots.txt as part of responsible crawling.

Keep the boundary clear: robots.txt is not a login system, permission grant, or legal review. RFC 9309 explicitly says, “These rules are not a form of access authorization.” A disallowed path is a crawler coordination signal; an allowed path does not establish that you have permission to access it or use its contents. Check applicable terms, contracts, and laws separately.
Use the file to decide what your crawler should fetch, not as a way to probe protected content. If a resource requires authentication, do not treat an Allow rule as authorization to bypass that control.
2. Identify your crawler clearly
Send a descriptive HTTP User-Agent header so operators can distinguish your crawler from ordinary browser traffic. AWS recommends identifying the crawler and says contact information is commonly included. For example, a team might use a name and a contact page or email address it controls. Do not imply that identification guarantees access; it makes the traffic more transparent and gives operators a route to reach you.
Keep the identifier stable across runs. If the crawler’s purpose or contact changes, update it. Avoid impersonating a popular browser or another crawler to get around a site’s controls. If the site asks your crawler to stop, honor that signal rather than rotating identities to continue.
3. Pace requests and respond to site load
Choose a conservative request rate, then adapt to the target’s responses. A fast scraper can turn a small collection task into unnecessary load, while a pace that is too aggressive can produce throttling, errors, and incomplete data.
AWS gives contextual examples in its guidance for ethical web crawlers: one request every 10–15 seconds for small or medium-sized websites, and 1–2 requests per second for larger websites or sites with explicit crawl permission. These are examples, not standards or safe limits for every site. Start lower when you do not know the site’s capacity, and follow any more restrictive instructions the operator publishes.
- HTTP 429, “Too Many Requests”: pause requests. Do not keep sending at the same rate. Resume only after a meaningful pause and at a lower rate, while respecting any retry guidance the site provides.
- HTTP 403, “Forbidden”: treat repeated responses as a reason to consider stopping. Do not attempt to evade the site’s restriction.
- Rising latency or 5xx responses: reduce load or pause and investigate. These are useful operational signals that the site may be under strain or having trouble.
Google documents that slower responses, 5xx errors, and rate-limit signals such as 429 can reduce its own crawler capacity. That describes Google’s crawler behavior; it does not define a universal threshold for independent scrapers. Use the signals to make your job more considerate, not to infer a general permitted rate. See Google’s crawl budget guidance for that crawler-specific context.
4. Use sitemaps to focus collection
When a site publishes a sitemap, use it as a discovery aid for the pages relevant to your task. AWS recommends sitemaps to focus collection on important pages. A sitemap can help avoid broad, wasteful URL guessing and make the intended scope easier to review.
A sitemap is not a complete inventory guarantee or permission signal. It may include pages outside your project, omit pages you need, or change over time. Check the URLs against the collection purpose and the host’s crawler rules before fetching. If the sitemap references additional sitemap files, inspect them under the same scope and pacing policy.
Record where each discovered URL came from and the time the discovery list was produced. That makes it easier to understand why a URL was included when the site’s structure changes between runs.
5. Divide large jobs into batches
Do not make a large collection one opaque run. Divide the URL set into smaller batches, as AWS recommends, so each unit of work is easier to monitor and less likely to run into a single timeout or resource constraint. Batching also gives an operator clearer checkpoints for resuming a long job; that is a practical consequence of the approach, not a guaranteed outcome.

- Build and review the candidate URL list.
- Group URLs into bounded batches that suit your available resources and the site’s request pace.
- Record which URLs were attempted and the outcome of each request.
- Review failures before continuing. Retry only when appropriate and without increasing load on a struggling site.
- Keep completed batches so an interruption does not require starting the whole collection over.
Batch size and request rate are separate controls: a small batch can still send requests too quickly, and a large batch can be processed at a considerate pace. Apply both limits deliberately.
6. Handle robots.txt fetch outcomes deliberately
Do not silently treat every robots.txt fetch failure as permission to crawl. RFC 9309 specifies different handling depending on what happened:
| Fetch outcome | What RFC 9309 says | Practical response |
|---|---|---|
| Successfully retrieved | Crawlers must follow parseable rules. | Parse the file and apply its rules to the URLs in scope. |
| Unreachable due to server or network errors | Crawlers must assume complete disallow. | Pause collection for that scope; do not proceed as if no rules exist. |
| Unavailable response in the 4xx range | Crawlers may access resources on the server. | Apply the standard carefully, preserve the observed outcome, and still consider the site’s other instructions and access restrictions. |
Use the RFC’s categories rather than collapsing all errors into “file missing.” A timeout, a server error, and a 404 are distinct outcomes. Retain the status and fetch time in your run records so an operator can understand which rule decision was made.
RFC 9309 also says crawlers should follow at least five consecutive redirects when retrieving robots.txt. It says a cached robots.txt should not be used for more than 24 hours unless the file is unreachable, and sets a minimum parsing limit of 500 KiB. These are protocol details that matter when building a robust rules fetcher; do not assume an arbitrarily stale cached file remains current.
7. Check the scope of the rules you read
A robots.txt file applies to a particular origin scope. Google documents that its robots.txt interpretation is limited to the host, protocol, and port where the file is hosted: a file on www.example.com does not automatically cover example.com, another subdomain, HTTP instead of HTTPS, or a different port. RFC 9309 likewise makes the rules you retrieve relevant to the authority they belong to.
Before applying a rules file, compare its origin with the URLs you intend to request. For example, do not assume rules fetched from https://www.example.com/robots.txt apply to https://shop.example.com/. Fetch and evaluate the appropriate file for each distinct host, protocol, and port in scope. Google’s explanation is at How Google interprets the robots.txt specification; its description is authoritative for Google’s crawler, while RFC 9309 is the general protocol standard.
Make a run observable
Responsible request behavior is only part of reliability. A collection can finish without an obvious crash and still omit pages or contain failed responses. Keep enough per-request information to explain what happened, while avoiding unnecessary storage of sensitive response content.
- Record the requested URL, request time, response status, and elapsed time.
- Track robots.txt fetch results and the applicable origin scope.
- Count successes, 429s, 403s, 5xx responses, timeouts, and other failures separately.
- Compare attempted URLs with completed results at each batch checkpoint.
- Review changes in latency and error rates before starting the next batch.
These are operational practices for making omissions visible; the cited sources do not prescribe a complete data-quality workflow or claim a particular improvement rate. Choose checks that fit your data and downstream use. If a missing page would change a decision, define how the run will flag that gap before collection begins.
Common problems and fixes
| Symptom | Likely cause | What to do |
|---|---|---|
| Repeated 429 responses | The site is rate-limiting requests or the crawl pace is too high. | Pause. Reduce the request rate before any resumption, and follow site-provided guidance. |
| Repeated 403 responses | The site is refusing the request or restricting the crawler. | Consider stopping. Do not try to evade the restriction. |
| Robots file times out or returns a server error | The rules file is unreachable because of a server or network problem. | Under RFC 9309, assume complete disallow while it is unreachable. |
| Rules appear inconsistent across URLs | The URLs may use a different host, protocol, or port. | Fetch and apply the robots.txt for the actual origin of each URL. |
| Long run stops partway through | The job may be too large or constrained by time or resources. | Use smaller batches with recorded checkpoints so completed work is identifiable. |
| Slow responses or 5xx errors increase | The target may be struggling or the request load may be too high. | Reduce load or pause, then review before continuing. |
| Some expected pages are absent | Discovery may be incomplete, a batch may have failed, or rules may exclude those paths. | Compare discovered and attempted URLs, inspect outcomes, and revisit scope and rules. |
Performance, reliability, and cost
For a self-managed crawler, the main trade-off is throughput against considerate load and the time required to recover from a partial run. Increasing concurrency can shorten collection time, but it also increases simultaneous demand on the target. Set the pace for the site, not just for your own compute capacity. A slow, checkpointed run can be more useful than a fast run whose failures and omissions are hard to reconstruct.
Cost depends on the infrastructure and labor you use, plus the cost of storing and processing collected data. The sources here do not provide a universal cost estimate, crawler success rate, or benchmark. Avoid using a nominal requests-per-second figure as a proxy for completeness: inspect the actual statuses, latency, and missing work.
If you need screenshots of rendered pages for visual review or documentation alongside a collection, a screenshot API can capture that evidence without making it a substitute for responsible crawling. For a screenshot-only task, ScreenshotNeo is a website screenshot API and MCP server: it returns PNG, JPEG, WebP, or PDF from one GET request. Its clean-shot flow accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; those steps can be turned off. Only clean shots are billed: bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with the outcome reported in response headers. Every feature is on every plan. Pricing starts with 1,000 free shots per month with no card, then $5 for 3,000 shots; see the ScreenshotNeo documentation for the API details.
Or skip the browser setup
If you need a screenshot of a page as it renders, one GET request can return the image. This is useful for visual checks alongside a crawler run; it does not replace the crawler rules, pacing, or access review described above.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Equivalent Python and Node.js calls, plus the available parameters, are in the ScreenshotNeo API documentation.
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners, popups, and chat widgets are removed before the shot.
- Bot checks, blank pages, and failed loads are never billed.
- An MCP server lets AI agents use screenshot tools.
- 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000.
Sign up for 1,000 free screenshots a month, with no card required.
FAQ
Does an Allow rule in robots.txt mean I have permission?
No. Robots.txt is a crawler coordination protocol, not access authorization. Check the site’s applicable terms, contracts, and laws separately.
How fast should my scraper make requests?
There is no universal safe rate. AWS offers contextual examples, but the right pace depends on the site and its instructions. Start conservatively, monitor responses, and pause on 429.
Should I continue if robots.txt cannot be reached?
RFC 9309 says to assume complete disallow when the file is unreachable because of server or network errors. Handle an unavailable 4xx response separately under the standard.
Does one robots.txt cover all subdomains?
No. Check the file for the URL’s actual host, protocol, and port. A file on one origin does not automatically apply to another.
Are AWS’s example request rates limits?
No. They are examples in AWS guidance, not universal standards or guarantees that a particular target permits that rate.


