ScreenshotNeo

BlogGuides

How to Choose a Web Scraping Service

Choose a scraping service by testing it against your real pages, fields, and volume. Compare valid results, integration effort, reliability, and total cost.

By the ScreenshotNeo team4 October 20269 min read

The right web scraping service is the one that returns complete, correct, usable records for your actual target pages at a predictable total cost. Start with a representative workload, define what counts as a valid result, and compare providers using cost per validated result—not a headline price per request, credit, or record.

First decide what kind of service you need. A hosted scraping API may handle rendering, retries, and extraction; a no-code tool may suit one-off collection; a managed pipeline may deliver recurring datasets; proxy or browser infrastructure still leaves you responsible for the scraper and its maintenance.

1. Define the workload before shopping

Write down the job you need the service to do. A vendor cannot meaningfully answer “Can you scrape this site?” without knowing which pages, which fields, how often, and what a correct result looks like.

Specify Questions to answer
Targets Which exact domains, URL patterns, page types, and representative URLs are in scope?
Fields Which values must be extracted? Which may be missing? What formats and units should they use?
Validity What makes a record usable: required fields present, parsing correct, fresh enough, no duplicate?
Volume and cadence How many pages per run and per month? How often must data refresh? Are bursts expected?
Page behavior Are pages static, JavaScript-rendered, session-dependent, or region-specific? Do pages paginate or load content on scroll?
Geography Must pages be fetched from a specific country, region, or locale?
Latency Is a result needed interactively, within minutes, or in a scheduled batch?
Delivery Do you need JSON, CSV, files, a webhook, a storage destination, or an API response?
Operations Who handles schema changes, failed runs, retries, monitoring, and site changes?

Use URLs you are permitted to collect and that represent normal production variation: multiple page templates, missing optional fields, different locales if relevant, and both easy and difficult pages. Avoid choosing only a provider’s demo URL.

2. Choose the service type that fits

Hosted scraping API

Choose this when your application needs a programmable endpoint and you want a provider to handle some combination of rendering, extraction, retries, or access infrastructure. Verify exactly which of those are included and how they affect metering.

No-code scraper

Choose a visual interface for small projects, analyst-led work, or a prototype where a developer-owned integration is unnecessary. Confirm scheduling, export formats, and whether the workflow can be maintained when the target page changes.

Managed data pipeline

Choose a managed delivery when you need ongoing datasets and want an owner for extraction and delivery operations. Clarify schema ownership, freshness, backfills, change requests, retention, and support responsibilities in writing.

Proxy or browser infrastructure

This is a building block, not a complete extraction service. Your team still builds request logic or browser automation, parses pages, validates records, retries failures, monitors changes, and maintains the system.

3. Evaluate output quality and coverage

Request a pilot on representative URLs and score the returned records against a small reference set you inspect yourself. Record the denominator: attempted pages, completed pages, and pages that passed your validity rules. A successful HTTP response is not necessarily a successful extraction.

  • Completeness: Are required fields present? Are optional omissions represented consistently?
  • Correctness: Do values match the page, including currency, dates, units, and nested variants?
  • Freshness: Does the extraction reflect the page state at the time and place your use case requires?
  • Duplicates: Can retries or pagination create duplicate records? Is there a stable key for deduplication?
  • Failure visibility: Can you distinguish unavailable pages, timeouts, parsing failures, and valid pages with empty fields?
  • Change behavior: What happens when the page layout or target fields change, and who notices?

Ask providers how they define any advertised success figure and whether it applies to your URLs, geography, page types, and measurement window. Treat vendor performance claims as vendor claims until you validate them against your own workload.

4. Check the features your pages actually need

Capability Check before buying
JavaScript rendering Does the service wait for client-rendered fields? Can you control wait conditions, and is rendering included or separately metered?
Proxy and geography Can it fetch from the locations you need? Ask how location options affect price and whether your pilot confirms the expected localized page.
Structured extraction Does it return the fields and schema you need, or only page content for your code to parse?
Retries and error detail Can you configure retry limits and backoff? Are errors and partial results visible rather than silently converted to empty data?
Concurrency and rate controls What are the limits, how are bursts handled, and can you slow collection to fit your workload?
Scheduling and delivery Are scheduled jobs, webhooks, storage destinations, and the formats you need available on the relevant plan?
Sessions and authentication Can it support the legitimate access context your use case requires? Check the provider’s policy and the target’s rules.
Support and maintenance Who investigates failed runs and extraction changes? What response path and service commitments are actually in the contract?

Bright Data’s own Web Scraper API page describes API and no-code workflows, JavaScript rendering, proxy management, concurrency, and JSON, NDJSON, or CSV delivery. These are vendor-described capabilities, not independent proof that a service will succeed on your targets. Check the [official product page](https://brightdata.com/products/web-scraper) and test your URLs.

5. Compare total cost per valid result

Providers may meter records, requests, page loads, credits, bandwidth, or runtime. Those units are not interchangeable. A low unit price can still produce a high bill if your workload needs rendering, geographic access, retries, or other metered options.

Use one equivalent pilot workload for each candidate. For each provider, calculate:

cost_per_valid_result = total_pilot_charge / number_of_validated_results

Include the complete charge, including plan minimums, metering multipliers, overages, storage or retention, support, and any required infrastructure. Also track engineering hours to integrate and maintain the workflow; a cheaper service can cost more overall if it leaves substantial operations work with your team.

At the time of research, Bright Data’s official pricing page listed 5,000 records per month in a free tier, pay-as-you-go at $1.50 per 1,000 records, and a Scale plan at $499 monthly including 384,000 records, with additional records listed at $1.30 per 1,000. This is a vendor-specific snapshot, not a market benchmark; prices and definitions can change. Confirm current USD pricing, what counts as a record, and contract terms on the [official pricing page](https://brightdata.com/pricing/web-scraper) before purchase.

Do not compare a provider’s “record” directly with another service’s credit, request, page load, bandwidth, or runtime. Run the same sample and compute the cost for results that pass the same validation rules.

6. Run a controlled pilot

  1. Freeze the sample. Keep the same representative URLs, required fields, region, and collection window for each candidate.
  2. Write validation rules. Specify required fields, acceptable formats, freshness, deduplication, and how partial results count.
  3. Run production-like settings. Match expected concurrency, schedule, rendering, and delivery method. Record every option that changes metering.
  4. Measure outcomes. Track valid records, completeness, correctness, duplicates, failures by cause, elapsed time, and total charge.
  5. Test recovery. Retry a controlled failed or delayed job and confirm whether the system exposes the failure and avoids duplicate downstream writes.
  6. Estimate at your real volume. Apply the observed valid-result rate and metering to expected monthly volume, then ask the vendor to confirm the estimate.
  7. Decide with evidence. Pick the candidate that meets quality, latency, and operational requirements within budget. If none does, revise the design or narrow the workload.

7. Build reliability into the integration

Even a hosted service does not remove the need for application-side safeguards. Treat extraction as an unreliable external dependency and make failure visible.

  • Use bounded retries with backoff for transient timeouts or temporary service errors; do not retry indefinitely.
  • Separate transport success from extraction validity. Validate required fields before writing a result as complete.
  • Make downstream writes idempotent, using a stable record key or job identifier where available.
  • Log the target URL, job or request ID, timestamp, status, retry count, and validation outcome. Avoid logging secrets or unnecessary personal data.
  • Alert on changes in failure rate, missing required fields, stale data, or unusual cost per valid result.
  • Set concurrency and schedule to your actual need. More parallel work can increase load, rate-limit failures, and cost without improving useful throughput.
  • Plan for schema changes with versioned output or compatibility checks, and keep a small regression set of representative pages.

8. Collect responsibly

Review the target site’s terms, robots rules, privacy and data-protection obligations, and the scraping provider’s acceptable-use policy for the specific data, collection location, and intended use. A paid service does not automatically make a collection permissible.

Google describes robots.txt as a way for site owners to communicate crawler access preferences and manage crawl traffic. It is not a security mechanism, does not enforce behavior for every crawler, and a blocked URL can still appear in Google Search if discovered through links. These are statements about Google’s crawler documentation, not a universal legal ruling. Read [Google’s robots.txt guidance](https://developers.google.com/search/docs/crawling-indexing/robots/intro) and assess the rules that apply to your own collection.

9. Troubleshooting during evaluation

Symptom Likely cause What to do
HTTP success, but required fields are empty The page has not rendered, the selector or extraction schema is wrong, or the page returned a different state. Inspect the returned content and error metadata; test the same URL with JavaScript rendering and a suitable wait condition, then validate fields explicitly.
Results work for some URLs but fail for others Different templates, locale, pagination, or optional page states are not covered by the pilot configuration. Group URLs by page type and test each group. Add explicit handling for optional fields and pagination.
Frequent timeouts Slow page resources, overly strict waits, overloaded concurrency, or a target that is unavailable. Compare elapsed time by URL, tune wait conditions, reduce concurrency, and establish a timeout appropriate to the workload.
Duplicates after retries A retry repeated a completed operation, or downstream writes are not idempotent. Use stable keys or job IDs, deduplicate before publishing records, and distinguish retry attempts from completed results.
Bill is higher than the estimate The estimate used the wrong billing unit or omitted rendering, retries, geography, overages, or minimums. Reconcile usage against the provider’s meter, identify each option’s multiplier, and recalculate using valid-result count.
Data becomes stale or changes shape The schedule is too infrequent, a page changed, or schema assumptions are brittle. Monitor freshness and required fields, add a regression sample, and agree who owns extraction updates.
Access is denied or collection is disallowed The target or provider policy does not permit the requested collection. Stop the affected collection and review the target rules, provider acceptable-use policy, and applicable obligations. Do not treat changing access infrastructure as a policy solution.

10. When a screenshot API is the better fit

If your output is a visual page capture rather than structured records, use a screenshot API instead of buying a scraping pipeline. ScreenshotNeo is a website screenshot API and MCP server for developers. It returns PNG, JPEG, WebP, or PDF from a URL and supports options such as full-page capture, element selection, JavaScript waits, custom CSS, and device viewports. It is an alternative for visual capture tasks, not a substitute for extracting structured fields from pages.

Or skip the browser setup

Make one GET request with a URL; see the ScreenshotNeo API documentation for options and formats.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie and consent banners, newsletter popups, and chat widgets are removed before the screenshot; each step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits are never billed; response headers identify the page verdict and billing status.
  • An MCP server lets AI agents use screenshot, page-info, and PDF-capture tools.
  • 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000.

Sign up free for 1,000 screenshots a month, no card required.

Frequently asked questions

Should I build my own scraper or buy a service?

Compare the full operating cost, not just the initial code. If your pages are stable and volume is limited, a small in-house scraper may be sufficient; if rendering, access variation, monitoring, or maintenance dominate, evaluate hosted or managed options with a pilot.

Can I choose based on the largest free tier?

Use a free tier to validate integration, but confirm the same workload’s quality, limits, and paid pricing before deciding. A free allowance does not establish production suitability.

Is a scraping service automatically compliant?

No. Evaluate the target’s rules, applicable obligations, and the service’s acceptable-use policy for your particular collection and intended use.

How many URLs should a pilot include?

Enough to cover each important page type and normal variation, including difficult cases. Representativeness matters more than a large count of nearly identical URLs.