ScreenshotNeo

BlogGuides

A Complete Guide to Using Proxies for Web Scraping

Learn how proxies change scraper traffic, choose datacenter or residential IPs, configure rotation, validate results, and avoid common failures.

By the ScreenshotNeo team29 September 20269 min read

A Complete Guide to Using Proxies for Web Scraping

Short answer: a proxy sends your scraper’s request through an intermediary so the target sees the proxy’s exit IP instead of your server’s address. That can provide controlled egress, geographic variation, or request distribution. It does not fix bad selectors, missing JavaScript rendering, throttling, access policy, or legal restrictions.

The reliable way to use proxies is to start with a small permitted workload, select the least complex proxy type that fits the target, configure explicit timeouts and bounded retries, and inspect the returned content. A changed IP or an HTTP 200 response is not proof that you received the right page.

How a proxy fits into a scraper

Without a proxy, your HTTP client connects directly to the destination:

scraper -> destination website

With a proxy, the route becomes:

scraper -> proxy endpoint -> destination website

The destination normally records the proxy’s exit IP. A provider may also offer authentication, a pool of addresses, session controls, protocol choices, and country or region targeting. Those controls affect routing; your scraper still has to send the correct request, parse the response, and respect the destination’s policies and limits.

When a scrape is empty, check the response body, status code, redirect chain, cookies, selectors, sitemap, and JavaScript requirements before changing proxy settings. Documentation for Web Scraper specifically recommends inspecting the returned page or screenshot and checking the sitemap and driver during troubleshooting. A proxy cannot repair an incorrect selector or an application error.

Choose the proxy type

Type Network origin When to evaluate it Trade-offs
Datacenter Hosting or datacenter infrastructure Fast, cost-sensitive workloads and targets that permit datacenter traffic Some sites restrict known datacenter ranges; performance and acceptance vary by provider and target
Residential Consumer ISP networks When datacenter traffic is challenged or a consumer geography is required Can add latency and usually has a different pricing model; it is not a guarantee of access
ISP Provider-specific ISP allocations Only when your provider documents a requirement for this category Capabilities, trust, speed, and cost are provider-specific
Mobile Provider-specific mobile networks Only when the target or location requirement calls for it Do not assume it bypasses controls or improves success for every target

Datacenter proxies are generally faster, while residential proxies may help when a target challenges known datacenter ranges or when location-specific content matters. These are general characteristics, not guarantees. Test the actual target with a small, permitted sample.

A proxy changes the network exit point, while the scraper still controls requests and parsing.
A proxy changes the network exit point, while the scraper still controls requests and parsing.

Use a practical selection sequence

  1. Define the data and geography. Decide whether you need a country, region, city, language, currency, or catalog variation. A location can change prices, availability, consent screens, and even page structure.
  2. Start with the least complex option. A datacenter proxy may be a reasonable first test for a target that permits it. Consider another type only after an observed requirement such as location variation or datacenter blocking.
  3. Check compatibility. Confirm HTTP/HTTPS or SOCKS5 support, authentication, concurrency and bandwidth limits, session controls, and the billing model.
  4. Validate content. Compare status codes, response bodies, language, currency, required fields, and selector behavior. Do not measure success only by whether the IP changed.

Rotation versus sticky sessions

A rotating proxy changes the exit IP according to the provider’s policy. This can suit independent requests where each page is self-contained.

A sticky session keeps the same exit IP for a configured period. It is useful for a stateful sequence in which several requests must share cookies or continuity, such as a multi-page flow. Provider semantics differ: verify the session lifetime, binding behavior, and failure handling in the current documentation.

Workflow Usually evaluate Validation
Independent product pages Rotation Confirm each response contains the expected product data
Pagination tied to a session Sticky session Verify cookies, ordering, and continuity across pages
Login-free dashboard or multi-step form Sticky session, if permitted Check that the sequence remains associated with one session

Rotation is not a responsible request schedule. Use modest concurrency, explicit timeouts, bounded retries, and backoff for errors. Never treat changing IPs as permission to ignore a site’s limits.

Configure an HTTP scraper with a proxy

Keep the endpoint and credentials outside source control. The examples below use PROXY_URL, such as a provider-issued URL in the form http://user:password@proxy.example:8080. Replace the placeholder with the endpoint supplied by your provider; do not copy credentials into a repository.

Rotation suits independent requests; sticky sessions preserve continuity for stateful flows.
Rotation suits independent requests; sticky sessions preserve continuity for stateful flows.

Python with requests

import os
import time
import requests

TARGET = "https://example.com/"
PROXY_URL = os.environ["PROXY_URL"]

proxies = {
    "http": PROXY_URL,
    "https": PROXY_URL,
}

for attempt in range(3):
    try:
        response = requests.get(
            TARGET,
            proxies=proxies,
            timeout=(10, 45),  # connect timeout, read timeout
            headers={"User-Agent": "permitted-research-bot/1.0"},
        )
        response.raise_for_status()
        print("status:", response.status_code)
        print("bytes:", len(response.content))
        print(response.text[:500])
        break
    except (requests.exceptions.Timeout,
            requests.exceptions.ProxyError,
            requests.exceptions.ConnectionError) as exc:
        if attempt == 2:
            raise
        time.sleep(2 ** attempt)

Install the client with python -m pip install requests, set PROXY_URL, and run the file. The retry loop handles transient network failures only. Do not blindly retry an authorization failure, a policy rejection, or a malformed request.

cURL

curl --fail --silent --show-error \
  --proxy "$PROXY_URL" \
  --connect-timeout 10 \
  --max-time 45 \
  -A 'permitted-research-bot/1.0' \
  'https://example.com/'

For a SOCKS5 endpoint, use the provider’s documented scheme and cURL option, commonly --socks5-hostname. Confirm the exact authentication format with the provider.

Node.js

Node’s built-in fetch does not automatically apply an HTTP proxy from PROXY_URL. Use an agent library supported by your Node version and provider, then pass that agent to the request. The following example uses https-proxy-agent:

import { HttpsProxyAgent } from 'https-proxy-agent';

const target = 'https://example.com/';
const proxyUrl = process.env.PROXY_URL;
if (!proxyUrl) throw new Error('Set PROXY_URL');

const agent = new HttpsProxyAgent(proxyUrl);
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 45_000);

try {
  const res = await fetch(target, {
    dispatcher: agent,
    signal: controller.signal,
    headers: { 'user-agent': 'permitted-research-bot/1.0' }
  });
  if (!res.ok) throw new Error(`HTTP ${res.status}`);
  const body = await res.text();
  console.log({ status: res.status, bytes: body.length });
} finally {
  clearTimeout(timer);
}

Check the agent library’s current API for your Node release before shipping. Proxy support is a client concern; a browser automation library has a separate proxy configuration and is needed when content appears only after JavaScript or interaction.

When a proxy is not enough

Choose the architecture that matches the page:

  • Proxy plus your scraper: you control HTTP requests, parsing, retries, and storage. This is suitable for static responses and teams that want control over the request stack.
  • Managed scraping API: you provide a URL while the service may handle proxies, retries, and rendering. Compare response format, rendering support, limits, and cost in the provider’s current documentation.
  • Browser automation: use a hosted or local browser when JavaScript execution, clicking, typing, scrolling, or other interaction is required. A browser is not the same as a proxy, even when a vendor bundles both.

For JavaScript-dependent pages, first confirm that the data is absent from the raw response. If it is rendered client-side, configure a browser or rendering service; adding another IP alone will not create the missing content.

Geo-targeting changes the data

Location targeting can change the page itself. Currency, language, prices, inventory, consent dialogs, regional catalogs, and HTML structure may differ by country or city. Treat location as an input to your data model.

  1. Record the requested location with each result.
  2. Validate language and currency before parsing numeric values.
  3. Keep selectors flexible enough for regional variants, or maintain separate selectors.
  4. Compare a known page from each location before increasing volume.

Reliability, performance, and cost

Reliability checklist

  • Set connect and read timeouts separately.
  • Retry only transient connection, timeout, or gateway errors.
  • Use exponential backoff with a maximum attempt count.
  • Log proxy region, session identifier, status, response size, and elapsed time.
  • Detect consent pages, bot checks, empty bodies, and unexpected redirects as data-quality failures.
  • Stop or slow the job when the destination reports rate limits.

Performance considerations

Measure end-to-end time, not only proxy connection time. Residential routes can add latency; DNS, TLS negotiation, destination processing, JavaScript rendering, and response size can dominate total time. Keep concurrency within the provider’s documented limit and the destination’s acceptable rate. A larger proxy pool does not automatically improve throughput if the target or your parser is the bottleneck.

Cost considerations

Providers may charge by bandwidth, requests, ports, concurrency, or subscription. Compare the cost of failed requests, browser minutes, and storage as well as the nominal proxy rate. Run a small sample and calculate cost per valid record, because a successful HTTP response containing the wrong page has no data value.

Troubleshooting common failures

Symptom Likely cause Fix
Proxy authentication error Wrong credentials, expired key, or unsupported auth format Copy the provider’s current endpoint format, URL-encode special characters, and test one request
Connection timeout Unavailable endpoint, overloaded route, firewall, or excessive timeout Check the endpoint and allow outbound traffic; test another permitted route and set bounded timeouts
HTTP 403 or 429 Destination policy, rate limit, or blocked network origin Reduce concurrency, back off, review site policy, and validate whether the task is permitted; do not assume rotation solves it
HTTP 200 but empty fields Wrong selector, consent wall, bot page, or JavaScript-rendered content Save and inspect the body or screenshot; fix parsing or use a browser/rendering path
Wrong language or prices Geo-targeted response differs from expectations Choose the required location and validate currency, language, and selectors
Login or checkout sequence breaks IP changed between requests or cookies were not retained Use a documented sticky session, preserve cookies, and confirm that the workflow is allowed
Only some requests fail Provider pool quality, destination variance, or intermittent network errors Log exit and response details, retry transient failures with backoff, and measure valid-result rate

Compliance and responsible collection

robots.txt is a crawler convention. RFC 9309 asks crawlers to honor its rules and expressly states: “These rules are not a form of access authorization.” Read the IETF Robots Exclusion Protocol, the site’s terms, and applicable privacy, data-protection, contract, and intellectual-property requirements separately.

A proxy changes routing. It does not make restricted information public, grant permission, or settle the legal analysis for your jurisdiction and purpose. Use public data where appropriate, avoid private or sensitive personal data without permission, identify your crawler, and seek qualified advice for a consequential or uncertain use.

Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. Use the do-it-yourself proxy approach above when you need to own the request and parsing stack. For visual capture, one request returns a PNG, JPEG, WebP, or PDF, with options for full-page shots, CSS-element capture, device presets, custom viewports, dark mode, retina scale, waits, custom CSS and JavaScript, headers, cookies, geolocation, caching, and more.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for the full parameter set. Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server lets AI agents such as Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.

Start with 1,000 free screenshots a month—no card required.

FAQ

Are residential proxies good for web scraping?

They can be useful when a target challenges datacenter traffic or when you need consumer ISP geography. They may add latency, and they do not guarantee access or permission. Test the target and validate the returned data.

Does changing the proxy IP prevent blocking?

No. A destination can use rate limits, fingerprints, cookies, behavior, and other controls. Rotation changes routing only.

Should I use a proxy or a headless browser?

Use a proxy when controlled egress or location is the requirement. Use a browser when JavaScript execution or interaction is required. Some managed products combine both, but they solve different layers.

How many requests can I send?

There is no universal safe rate. Follow the destination’s limits and your provider’s concurrency terms, then increase volume only after a small permitted test shows correct content and stable behavior.

Can robots.txt authorize scraping?

No. RFC 9309 describes robots.txt as a crawler convention and says its rules are not access authorization. Review all other applicable policies and requirements.