A Complete Guide to Using Proxies for Web Scraping
Learn how proxies change scraper traffic, choose datacenter or residential IPs, configure rotation, validate results, and avoid common failures.

Short answer: a proxy sends your scraper’s request through an intermediary so the target sees the proxy’s exit IP instead of your server’s address. That can provide controlled egress, geographic variation, or request distribution. It does not fix bad selectors, missing JavaScript rendering, throttling, access policy, or legal restrictions.
The reliable way to use proxies is to start with a small permitted workload, select the least complex proxy type that fits the target, configure explicit timeouts and bounded retries, and inspect the returned content. A changed IP or an HTTP 200 response is not proof that you received the right page.
How a proxy fits into a scraper
Without a proxy, your HTTP client connects directly to the destination:
scraper -> destination website
With a proxy, the route becomes:
scraper -> proxy endpoint -> destination website
The destination normally records the proxy’s exit IP. A provider may also offer authentication, a pool of addresses, session controls, protocol choices, and country or region targeting. Those controls affect routing; your scraper still has to send the correct request, parse the response, and respect the destination’s policies and limits.
When a scrape is empty, check the response body, status code, redirect chain, cookies, selectors, sitemap, and JavaScript requirements before changing proxy settings. Documentation for Web Scraper specifically recommends inspecting the returned page or screenshot and checking the sitemap and driver during troubleshooting. A proxy cannot repair an incorrect selector or an application error.
Choose the proxy type
| Type | Network origin | When to evaluate it | Trade-offs |
|---|---|---|---|
| Datacenter | Hosting or datacenter infrastructure | Fast, cost-sensitive workloads and targets that permit datacenter traffic | Some sites restrict known datacenter ranges; performance and acceptance vary by provider and target |
| Residential | Consumer ISP networks | When datacenter traffic is challenged or a consumer geography is required | Can add latency and usually has a different pricing model; it is not a guarantee of access |
| ISP | Provider-specific ISP allocations | Only when your provider documents a requirement for this category | Capabilities, trust, speed, and cost are provider-specific |
| Mobile | Provider-specific mobile networks | Only when the target or location requirement calls for it | Do not assume it bypasses controls or improves success for every target |
Datacenter proxies are generally faster, while residential proxies may help when a target challenges known datacenter ranges or when location-specific content matters. These are general characteristics, not guarantees. Test the actual target with a small, permitted sample.

Use a practical selection sequence
- Define the data and geography. Decide whether you need a country, region, city, language, currency, or catalog variation. A location can change prices, availability, consent screens, and even page structure.
- Start with the least complex option. A datacenter proxy may be a reasonable first test for a target that permits it. Consider another type only after an observed requirement such as location variation or datacenter blocking.
- Check compatibility. Confirm HTTP/HTTPS or SOCKS5 support, authentication, concurrency and bandwidth limits, session controls, and the billing model.
- Validate content. Compare status codes, response bodies, language, currency, required fields, and selector behavior. Do not measure success only by whether the IP changed.
Rotation versus sticky sessions
A rotating proxy changes the exit IP according to the provider’s policy. This can suit independent requests where each page is self-contained.
A sticky session keeps the same exit IP for a configured period. It is useful for a stateful sequence in which several requests must share cookies or continuity, such as a multi-page flow. Provider semantics differ: verify the session lifetime, binding behavior, and failure handling in the current documentation.
| Workflow | Usually evaluate | Validation |
|---|---|---|
| Independent product pages | Rotation | Confirm each response contains the expected product data |
| Pagination tied to a session | Sticky session | Verify cookies, ordering, and continuity across pages |
| Login-free dashboard or multi-step form | Sticky session, if permitted | Check that the sequence remains associated with one session |
Rotation is not a responsible request schedule. Use modest concurrency, explicit timeouts, bounded retries, and backoff for errors. Never treat changing IPs as permission to ignore a site’s limits.
Configure an HTTP scraper with a proxy
Keep the endpoint and credentials outside source control. The examples below use PROXY_URL, such as a provider-issued URL in the form http://user:password@proxy.example:8080. Replace the placeholder with the endpoint supplied by your provider; do not copy credentials into a repository.

Python with requests
import os
import time
import requests
TARGET = "https://example.com/"
PROXY_URL = os.environ["PROXY_URL"]
proxies = {
"http": PROXY_URL,
"https": PROXY_URL,
}
for attempt in range(3):
try:
response = requests.get(
TARGET,
proxies=proxies,
timeout=(10, 45), # connect timeout, read timeout
headers={"User-Agent": "permitted-research-bot/1.0"},
)
response.raise_for_status()
print("status:", response.status_code)
print("bytes:", len(response.content))
print(response.text[:500])
break
except (requests.exceptions.Timeout,
requests.exceptions.ProxyError,
requests.exceptions.ConnectionError) as exc:
if attempt == 2:
raise
time.sleep(2 ** attempt)
Install the client with python -m pip install requests, set PROXY_URL, and run the file. The retry loop handles transient network failures only. Do not blindly retry an authorization failure, a policy rejection, or a malformed request.
cURL
curl --fail --silent --show-error \
--proxy "$PROXY_URL" \
--connect-timeout 10 \
--max-time 45 \
-A 'permitted-research-bot/1.0' \
'https://example.com/'
For a SOCKS5 endpoint, use the provider’s documented scheme and cURL option, commonly --socks5-hostname. Confirm the exact authentication format with the provider.
Node.js
Node’s built-in fetch does not automatically apply an HTTP proxy from PROXY_URL. Use an agent library supported by your Node version and provider, then pass that agent to the request. The following example uses https-proxy-agent:
import { HttpsProxyAgent } from 'https-proxy-agent';
const target = 'https://example.com/';
const proxyUrl = process.env.PROXY_URL;
if (!proxyUrl) throw new Error('Set PROXY_URL');
const agent = new HttpsProxyAgent(proxyUrl);
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 45_000);
try {
const res = await fetch(target, {
dispatcher: agent,
signal: controller.signal,
headers: { 'user-agent': 'permitted-research-bot/1.0' }
});
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const body = await res.text();
console.log({ status: res.status, bytes: body.length });
} finally {
clearTimeout(timer);
}
Check the agent library’s current API for your Node release before shipping. Proxy support is a client concern; a browser automation library has a separate proxy configuration and is needed when content appears only after JavaScript or interaction.
When a proxy is not enough
Choose the architecture that matches the page:
- Proxy plus your scraper: you control HTTP requests, parsing, retries, and storage. This is suitable for static responses and teams that want control over the request stack.
- Managed scraping API: you provide a URL while the service may handle proxies, retries, and rendering. Compare response format, rendering support, limits, and cost in the provider’s current documentation.
- Browser automation: use a hosted or local browser when JavaScript execution, clicking, typing, scrolling, or other interaction is required. A browser is not the same as a proxy, even when a vendor bundles both.
For JavaScript-dependent pages, first confirm that the data is absent from the raw response. If it is rendered client-side, configure a browser or rendering service; adding another IP alone will not create the missing content.
Geo-targeting changes the data
Location targeting can change the page itself. Currency, language, prices, inventory, consent dialogs, regional catalogs, and HTML structure may differ by country or city. Treat location as an input to your data model.
- Record the requested location with each result.
- Validate language and currency before parsing numeric values.
- Keep selectors flexible enough for regional variants, or maintain separate selectors.
- Compare a known page from each location before increasing volume.
Reliability, performance, and cost
Reliability checklist
- Set connect and read timeouts separately.
- Retry only transient connection, timeout, or gateway errors.
- Use exponential backoff with a maximum attempt count.
- Log proxy region, session identifier, status, response size, and elapsed time.
- Detect consent pages, bot checks, empty bodies, and unexpected redirects as data-quality failures.
- Stop or slow the job when the destination reports rate limits.
Performance considerations
Measure end-to-end time, not only proxy connection time. Residential routes can add latency; DNS, TLS negotiation, destination processing, JavaScript rendering, and response size can dominate total time. Keep concurrency within the provider’s documented limit and the destination’s acceptable rate. A larger proxy pool does not automatically improve throughput if the target or your parser is the bottleneck.
Cost considerations
Providers may charge by bandwidth, requests, ports, concurrency, or subscription. Compare the cost of failed requests, browser minutes, and storage as well as the nominal proxy rate. Run a small sample and calculate cost per valid record, because a successful HTTP response containing the wrong page has no data value.
Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Proxy authentication error | Wrong credentials, expired key, or unsupported auth format | Copy the provider’s current endpoint format, URL-encode special characters, and test one request |
| Connection timeout | Unavailable endpoint, overloaded route, firewall, or excessive timeout | Check the endpoint and allow outbound traffic; test another permitted route and set bounded timeouts |
| HTTP 403 or 429 | Destination policy, rate limit, or blocked network origin | Reduce concurrency, back off, review site policy, and validate whether the task is permitted; do not assume rotation solves it |
| HTTP 200 but empty fields | Wrong selector, consent wall, bot page, or JavaScript-rendered content | Save and inspect the body or screenshot; fix parsing or use a browser/rendering path |
| Wrong language or prices | Geo-targeted response differs from expectations | Choose the required location and validate currency, language, and selectors |
| Login or checkout sequence breaks | IP changed between requests or cookies were not retained | Use a documented sticky session, preserve cookies, and confirm that the workflow is allowed |
| Only some requests fail | Provider pool quality, destination variance, or intermittent network errors | Log exit and response details, retry transient failures with backoff, and measure valid-result rate |
Compliance and responsible collection
robots.txt is a crawler convention. RFC 9309 asks crawlers to honor its rules and expressly states: “These rules are not a form of access authorization.” Read the IETF Robots Exclusion Protocol, the site’s terms, and applicable privacy, data-protection, contract, and intellectual-property requirements separately.
A proxy changes routing. It does not make restricted information public, grant permission, or settle the legal analysis for your jurisdiction and purpose. Use public data where appropriate, avoid private or sensitive personal data without permission, identify your crawler, and seek qualified advice for a consequential or uncertain use.
Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server. Use the do-it-yourself proxy approach above when you need to own the request and parsing stack. For visual capture, one request returns a PNG, JPEG, WebP, or PDF, with options for full-page shots, CSS-element capture, device presets, custom viewports, dark mode, retina scale, waits, custom CSS and JavaScript, headers, cookies, geolocation, caching, and more.
cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for the full parameter set. Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server lets AI agents such as Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.
Start with 1,000 free screenshots a month—no card required.
FAQ
Are residential proxies good for web scraping?
They can be useful when a target challenges datacenter traffic or when you need consumer ISP geography. They may add latency, and they do not guarantee access or permission. Test the target and validate the returned data.
Does changing the proxy IP prevent blocking?
No. A destination can use rate limits, fingerprints, cookies, behavior, and other controls. Rotation changes routing only.
Should I use a proxy or a headless browser?
Use a proxy when controlled egress or location is the requirement. Use a browser when JavaScript execution or interaction is required. Some managed products combine both, but they solve different layers.
How many requests can I send?
There is no universal safe rate. Follow the destination’s limits and your provider’s concurrency terms, then increase volume only after a small permitted test shows correct content and stable behavior.
Can robots.txt authorize scraping?
No. RFC 9309 describes robots.txt as a crawler convention and says its rules are not access authorization. Review all other applicable policies and requirements.


