ScreenshotNeo

BlogGuides

Residential Proxying for Web Scraping and Browser Automation

Learn how residential proxies route browser traffic, configure them with Playwright, choose rotating or sticky sessions, and keep collection authorized.

By the ScreenshotNeo team4 October 202612 min read

A residential proxy routes a browser or other client through a proxy server whose outgoing connection uses a residential-network IP address. To use one with Playwright, configure the provider’s HTTP(S) or SOCKSv5 endpoint at browser launch or on a browser context, then run only collection that the site owner and applicable rules permit. A proxy changes where traffic exits; it does not provide authorization, make restricted collection permissible, or guarantee access.

1. What a residential proxy does

The client connects to the proxy endpoint. The proxy forwards the request to the destination, which sees the proxy’s egress IP rather than the client’s direct egress IP. Depending on the provider and plan, controls may include a target location, rotating egress, or a persistent session. Those controls vary: check the provider’s endpoint and authentication documentation.

A proxy is one part of a collection system. It does not fetch and parse data by itself, operate a browser, decide whether collection is allowed, or guarantee that a request will succeed. Vendor pages often describe scraping and browser automation as use cases; those are vendor descriptions, not independent evidence of access or performance.

Typical request path:

Playwright page or API client
          |
          v
Provider proxy endpoint
          |
          v
Destination website

The provider forwards traffic and presents its residential-network egress to the destination. The site can still return a block, challenge, error, or ordinary page. Do not use proxy routing to evade authentication, access controls, rate limits, CAPTCHAs, anti-bot measures, or other safety restrictions.

2. Decide whether proxy routing is appropriate

First check whether the site offers an official API, data export, or another authorized access method. A residential proxy may be relevant when an authorized workflow needs requests to exit through a particular residential network or location, or when a provider-supported session model fits the workflow. It is not a workaround for missing permission.

ResidentialProxy.io describes proxy use with scraping and browser automation. That establishes topical relevance, not a neutral performance result. The sources reviewed for this guide do not establish an independent provider ranking, response-time comparison, success-rate comparison, or current market-price comparison.

Before collecting data, identify the site owner’s rules, any account or contract restrictions, applicable privacy obligations, your organization’s policies, and the proxy provider’s acceptable-use terms. Ask for permission where needed. Narrow the dataset, reduce request volume, and stop if the owner or provider asks you to stop.

3. Check authorization before running a workflow

A proxy does not grant authorization. Infatica’s policy, for example, identifies some potentially legitimate business uses subject to agreement and applicable law, while prohibiting unauthorized access and attempts to evade access controls, authentication, rate limits, anti-bot systems, CAPTCHAs, or safety restrictions. Policies differ by provider; read the current terms for the service and target workflow.

The Internet Engineering Task Force’s RFC 9309 says of robots.txt: “These rules are not a form of access authorization.” Robots.txt communicates crawler preferences. It neither grants access nor replaces permission, access controls, contractual terms, or legal review. Amazon documents behavior for its named crawlers, including Amazonbot, Amzn-SearchBot, and Amzn-User; those Amazon-specific rules should not be generalized to other automated clients.

  • Prefer an official API or an explicit agreement when available.
  • Review robots.txt and site terms, while treating them as separate from authorization.
  • Do not try to bypass login, rate limits, CAPTCHAs, anti-bot systems, or safety controls.
  • Follow the provider’s acceptable-use policy and its IP sourcing and consent disclosures.
  • For a legally uncertain use or jurisdiction, get qualified advice before collecting.

4. Configure a residential proxy in Playwright

Playwright documents HTTP(S) and SOCKSv5 proxy support. It permits configuration for the whole browser or for an individual browser context, with optional username, password, and bypass hosts. Use the exact server URL and credential format supplied by your provider. Keep credentials in a secret store or environment variables; do not commit them to source control or print them in logs.

Python: browser-wide proxy configuration

Install Playwright and its browser once in your environment:

python -m pip install playwright
python -m playwright install chromium

Set PROXY_SERVER to the provider’s endpoint, such as its documented HTTP or SOCKS endpoint. Set credentials only if the provider requires them.

import os
from playwright.sync_api import sync_playwright

proxy = {
    "server": os.environ["PROXY_SERVER"],
}
if os.environ.get("PROXY_USERNAME"):
    proxy["username"] = os.environ["PROXY_USERNAME"]
if os.environ.get("PROXY_PASSWORD"):
    proxy["password"] = os.environ["PROXY_PASSWORD"]
if os.environ.get("PROXY_BYPASS"):
    proxy["bypass"] = os.environ["PROXY_BYPASS"]

with sync_playwright() as p:
    browser = p.chromium.launch(proxy=proxy, headless=True)
    page = browser.new_page()
    response = page.goto("https://example.com", wait_until="domcontentloaded", timeout=45000)
    print({"status": response.status if response else None, "title": page.title()})
    browser.close()

Save the script as capture.py, set the environment variables for your shell or secret manager, and run python capture.py. This example prints the HTTP response status when available and the page title; it does not assert that a particular status means the collection is authorized.

Python: context-level proxy configuration

Use a context-level setting when separate contexts need different provider endpoints or session credentials. Confirm context-level proxy support and the exact API syntax in the Playwright version you install.

import os
from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    context = browser.new_context(proxy={
        "server": os.environ["PROXY_SERVER"],
        "username": os.environ.get("PROXY_USERNAME"),
        "password": os.environ.get("PROXY_PASSWORD"),
        "bypass": os.environ.get("PROXY_BYPASS"),
    })
    page = context.new_page()
    page.goto("https://example.com", wait_until="domcontentloaded", timeout=45000)
    print(page.title())
    browser.close()

Omit optional keys when they are not needed; some providers use a username format to select a region or session, so do not assume the example credential fields alone express those controls. Use the provider’s instructions.

Node.js: browser-wide proxy configuration

Install the Playwright package and browser in your project:

npm install playwright
npx playwright install chromium
const { chromium } = require('playwright');

(async () => {
  const proxy = { server: process.env.PROXY_SERVER };
  if (process.env.PROXY_USERNAME) proxy.username = process.env.PROXY_USERNAME;
  if (process.env.PROXY_PASSWORD) proxy.password = process.env.PROXY_PASSWORD;
  if (process.env.PROXY_BYPASS) proxy.bypass = process.env.PROXY_BYPASS;

  const browser = await chromium.launch({ headless: true, proxy });
  try {
    const page = await browser.newPage();
    const response = await page.goto('https://example.com', {
      waitUntil: 'domcontentloaded',
      timeout: 45000,
    });
    console.log({ status: response ? response.status() : null, title: await page.title() });
  } finally {
    await browser.close();
  }
})().catch((error) => {
  console.error(error.message);
  process.exitCode = 1;
});

Node.js: context-level proxy configuration

const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch({ headless: true });
  try {
    const context = await browser.newContext({
      proxy: {
        server: process.env.PROXY_SERVER,
        username: process.env.PROXY_USERNAME,
        password: process.env.PROXY_PASSWORD,
        bypass: process.env.PROXY_BYPASS,
      },
    });
    const page = await context.newPage();
    await page.goto('https://example.com', {
      waitUntil: 'domcontentloaded',
      timeout: 45000,
    });
    console.log(await page.title());
  } finally {
    await browser.close();
  }
})().catch((error) => {
  console.error(error.message);
  process.exitCode = 1;
});

As with Python, remove optional proxy fields when unused and follow the provider’s session and location syntax. If multiple contexts run concurrently, keep their session identities and task state explicit.

cURL: route an HTTP request through a proxy

For non-browser HTTP work, cURL can use a proxy endpoint. This does not create a browser session or reproduce JavaScript execution. The following form uses an environment variable so the credential is not written directly in the command:

curl --proxy "$PROXY_SERVER" --proxy-user "$PROXY_USERNAME:$PROXY_PASSWORD" \
  --fail --show-error --silent \
  https://example.com/ -o response.html

Adapt proxy flags to the provider’s documented scheme and authentication. Avoid shell history, process listings, and shared logs that expose credentials. Use an HTTP client rather than a browser when the authorized task only needs HTTP responses.

5. Choose rotating or sticky sessions

Session choice depends on whether the work needs continuity. Providers commonly describe rotation as useful for separate requests or workers that can operate independently, and sticky sessions as useful when a sequence of steps needs to retain one egress identity. This is provider guidance, not measured evidence that either mode avoids detection or improves success.

Workflow Session approach to evaluate Reason
Independent, authorized requests that do not share state Rotation may fit Each request or worker can be provisioned independently, if the provider supports the required behavior.
A multi-step flow whose state depends on continuity, such as an authorized cart or pagination task Sticky session may fit The workflow can keep using the same provider session for its steps.
Authenticated work tied to an account or explicit agreement Follow the site and provider’s approved session design Do not rotate or change identity to evade controls or account restrictions.

Confirm how a session is selected, how long it persists, what happens after a disconnect, and whether parallel requests share or replace session state. A rotation setting is a routing option, not permission to defeat a site’s controls.

6. Select a provider and deployment design

No independent comparative benchmark or provider winner was established in the research for this guide. Compare services against your authorized workload and verify current terms and capabilities before purchasing.

Criterion Questions to verify
Protocol and compatibility Does the service support the HTTP(S) or SOCKSv5 protocol your client needs? Does its documented setup match Playwright’s supported configuration?
Location targeting Which countries or regions are available, at what granularity, and under what plan terms? Verify current availability instead of relying on an old claim.
Session controls How are rotating and sticky sessions requested? What is the session lifetime, and what happens when a connection drops?
Authentication and account security What credential format is used? Can credentials be scoped or rotated? How should secrets be stored?
IP sourcing and consent How does the provider explain sourcing, participant consent, abuse remediation, and complaint handling? Read policies rather than relying on a label.
Acceptable use Does the provider permit the planned workload, and is that workload authorized by the site owner? What restrictions apply?
Operations and support Can you observe proxy errors and session behavior without logging secrets? Is support available for your deployment needs?
Total cost What is billed: traffic, ports, requests, time, or another unit? Include location, session, support, and overage terms in the estimate.

Choose the smallest authorized workflow that answers the business question. Record the target, permission basis, request scope, session policy, and stop conditions so operators can detect scope drift.

7. Keep browser workflows reliable and efficient

  • Use the least complex client that works. Playwright is useful when the authorized task depends on browser rendering or interaction. For static HTTP content, a direct HTTP client may use fewer resources.
  • Set explicit timeouts. Browser launch, navigation, and page readiness are separate operations. Bound each one and report which stage failed.
  • Choose readiness carefully. domcontentloaded can return before images or later page activity finish. Use a condition appropriate to the task, such as a documented selector, when the authorized output depends on it.
  • Limit concurrency. Match worker counts to the provider’s plan, the target’s permission and rate expectations, and your own memory and connection limits. More concurrency can increase failures and cost.
  • Reuse only the state that should persist. Browser contexts isolate cookies and storage. Keep context and proxy session lifetimes aligned with the workflow’s intended continuity.
  • Retry narrowly. Retry transient connection failures with bounded backoff. Do not retry in a loop against access denials, rate limits, challenges, or explicit stop signals.
  • Track outcomes. Record target, time, task identifier, status, latency, and a sanitized error category. Never log proxy passwords, authorization headers, or sensitive page data by default.
  • Measure total cost. Include provider billing units, browser compute, bandwidth, storage, support, and engineering time. Prices and billing rules can change; use current provider terms.

Residential routing can add a network hop and browser rendering adds its own work. The sources reviewed do not support numerical latency, uptime, or success-rate claims for providers. Measure your authorized workload in your own environment and compare results under the same target, region, session behavior, and browser configuration.

8. Troubleshoot common failures

Symptom Likely cause Fix
Proxy connection refused or timed out Wrong endpoint or port, network egress restriction, provider outage, or unavailable session. Check the endpoint and scheme against provider instructions, verify outbound connectivity, and inspect provider status or logs. Use a bounded retry only for transient errors.
Proxy authentication error Missing, expired, malformed, or incorrectly escaped credentials. Check the required auth format and credential status. Store secrets safely and avoid embedding special characters incorrectly in a URL.
Playwright rejects the proxy configuration Unsupported protocol spelling, malformed server URL, invalid option placement, or version mismatch. Use a documented HTTP(S) or SOCKSv5 endpoint and the installed Playwright API’s proxy configuration shape. Check the official Network guide.
Navigation times out Slow destination, proxy connection trouble, or a readiness condition that never completes. Identify whether launch, connection, navigation, or selector wait timed out. Choose a task-appropriate readiness condition and a bounded timeout. Stop if the target denies access.
Page is blank or incomplete Client-side rendering has not completed, a resource failed, or the destination returned an error page. Inspect the response status and sanitized browser errors. Wait for the specific authorized content condition. Do not interpret a blank page as permission to bypass controls.
Session changes during a multi-step task Rotation policy or session lifetime conflicts with workflow continuity. Review provider session selection and lifetime. Use a provider-supported sticky session only when continuity is required and allowed by the target.
Different workers appear to share state Contexts, cookies, or provider session identifiers are being reused unintentionally. Make context and session ownership explicit. Isolate independent tasks and share state only where the authorized workflow requires it.
Requests are denied, challenged, or rate limited The site is enforcing access or traffic rules. Stop automated attempts, reduce scope, request permission, or use an official access method. Do not rotate proxies to evade the restriction.

9. Or skip the browser setup

If the task is to capture a website screenshot or PDF, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request returns an image or PDF. Its clean capture flow accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses report the page verdict and billing status.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for request options. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. A screenshot API is for rendering captures, not a substitute for authorization to collect data.

Sign up free for 1,000 screenshots a month, with no card required.

10. Quick decision checklist

  • Is the collection authorized by the site owner and consistent with applicable rules?
  • Have you checked official APIs, site terms, robots.txt preferences, and provider acceptable-use terms?
  • Does the task need browser rendering, or will a direct HTTP client suffice?
  • Does the workflow need independent egress per task or continuity across steps?
  • Have you verified endpoint protocol, location availability, session controls, credential format, sourcing disclosures, and current billing terms?
  • Are timeouts, concurrency, retries, logging, and stop conditions bounded and documented?

11. Frequently asked questions

How do I use a residential proxy for web scraping?

Use an authorized access method, obtain a provider endpoint and credentials, configure your HTTP client or browser to use that endpoint, and keep request scope and rate within the site’s and provider’s rules. A proxy only changes routing.

How do I use a residential proxy with Playwright?

Set the provider’s supported HTTP(S) or SOCKSv5 server in the browser launch or browser context proxy options. Add credentials and bypass hosts only as required by the provider. Keep credentials secret.

Should I use rotating or sticky proxies with browser automation?

Evaluate rotation for independent authorized tasks and sticky sessions for multi-step work that needs continuity. The distinction is provider guidance, not a guarantee of better access or performance.

That depends on the facts, jurisdiction, target, data, permission, and terms. A proxy does not settle the question. Robots.txt is a crawl preference mechanism, not authorization. Get permission and qualified legal advice where needed.

Does a residential proxy make a scraper anonymous?

No. It changes the network egress path, but does not guarantee anonymity or hide all identifying signals. Do not use it to conceal unauthorized activity.

Can I use a proxy to get around a CAPTCHA or rate limit?

No. Treat a challenge, denial, or limit as a stop condition. Seek permission or an official access route instead of attempting to evade it.