How to Capture Websites That Block Apify with Proxies
Diagnose Apify proxy failures, choose a proxy group, and configure sessions or rotation in Actors and external clients.
When a website blocks an Apify crawler, first determine whether the proxy connection itself failed or the target site returned a block page. Test Apify Proxy, check the apparent client IP, and inspect the target response. Then choose an available proxy group and decide whether requests need a stable session or a changing IP. A proxy can change the route and IP characteristics; it cannot guarantee access through every protection mechanism.
For an Actor on Apify, configure the proxy through the Apify SDK and pass it to the crawler. For a client outside Apify, use the documented external proxy hostname, port, username parameters, and proxy password. Keep credentials private: Apify says the HTTP proxy password is sent unencrypted by the HTTP protocol and proxy use is charged to your account. See the Apify Proxy documentation before deploying.
1. Diagnose whether Apify Proxy or the target site is failing
- Check the proxy connection. Open
http://proxy.apify.com/through the configured proxy. Apify documents this as a connection-status check. - Check the apparent IP. Request
https://api.apify.com/v2/browser-info/through that same proxy. This helps confirm that the request is using a proxy and lets you inspect rotation. - Inspect the target response. Record the status code, response body, final URL, and (for browser crawling) a screenshot. A CAPTCHA or challenge page from the target means the proxy may be working even though the page is not usable.
- Compare direct and proxied requests only where authorized. A difference helps isolate a route or IP-related block; it does not establish permission to access the content.
- Review the Actor logs and proxy error code. Separate authentication, DNS, connection, and upstream errors before changing proxy groups.
Apify’s documentation describes Proxy as a way to rotate IP addresses when scraping to avoid geographic blocking. That description is not a guarantee that a particular target will allow a request.
Interpret Apify Proxy’s 590–599 codes
| Code | Documented meaning | What to check |
|---|---|---|
| 590 | Upstream returned a non-200 status | Inspect the upstream response and target behavior. |
| 592 | Upstream returned a status outside the supported range | Check the upstream server and response handling. |
| 593 | DNS lookup failed | Check the target hostname and proxy configuration. |
| 594 | Connection refused | Check the destination host, port, and whether the service accepts connections. |
| 595 | Connection reset | Check for transient network loss or timeouts; retry conservatively. |
| 596 | Broken pipe | Check whether the connection closed before the request finished. |
| 597 | Upstream authentication failed | Verify the upstream credentials and secret configuration. |
| 599 | Generic upstream error | Use logs and a small repeatable request to narrow down the cause. |
These errors are distinct from an ordinary target-site denial or challenge page. Use the proxy troubleshooting reference when interpreting codes.
2. Configure an Apify Actor with a proxy
For an Actor running on Apify, use the SDK proxy configuration path. The SDK checks access and can load the proxy password from the Actor environment. Supply a group and, when needed, a country code; then pass the returned configuration to your crawler. Group access depends on your account.
JavaScript Actor with CheerioCrawler
This example uses Apify SDK and Crawlee. Install the packages in your Actor project as appropriate for your project’s SDK version. Set TARGET_URL in the Actor environment or replace the example URL.
import { Actor } from 'apify';
import { CheerioCrawler } from 'crawlee';
await Actor.init();
try {
const input = (await Actor.getInput()) ?? {};
const targetUrl = input.url ?? 'https://example.com/';
const proxyConfiguration = await Actor.createProxyConfiguration({
// Omit groups to let Apify choose automatically, or specify a group
// your account can use, such as ['RESIDENTIAL'].
groups: input.proxyGroups,
countryCode: input.countryCode,
});
const crawler = new CheerioCrawler({
proxyConfiguration,
maxRequestsPerCrawl: 1,
requestHandler: async ({ request, response, body, proxyInfo, log }) => {
log.info('Fetched target', {
url: request.url,
statusCode: response?.statusCode,
proxyHost: proxyInfo?.hostname,
});
await Actor.pushData({
url: request.url,
statusCode: response?.statusCode,
html: body,
});
},
failedRequestHandler: async ({ request, error, log }) => {
log.error('Request failed', { url: request.url, message: error.message });
},
});
await crawler.run([targetUrl]);
} finally {
await Actor.exit();
}
Use the option names supported by the SDK/Crawlee version installed in your Actor. The JavaScript ProxyConfiguration reference describes group, country, custom proxy URL, and session-related behavior. In current crawler setups, a SessionPool can manage continuity and cookies; check the documentation for your installed major version because session APIs can change.
Python Actor with a proxy configuration
For an Apify Python Actor, create the configuration with the SDK and pass it to a supported crawler. The snippet shows the configuration step; wire it into the crawler class and request handler used by your project.
from apify import Actor
async def main():
async with Actor:
proxy_configuration = await Actor.create_proxy_configuration(
groups=['RESIDENTIAL'], # Use a group available to your account.
country_code='US', # Optional; choose the location your task needs.
)
proxy_url = await proxy_configuration.new_url()
# Pass proxy_configuration to your Apify/Crawlee crawler.
# Use proxy_url directly only when configuring a compatible HTTP client.
print('Proxy configuration is ready:', proxy_url.split('@')[-1])
if __name__ == '__main__':
import asyncio
asyncio.run(main())
For a Python crawler, use the Python ProxyConfiguration reference and the documentation for the crawler version you have installed. Avoid printing full proxy URLs: they can contain credentials.
Relevant Actor configuration choices
groups: select one or more proxy groups that your account can access. If omitted, Apify can select groups automatically. A documented group name for Unblocker isUNBLOCKER; confirm current availability and access in your account.countryCode: request a location when geography matters and the selected group supports it. Location availability depends on group and account.proxyUrls: use your own proxy URLs instead of Apify Proxy. This is a configuration option, not evidence that those proxies will work against a particular site.sessionId/ session management: keep requests associated with a session when continuity matters. For browser workflows, use the crawler’s session management and cookie persistence options where supported. Check your installed SDK version: session behavior differs between major versions.checkAccess: some SDK versions expose this option to skip the initialization access check. The check can add startup time; skipping it removes an early configuration validation step.
3. Choose a proxy group based on the failure
| Choice | Consider it when | Tradeoffs |
|---|---|---|
| Automatic or datacenter | You need ordinary proxy routing, speed, or relatively low cost. | Apify describes datacenter proxies as fast and relatively inexpensive, with health checks, rotation, country selection, persistent sessions, and shared or dedicated groups. Shared activity can contribute to blocks. |
| Residential | The task has a justified need for IPs assigned by ISPs to homes or offices. | Traffic-based pricing applies; speed can vary across host devices, and a session may end if a device disconnects. Check current prices, account terms, and group availability. |
| Unblocker | You want Apify’s managed routing for common protections and CAPTCHA challenges. | Apify documents the UNBLOCKER group and unit-based billing. It does not provide a session parameter. Current availability and pricing can change; verify them in official documentation or your account. |
| Your own proxy URLs | You already have proxy infrastructure or a provider and need to configure those URLs in an Actor or SDK. | Apify’s configuration does not validate that a third-party proxy is suitable for a specific target. You own its credentials, performance, and costs. |
Choose by observed response, continuity needs, geography, latency, traffic or unit cost, group access, and whether your crawler is a browser or makes individual HTTP requests. Proxy group characteristics and pricing can change; consult Apify’s current proxy documentation and account UI.
4. Decide whether to keep a session or rotate IPs
A session is useful when multiple requests need continuity, such as a workflow that carries cookies or maintains a login. Rotation is appropriate when changing IPs is part of the request strategy. Neither choice fixes unrelated failures such as invalid URLs, a broken selector, or a target challenge that checks signals beyond IP address.
- Keep continuity: use a consistent session and the crawler’s session/cookie management when the workflow depends on state.
- Change IPs: allow the proxy configuration to rotate or use a new session where supported. Do not assume a browser crawler changes IP on every page request; inspect its session and browser lifecycle.
- Unblocker: Apify’s documented Unblocker group does not support a session parameter, so it is not the choice for a workflow requiring that particular session control.
- Handle failures deliberately: retire a session after a clearly unusable response only if your crawler design and target rules support that behavior. Avoid rapid repeated retries against a challenge page.
Apify’s SDK guidance explains proxy configuration and SessionPool integration; consult the proxy management guide and the version-specific references before relying on a particular session API.
5. Connect external clients through Apify Proxy
External connections use proxy.apify.com:8000, a proxy username containing parameters such as group, session, or location, and the Apify Proxy password from your Console. Apify documents that external connection access requires a paid plan. Do not confuse the proxy password with your Apify account password. Store credentials in environment variables or a secret manager.
cURL: status and target checks
# Set these in your shell without committing them to source control.
export APIFY_PROXY_USER='auto'
export APIFY_PROXY_PASSWORD='YOUR_APIFY_PROXY_PASSWORD'
export PROXY="http://${APIFY_PROXY_USER}:${APIFY_PROXY_PASSWORD}@proxy.apify.com:8000"
# Check the proxy connection.
curl --proxy "$PROXY" http://proxy.apify.com/
# Check the apparent client IP through the proxy.
curl --proxy "$PROXY" https://api.apify.com/v2/browser-info/
# Request the target through the proxy.
curl --proxy "$PROXY" --dump-header response-headers.txt \
--output response.html --write-out '\nHTTP %{http_code}\n' \
https://example.com/
For a group or session, use the username parameter syntax documented by Apify and available to your account; do not guess a parameter format. The proxy password is sensitive, and the HTTP proxy protocol does not encrypt it in transit according to Apify.
Python requests: proxied GET
import os
import requests
from urllib.parse import quote
proxy_user = os.environ.get('APIFY_PROXY_USER', 'auto')
proxy_password = os.environ['APIFY_PROXY_PASSWORD']
# URL-encode credentials in case they contain reserved characters.
proxy = (
f"http://{quote(proxy_user, safe='')}:{quote(proxy_password, safe='')}"
"@proxy.apify.com:8000"
)
proxies = {'http': proxy, 'https': proxy}
for url in (
'http://proxy.apify.com/',
'https://api.apify.com/v2/browser-info/',
'https://example.com/',
):
response = requests.get(url, proxies=proxies, timeout=(15, 45))
print(url, response.status_code, response.url)
response.raise_for_status() if 'browser-info' in url else None
For production, handle target error statuses explicitly rather than treating every non-200 response as a proxy failure. If your credentials contain reserved characters, URL-encode them as shown. Avoid logging the assembled proxy URL.
Node.js: proxied requests with an HTTP agent
Node’s built-in fetch does not by itself accept a proxy URL in the same way across Node versions. Use a maintained proxy agent compatible with your runtime, install it through your project’s normal package manager, and keep credentials in environment variables.
import { ProxyAgent, fetch as proxyFetch } from 'undici';
const user = process.env.APIFY_PROXY_USER ?? 'auto';
const password = process.env.APIFY_PROXY_PASSWORD;
if (!password) throw new Error('Set APIFY_PROXY_PASSWORD');
const proxyUrl = new URL(`http://${encodeURIComponent(user)}:${encodeURIComponent(password)}@proxy.apify.com:8000`);
const dispatcher = new ProxyAgent(proxyUrl);
for (const url of [
'http://proxy.apify.com/',
'https://api.apify.com/v2/browser-info/',
'https://example.com/',
]) {
const response = await proxyFetch(url, { dispatcher, signal: AbortSignal.timeout(45000) });
console.log(url, response.status, response.url);
await response.body?.cancel();
}
await dispatcher.close();
Install and pin a compatible undici version for your Node runtime. If your application uses another HTTP client, follow that client’s proxy-agent configuration and TLS behavior. A browser crawler may require proxy configuration at browser launch or crawler level rather than setting a proxy for one fetch call.
6. Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Proxy status page cannot load | Proxy URL, port, credentials, network route, or account access is wrong. | Use the documented host and port; verify the proxy password and external-plan access. Test the status endpoint before the target. |
| 597 Auth Failed | Incorrect upstream credentials. | Check the upstream username/password and secret source. Do not use your Apify account password as the proxy password. |
| 593 Not Found | DNS lookup failed, potentially due to hostname or proxy configuration. | Check the destination hostname spelling and the proxy-chain/client configuration. |
| 594, 595, or 596 | Connection refused, reset, or broken pipe. | Check the destination port and service availability. Retry transient failures with bounded backoff. |
| 599 Upstream Error | Generic upstream failure. | Record time, target, request type, and error; test a simple known URL and consult Apify support if it persists. |
| Target returns 403, CAPTCHA, or challenge HTML | The target is responding, but its controls reject or challenge the request. | Confirm the page is not being mistaken for usable content. Review permission and target rules, reduce request pressure, and choose an appropriate group only if justified. No proxy guarantees success. |
| Direct requests work but proxied ones fail | Proxy authentication, account eligibility, location restrictions, or target response differs by route. | Compare status endpoint, browser-info response, and target status separately. Verify group and country access. |
| Actor starts slowly or fails at initialization | Proxy access validation fails or selected group/location is unavailable. | Verify account access and configuration. Only disable an SDK access check when you understand the startup-validation tradeoff and your installed version supports it. |
| Browser pages switch identity unexpectedly | Proxy/session lifecycle does not match the crawler’s browser lifecycle. | Inspect SessionPool, cookie persistence, and crawler settings; do not assume each navigation gets a new proxy IP. |
| Proxy credentials appear in logs or source | Full proxy URL was printed or committed. | Rotate exposed credentials and move secrets to environment variables or a secret store. Log only non-sensitive host and status data. |
7. Reliability, performance, and cost
- Measure the right stages: track proxy connection failures separately from target status, navigation timeout, content extraction failure, and challenge responses.
- Use bounded retries: retry transient network errors with a small limit and backoff. Do not retry a persistent block page in a tight loop.
- Control concurrency: begin conservatively and increase only while the target and your authorization permit it. More parallelism can raise load and does not ensure more successful pages.
- Choose a fit-for-purpose group: datacenter options are positioned as fast and relatively inexpensive; residential is traffic-priced and its speed can vary; Unblocker uses unit-based billing. Current costs and availability are account- and time-dependent, so check official Apify sources.
- Account for external traffic: Apify documents external proxy connections as requiring a paid plan. Actors can use the SDK path to connect through Apify infrastructure; consult the current docs for billing details.
- Keep useful diagnostics: save status code, elapsed time, final URL, proxy group/location where exposed, retry count, and a redacted error. Do not persist credentials or unnecessary personal data.
- Respect boundaries: a proxy does not provide permission to access content. Follow the site’s terms, applicable law, and your authorization. The available research does not assess any particular target’s policy.
8. Or skip the browser setup
If you only need a screenshot rather than a crawler, ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF; its cleanup steps accept cookie and consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture, and each step can be turned off. A screenshot API is not a proxy configuration for Apify, and it does not promise access to every protected page.
See the ScreenshotNeo API documentation for request options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));
ScreenshotNeo also provides an MCP server for AI agents, with take_screenshot, get_page_info, and capture_pdf. Bot checks/CAPTCHAs, blank pages, failed loads, timeouts, and cache hits are not billed; the response indicates the page verdict and billing state. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card required.
9. FAQ
Why is Apify blocked?
“Blocked” can mean a proxy connection error, a target denial, a CAPTCHA, or a page that loaded without the expected content. Check the proxy status and browser-info endpoints first, then inspect the target response.
Can I make every blocked site work by changing proxies?
No. IP reputation and rate limits are only part of site defenses. A different proxy group may change the outcome, but success is not guaranteed.
Can I use my own proxies with an Apify Actor?
Yes. The SDK supports customer-supplied proxy URLs. Configure them using the version-appropriate SDK options and verify them independently; configuration compatibility does not guarantee target access.
Should I use residential proxies for every block?
No. Select a group based on the observed failure, geography and continuity requirements, account access, performance, and cost. Residential proxies have variable speeds and traffic-based pricing.
Does a screenshot API replace Apify crawling?
No. ScreenshotNeo is useful when the output you need is a rendered screenshot or PDF. Apify Actors are for crawler workflows that fetch and process pages or datasets.


