Automatic Configuration for Web Scraping
Learn how automatic scraping configuration escalates rendering and proxies, controls cost, handles failures, and reduces scraper maintenance.
Automatic configuration for web scraping means letting a hosted scraping service choose the least powerful request setup that can successfully retrieve a target page. The service starts with a simple HTTP request, then escalates to JavaScript rendering or more advanced proxies only when the cheaper configuration fails.
This approach turns scraper configuration into a controlled escalation strategy. You spend less time guessing whether a site needs a browser or a premium proxy, while still retaining explicit controls for cookies, headers, country, waits, custom JavaScript, extraction, and maximum cost.
How automatic configuration works
Most automatic systems make decisions across three dimensions:
| Dimension | What changes | Why it matters |
|---|---|---|
| JavaScript rendering | Plain HTTP retrieval can escalate to a real browser | Required when content is inserted after page scripts run |
| Proxy sophistication | Rotating proxies can escalate to premium or stealth proxies | Useful when a site blocks shared or datacenter traffic |
| Retry and escalation policy | The service tries configurations from inexpensive to advanced | Prevents every request from paying the highest tier |
A typical sequence is:
- Try a rotating proxy without JavaScript.
- If the response is unsuccessful, try JavaScript rendering with the rotating proxy.
- If needed, try a premium proxy with or without rendering.
- Use stealth browser infrastructure only when the configured cost limit permits it.
- Stop at the first successful response and charge for that successful tier.
Automatic configuration does not mean every scraper setting becomes automatic. Headers, cookies, authentication, country targeting, wait conditions, custom JavaScript, extraction schemas, and request bodies commonly remain caller-controlled.
ScrapingBee Auto Mode: controls and limits
ScrapingBee describes Auto Mode as starting with the simplest and least expensive configuration and gradually trying more advanced options until a request succeeds. It automatically chooses JavaScript rendering, premium proxies, or stealth proxies.
Auto Mode currently supports GET requests only. Do not combine it with render_js, premium_proxy, stealth_proxy, or transparent_status_code; those combinations return HTTP 400 responses. Cookies, custom headers, country selection, waiting instructions, JavaScript scenarios, and AI extraction can still be supplied explicitly.
ScrapingBee’s published credit tiers
| Configuration | Credits |
|---|---|
| Rotating proxy, no JavaScript | 1 |
| Rotating proxy with JavaScript | 5 |
| Premium proxy, no JavaScript | 10 |
| Premium proxy with JavaScript | 25 |
| Stealth proxy with JavaScript | 75 |
The max_cost setting limits the tiers Auto Mode may try. For example, a maximum of 25 credits excludes the 75-credit stealth tier. ScrapingBee says billing applies to the configuration that successfully returns the page rather than every failed attempt. AI extraction is charged separately.
Building automatic configuration yourself
If your provider does not offer Auto Mode, implement the same pattern as an ordered policy. Keep the policy data-driven so you can change cost limits without rewriting retry logic.
Python reference implementation
import os
import requests
ENDPOINT = os.environ["SCRAPER_ENDPOINT"]
API_KEY = os.environ["SCRAPER_API_KEY"]
TARGET_URL = "https://example.com/catalog"
TIERS = [
{"name": "rotating-http", "render_js": False, "proxy": "rotating", "cost": 1},
{"name": "rotating-browser", "render_js": True, "proxy": "rotating", "cost": 5},
{"name": "premium-http", "render_js": False, "proxy": "premium", "cost": 10},
{"name": "premium-browser", "render_js": True, "proxy": "premium", "cost": 25},
{"name": "stealth-browser", "render_js": True, "proxy": "stealth", "cost": 75},
]
MAX_COST = int(os.getenv("SCRAPER_MAX_COST", "25"))
for tier in TIERS:
if tier["cost"] > MAX_COST:
continue
params = {
"api_key": API_KEY,
"url": TARGET_URL,
"render_js": str(tier["render_js"]).lower(),
"proxy": tier["proxy"],
}
response = requests.get(ENDPOINT, params=params, timeout=90)
if response.ok and response.content:
print(f"Succeeded with {tier['name']} ({tier['cost']} credits)")
open("response.html", "wb").write(response.content)
break
else:
raise RuntimeError("No permitted scraping tier returned a usable response")
Replace SCRAPER_ENDPOINT and parameter names with those documented by your provider. The important behavior is the order, the success test, and the cost ceiling.
cURL pattern
curl --fail-with-body -G "$SCRAPER_ENDPOINT" \
--data-urlencode "api_key=$SCRAPER_API_KEY" \
--data-urlencode "url=https://example.com/catalog" \
--data-urlencode "render_js=false" \
--data-urlencode "proxy=rotating" \
-o response.html
Run the same command again with the next tier only when the first response fails your success checks. A HTTP 200 alone is not enough: verify that the body contains the expected content and is not a block page.
Node.js pattern
const endpoint = process.env.SCRAPER_ENDPOINT;
const apiKey = process.env.SCRAPER_API_KEY;
const target = 'https://example.com/catalog';
const tiers = [
{ name: 'rotating-http', render_js: false, proxy: 'rotating', cost: 1 },
{ name: 'rotating-browser', render_js: true, proxy: 'rotating', cost: 5 },
{ name: 'premium-http', render_js: false, proxy: 'premium', cost: 10 },
{ name: 'premium-browser', render_js: true, proxy: 'premium', cost: 25 },
{ name: 'stealth-browser', render_js: true, proxy: 'stealth', cost: 75 }
];
const maxCost = Number(process.env.SCRAPER_MAX_COST || 25);
for (const tier of tiers) {
if (tier.cost > maxCost) continue;
const q = new URLSearchParams({
api_key: apiKey,
url: target,
render_js: String(tier.render_js),
proxy: tier.proxy
});
const res = await fetch(`${endpoint}?${q}`);
const body = await res.text();
if (res.ok && body.length > 0) {
console.log(`Succeeded with ${tier.name} (${tier.cost} credits)`);
await import('node:fs/promises').then(fs => fs.writeFile('response.html', body));
break;
}
}
Which settings should be automatic?
Good candidates for automatic selection
- Whether JavaScript rendering is necessary.
- Whether rotating, premium, or stealth proxy infrastructure is needed.
- Retry order and stopping after the first usable response.
- Cost-limited escalation through increasingly expensive tiers.
Settings that usually remain explicit
- Cookies and authentication: required for sessions and logged-in pages.
- Headers and user agent: needed when the origin expects a particular client profile.
- Country: required for regional pricing or localized content.
- Wait conditions: use a selector, delay, or network-idle rule for late content.
- Custom JavaScript: useful for clicking controls or changing application state.
- Extraction schema: define the fields and output format your pipeline requires.
- Request method and body: especially important for POST, GraphQL, and form submissions.
When JavaScript rendering is necessary
Start without a browser when the data is present in the initial HTML response. Escalate to rendering when the HTML contains an application shell but not the records you need, when content appears only after scripts execute, or when interaction is required.
Rendering costs more and is slower because it starts a browser. Use a specific wait condition instead of a long fixed delay when possible. If the provider supports network-idle waits, confirm that the page actually reaches idle; analytics and streaming connections can keep a page busy indefinitely.
When premium or stealth proxies are necessary
Proxy escalation is appropriate when the target blocks the initial network identity, presents a challenge page, or returns different content to datacenter traffic. Do not automatically assume that stealth is required. A lower tier may work for most URLs, and a cost ceiling prevents rare difficult pages from consuming the most expensive tier.
Track the selected tier and the reason for escalation. This lets you identify a site that consistently needs browser rendering or a stronger proxy and decide whether a site-specific policy is cheaper than repeated automatic retries.
Cost, performance, and reliability
| Concern | Practical rule |
|---|---|
| Cost | Set a maximum tier; charge only after a usable response when the provider supports that model. |
| Latency | Order cheap, fast HTTP attempts before browser attempts and avoid unnecessary fixed waits. |
| Reliability | Validate body content, status, and block-page indicators instead of trusting HTTP 200. |
| Repeat traffic | Cache successful responses where freshness requirements allow it. |
| Observability | Log URL, selected tier, response status, elapsed time, and failure reason. |
| Scaling | Use queues and bounded concurrency so browser retries do not overwhelm your worker pool. |
Automatic configuration improves reliability when failures are caused by a missing capability. It cannot make an inaccessible, unauthorized, or legally restricted page collectible. Apply the target site’s restrictions, authentication rules, privacy requirements, and applicable law before collecting data.
Automatic configuration with Zyte API
Zyte presents automatic configuration as a default workflow for easy and difficult websites. Its API selects a lean proxy and technique set, manages rotation, rendering, extraction, and ban handling, and lets users override machine-selected decisions. Zyte also describes adaptive handling of layout and structure changes to reduce manual scraper maintenance.
When comparing this model with a manually scripted escalation loop, evaluate JavaScript-heavy and protected-site success, cost predictability, browser rendering, extraction quality, GET and POST support, sessions, headers, cookies, country targeting, observability, and compliance controls. Zyte notes that login restrictions, PII and copyrighted-data exclusions, KYC requirements for higher-trust infrastructure, and website restrictions remain part of the customer’s responsibility.
Or skip the browser setup
If your task is to save a visual record of a page rather than extract structured fields, ScreenshotNeo provides a single screenshot API request. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify the page verdict and billing state.
Use the ScreenshotNeo API documentation for all options. A basic request:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. It supports full-page and element capture, custom CSS and JavaScript, waits, blocking rules, headers, cookies, user agents, authorization, timezone, geolocation, caching, signed links, asynchronous jobs, webhooks, bulk capture, PDFs, HTML/CSS rendering, and image formats including PNG, JPEG, and WebP.
There are 1,000 free screenshots each month with no card. Paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Troubleshooting automatic scraping
HTTP 400 from Auto Mode
Cause: A manually selected option conflicts with automatic selection, such as render_js, premium_proxy, stealth_proxy, or transparent_status_code.
Fix: Remove the conflicting option and let Auto Mode choose it, or disable Auto Mode and configure the tier explicitly.
The response is HTTP 200 but contains no data
Cause: The server returned an application shell, challenge page, consent wall, or error document.
Fix: Validate expected selectors or fields, then retry with JavaScript rendering, an explicit wait, cookies, or a stronger proxy.
The page works in a browser but not through the first tier
Cause: Content is rendered client-side or the origin blocks the initial network identity.
Fix: Permit browser escalation and set a maximum cost that includes the tiers you are willing to use.
Every request reaches the expensive tier
Cause: The target may require a persistent session, country-specific traffic, authentication, or a browser interaction that Auto Mode cannot infer.
Fix: Add the required cookies, headers, country, waits, or JavaScript explicitly. Measure whether a site-specific configuration is more efficient.
Requests time out
Cause: Slow scripts, never-ending network connections, overloaded concurrency, or an overly long browser wait.
Fix: Use selector-based waits, cap concurrency, shorten timeouts, and record which tier and wait condition timed out.
POST or form submission is required
Cause: ScrapingBee Auto Mode supports GET only.
Fix: Use a provider and configuration that support the required method, or perform the request separately and apply automatic escalation to subsequent GET pages.
Implementation checklist
- Define what counts as a successful response.
- Order tiers from least expensive to most capable.
- Set a maximum cost before production traffic.
- Keep cookies, headers, country, waits, and custom scripts explicit.
- Log the selected configuration and escalation reason.
- Test JavaScript-heavy, blocked, empty, and slow pages.
- Respect authentication, privacy, copyright, and site restrictions.
- Cache results when freshness requirements allow.
FAQ
Does automatic configuration always use a browser?
No. It should begin with a plain HTTP request and render JavaScript only when a cheaper configuration fails.
Can automatic mode select a country?
Not necessarily. ScrapingBee’s Auto Mode chooses rendering and proxy sophistication; country selection remains an explicit option.
How do I prevent a difficult page from becoming too expensive?
Set a maximum cost such as ScrapingBee’s max_cost, and skip tiers above that limit.
Is automatic configuration the same as automatic extraction?
No. Configuration selects request infrastructure. Extraction still requires a parser, selector, schema, or provider feature.
Should I use automatic configuration for every URL?
Use it when site requirements vary or are unknown. For a stable site with known requirements, a fixed configuration can reduce latency and make costs easier to predict.


