ScreenshotNeo

BlogHow-to

How to Customize Web Scraping API Requests

Learn how to add headers, cookies, JavaScript rendering, proxies, waits, extraction, retries, caching and validation to scraping API requests.

By the ScreenshotNeo team1 October 20267 min read

Customize a scraping API request incrementally: authenticate on your server, send the target URL, add only the headers or cookies the target needs, enable JavaScript when content is client-rendered, choose proxy and country settings for access or localization, wait for dynamic content, and request the smallest useful output format.

This approach makes failures diagnosable. Start with one working request, then add one control at a time. Keep API credentials server-side, validate the returned data, and record the settings that affect reproducibility.

1. Start with the smallest working request

Most providers require an API key and a URL. The exact parameter names differ, but the basic shape is:

GET https://provider.example/scrape?api_key=SERVER_SIDE_SECRET&url=https%3A%2F%2Fexample.com

Use URL encoding for the target URL and never place secrets in browser JavaScript, public repositories, screenshots, shared notebooks or logs.

cURL baseline

curl -G 'https://provider.example/scrape' \
  --data-urlencode 'api_key=YOUR_SERVER_SIDE_KEY' \
  --data-urlencode 'url=https://example.com' \
  -o response.html

Python baseline

import os
import requests

params = {
    "api_key": os.environ["SCRAPING_API_KEY"],
    "url": "https://example.com",
}
response = requests.get("https://provider.example/scrape", params=params, timeout=90)
response.raise_for_status()
open("response.html", "wb").write(response.content)

Node.js baseline

const key = process.env.SCRAPING_API_KEY;
const q = new URLSearchParams({
  api_key: key,
  url: 'https://example.com'
});

const res = await fetch(`https://provider.example/scrape?${q}`);
if (!res.ok) throw new Error(`Scraper returned ${res.status}`);
await Bun.write('response.html', await res.text());

2. Add custom headers deliberately

Headers are useful when the target depends on a particular user agent, language, authorization value, referer or accepted response type. Provider schemas vary: some accept a JSON object in a POST body, while others expose individual query parameters.

curl -X POST 'https://provider.example/scrape' \
  -H 'Content-Type: application/json' \
  -d '{
    "api_key": "YOUR_SERVER_SIDE_KEY",
    "url": "https://example.com/account",
    "headers": {
      "User-Agent": "ExampleMonitor/1.0",
      "Accept-Language": "en-US,en;q=0.9",
      "Authorization": "Bearer TARGET_TOKEN",
      "Referer": "https://example.com/"
    }
  }'

Send only headers required by the workflow. Copying every browser header can make requests brittle and can expose credentials in logs. Redact authorization and cookie values before recording request diagnostics.

3. Send cookies and authenticated context

Cookies reproduce a consent choice, login session, region preference or feature flag. Prefer a provider’s documented cookie structure, commonly a name/value map or an array of cookie objects.

curl -X POST 'https://provider.example/scrape' \
  -H 'Content-Type: application/json' \
  -d '{
    "api_key": "YOUR_SERVER_SIDE_KEY",
    "url": "https://example.com/dashboard",
    "cookies": {
      "session_id": "REDACTED_SESSION",
      "locale": "en-US"
    }
  }'

Use short-lived credentials where possible. Do not pass a session cookie to a different host, and make sure your use complies with the target’s terms and applicable law.

4. Decide whether JavaScript rendering is required

Use a plain HTTP request when the data is present in the initial HTML. Enable rendering for single-page applications or pages that populate content after JavaScript executes. Providers expose different flags, such as dynamic=true, render=true or render_js=1.

curl -G 'https://provider.example/scrape' \
  --data-urlencode 'api_key=YOUR_SERVER_SIDE_KEY' \
  --data-urlencode 'url=https://example.com/catalog' \
  --data 'render=true'

Rendering alone may return before the content appears. Pair it with a selector wait when the provider supports one:

curl -G 'https://provider.example/scrape' \
  --data-urlencode 'api_key=YOUR_SERVER_SIDE_KEY' \
  --data-urlencode 'url=https://example.com/catalog' \
  --data 'render=true' \
  --data-urlencode 'wait_for_selector=.product-card'

Use a bounded delay only when no stable selector exists. A selector tied to the required content is usually more reliable than an arbitrary sleep.

5. Choose proxy type, country and session behavior

Setting Use it when Trade-off
Datacenter proxy Ordinary public pages work from server IPs Usually simpler and faster
Residential proxy The target requires a consumer-network origin Often slower or more expensive
Mobile proxy The workflow specifically depends on mobile-network identity Use only when necessary
Country or geo code Content varies by market, language, inventory or legal availability Results differ by location
Sticky session Several requests must appear to come from one client Session state must be managed and expired
curl -G 'https://provider.example/scrape' \
  --data-urlencode 'api_key=YOUR_SERVER_SIDE_KEY' \
  --data-urlencode 'url=https://example.com/pricing' \
  --data 'proxy_type=residential' \
  --data 'country=us' \
  --data 'session_number=42'

Record proxy type, country and session identifier with the result. This makes localized or multi-step captures reproducible.

6. Select output and extraction format

Request the smallest response your pipeline needs. Depending on the provider, you may be able to return raw HTML, links, Markdown, images, summaries or parsed JSON using extraction rules.

curl -X POST 'https://provider.example/scrape' \
  -H 'Content-Type: application/json' \
  -d '{
    "api_key": "YOUR_SERVER_SIDE_KEY",
    "url": "https://example.com/products",
    "output": "json",
    "extract": {
      "name": ".product-name",
      "price": ".price",
      "url": "a@href"
    }
  }'

Define required fields and validate them after every response. HTTP 200 only means the provider returned successfully; it does not prove that the intended page state or fields were captured. Preserve the raw response for debugging when storage and policy allow.

7. Combine options safely

Add controls in this order:

  1. Authentication and URL.
  2. Headers or cookies required to reach the page.
  3. JavaScript rendering.
  4. A selector wait or bounded delay.
  5. Proxy type, country and session.
  6. Extraction and output format.

After each change, compare status, response length, required fields and timing. This isolates the option that introduced a failure.

8. Reliability, retries and caching

Managed scraping APIs may rotate proxies, retry blocked requests, solve CAPTCHA challenges or render with a headless browser. Your client should still implement bounded retries for transient failures.

import time
import requests

retryable = {408, 425, 429, 500, 502, 503, 504}
for attempt in range(4):
    response = requests.get(
        "https://provider.example/scrape",
        params={"api_key": "YOUR_SERVER_SIDE_KEY", "url": "https://example.com"},
        timeout=90,
    )
    if response.status_code not in retryable:
        response.raise_for_status()
        break
    if attempt == 3:
        response.raise_for_status()
    time.sleep(2 ** attempt)

Cache idempotent requests when freshness permits. Include the URL, render mode, country, session, headers that affect content and extraction rules in the cache key. Provider cache controls and rate limits are specific to each service and can change; verify current documentation before setting capacity assumptions.

9. Security and observability checklist

  • Keep API keys in environment variables or a secret manager.
  • Redact authorization headers, cookies and tokens in logs.
  • Set connect and total timeouts.
  • Record provider status, latency, retry count, render mode, proxy country and cache state.
  • Validate required fields and content type.
  • Limit concurrency to the provider’s documented rate limit.
  • Respect the target site’s terms, robots guidance and applicable law.

10. Troubleshooting common errors

Symptom Likely cause Fix
401 or 403 from provider Missing, invalid or exposed API key Rotate the key, send it server-side and verify the provider’s authentication parameter.
Target returns a login page Missing cookies or authorization header Send the required session context and confirm it has not expired.
HTML lacks visible content Content is client-rendered Enable JavaScript rendering and wait for a content selector.
Selector wait times out Selector is unstable, incorrect or blocked by a consent layer Inspect the HTML, use a stable selector, handle consent and set a bounded fallback delay.
Wrong language or inventory Request exits from the wrong region Set country or geo code and send an appropriate Accept-Language header.
Intermittent blocks IP reputation, request volume or missing browser context Reduce concurrency, use the documented proxy tier, add only necessary headers and use bounded retries.
HTTP 200 but empty extraction Wrong page state or changed markup Validate required fields, capture the raw response and update extraction rules.
High cost or latency Rendering, premium proxies or repeated uncached requests Use static fetching where possible, selector waits instead of long delays, appropriate caching and the least expensive proxy that works.

11. Performance and cost decisions

  • Static requests are generally simpler and cheaper than browser rendering.
  • Rendering plus premium residential routing can consume provider-specific credits; for example, Scrapingdog documents different credit use for dynamic requests and premium residential proxies.
  • Selector waits reduce wasted browser time compared with an unnecessarily long fixed delay.
  • Batch independent URLs only when the provider supports it and your rate limit allows it.
  • Cache stable pages, but include every content-affecting option in the cache key.

12. Or skip the browser setup

If your goal is a clean visual capture rather than extracted HTML, ScreenshotNeo provides a single screenshot API request. Its cookie and consent handling accepts the banner like a visitor, then removes more than 60 known consent platforms, newsletter popups and chat widgets before capture. Failed loads, bot checks, blank pages, timeouts and cache hits are not billed.

See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

You can configure full-page or element capture, dark mode, device presets, viewport and retina scale, PDF output, custom CSS and JavaScript, clicks, waits, blocked resources, headers, cookies, user agent, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous jobs, bulk capture and usage reporting. An MCP server provides take_screenshot, get_page_info and capture_pdf tools for AI agents.

Every response includes X-Page-Verdict and X-Billed headers so you can see what happened. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

13. FAQ

Should I use headers or cookies first?

Use whichever the target workflow requires, and add one at a time. Headers identify the request; cookies carry browser state.

Is a proxy required for every request?

No. Start with a datacenter route for ordinary public pages. Move to residential or mobile routing only when access requires it.

How do I keep the same IP?

Use the provider’s sticky-session or reusable session option and keep its identifier across related requests.

Why does a successful response contain the wrong page?

The request may have captured a login, consent, challenge or pre-render state. Check cookies, rendering, waits, location and required-field validation.

When should I request JSON?

Use structured extraction when the fields and selectors are stable. Keep raw HTML available for audits and parser changes.