How to Use a Python Client for Web Scraping APIs
Learn how Python API clients authenticate, request pages, handle errors, retries, parsing, screenshots, and production limits.
Direct answer: use the provider’s maintained Python client when it supports your runtime and output needs. Install the documented package, load the API key from an environment variable, send the smallest request that answers your task, check the HTTP status and response body before parsing, then add bounded retries, timeouts, logging without secrets, and rate controls.
There is no universal Python interface for scraping APIs. ScrapingBee, Apify, and Zyte use different packages, authentication methods, parameters, and response formats. Treat each SDK as a wrapper around one provider’s API and verify the installed version’s documentation before deploying.
1. Define the data you need first
Write down the target pages, fields, output format, and freshness requirement before choosing a client. A request for ordinary HTML usually needs less configuration than a JavaScript-rendered page, a screenshot, or structured extraction.
- Raw HTML: suitable for server-rendered pages and your own parser.
- Rendered HTML: needed when important content appears after JavaScript runs.
- Structured extraction: useful when the provider returns selected fields instead of a full document.
- Screenshot or PDF: choose a capture API rather than parsing HTML.
Confirm that your use of the target site is permitted by applicable law, contracts, robots policies, and the site’s terms. The provider documentation does not decide those questions for your jurisdiction.
2. Install and pin the client
Install the package in the same dependency workflow as the rest of your application, then pin a version after reviewing its release notes.
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install scrapingbee
# or, for Apify:
python -m pip install apify-client
Apify documents apify-client as its official Python REST API client and requires Python 3.11 or newer. It provides synchronous and asynchronous interfaces for resources such as Actors, Datasets, and key-value stores. See the Apify Python client documentation.
3. Keep credentials out of code and logs
Use an environment variable or your secret manager. Do not commit keys, put them in a notebook shared publicly, include them in URLs, or print request headers.
export SCRAPINGBEE_API_KEY='replace-me'
export APIFY_API_TOKEN='replace-me'
export ZYTE_API_KEY='replace-me'
import os
scrapingbee_key = os.environ['SCRAPINGBEE_API_KEY']
# Fail early with a useful message in applications:
if not scrapingbee_key:
raise RuntimeError('SCRAPINGBEE_API_KEY is not configured')
4. Make a minimal request with ScrapingBee
ScrapingBee’s tutorial shows a client object, a target URL, a parameter dictionary, and an explicit response.ok check before using the content.
import os
from scrapingbee import ScrapingBeeClient
client = ScrapingBeeClient(api_key=os.environ['SCRAPINGBEE_API_KEY'])
response = client.get(
'https://example.com',
params={}
)
if response.ok:
print(response.status_code)
html = response.content
with open('page.html', 'wb') as output:
output.write(html)
else:
print(response.status_code, response.content)
This follows the vendor’s documented pattern; confirm method names and supported parameters against the version installed in your project. ScrapingBee’s Python SDK tutorial and HTML API documentation describe JavaScript rendering, proxy choices, forwarded headers, screenshots, and extraction options.
5. Use Apify’s synchronous or asynchronous client
Apify’s client exposes platform resources rather than one universal scrape method. The exact Actor, input schema, and returned dataset depend on the Actor you select, so read that Actor’s documentation before writing code.
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ['APIFY_API_TOKEN'])
# Replace the Actor ID and input keys with the values documented by that Actor.
run = client.actor('USERNAME/ACTOR-NAME').call(
run_input={
'startUrls': [{'url': 'https://example.com'}]
}
)
for item in client.dataset(run['defaultDatasetId']).iterate_items():
print(item)
For an async application, use the asynchronous client documented by Apify and await the same lifecycle operations. Apify’s default HTTP layer documents retries with exponential backoff for network errors, HTTP 429, and HTTP 5xx responses; do not assume another provider’s client has the same policy. See Apify HTTP client guidance.
6. Call a provider API directly with Python when no SDK fits
A direct HTTP request can be easier to audit and keeps your dependency surface small. Authentication is provider-specific: ScrapingBee recommends a Bearer header, while Zyte documents Basic authentication with the API key as the username and an empty password.
import os
import requests
url = 'https://example.com'
response = requests.get(
'https://api.example-provider.test/v1/extract',
params={'url': url},
headers={'Authorization': f"Bearer {os.environ['PROVIDER_API_KEY']}"},
timeout=(10, 60),
)
response.raise_for_status()
data = response.json()
print(data)
For Zyte, follow its documented Basic authentication and endpoint format instead of copying the Bearer example. The Zyte API reference is the source for its current request and response schema.
7. Parse only after validating the response
A successful connection does not prove that the provider returned complete or useful content. Check the status, content type, body size, and provider-specific error fields before parsing.
import requests
from bs4 import BeautifulSoup
response = requests.get(
'https://example.com',
timeout=(10, 60),
)
response.raise_for_status()
content_type = response.headers.get('content-type', '')
if 'html' not in content_type:
raise ValueError(f'Expected HTML, received {content_type!r}')
soup = BeautifulSoup(response.content, 'html.parser')
title = soup.title.get_text(strip=True) if soup.title else None
print({'title': title})
Keep transport errors, provider errors, and target-page errors separate in your logs. For example, an authentication failure belongs to your API configuration; a target page that returns an error document is a content problem; a 429 is a rate-limit event that may be retryable.
8. Add timeouts, bounded retries, and rate control
Set a finite connect and read timeout for every request. Retry only failures that are plausibly transient, such as network interruptions, 429 responses, and selected 5xx responses. Use exponential backoff with jitter and a maximum attempt count.
import random
import time
import requests
RETRYABLE = {429, 500, 502, 503, 504}
def get_with_retries(url, *, headers=None, attempts=4):
for attempt in range(attempts):
try:
response = requests.get(
url,
headers=headers,
timeout=(10, 60),
)
except requests.RequestException:
if attempt == attempts - 1:
raise
delay = min(30, 2 ** attempt) + random.random()
time.sleep(delay)
continue
if response.status_code not in RETRYABLE:
return response
if attempt == attempts - 1:
response.raise_for_status()
delay = min(30, 2 ** attempt) + random.random()
time.sleep(delay)
raise RuntimeError('retry loop ended unexpectedly')
response = get_with_retries('https://example.com')
response.raise_for_status()
Respect the provider’s documented limits and the target site’s capacity. A retry loop without a cap can multiply load and cost. Add a queue or token bucket when many workers share one API key.
9. Choose rendering and proxy options deliberately
- Start without JavaScript rendering when server HTML contains the fields you need.
- Enable rendering only when scripts create the required content.
- Use premium or geographic proxies only when the target and provider documentation justify them.
- Forward headers or cookies only when required, and redact them from logs.
- Request extraction fields instead of full pages when that reduces transfer and parsing work.
These options are provider-specific and can change usage, latency, and price. Recheck current documentation before relying on a parameter.
10. Save binary results safely
For screenshots, PDFs, or other binary responses, check response.ok before writing bytes. A JSON error saved as .png is difficult to diagnose later.
from pathlib import Path
import requests
response = requests.get(
'https://api.example-provider.test/v1/screenshot',
params={'url': 'https://example.com'},
timeout=(10, 90),
)
if not response.ok:
raise RuntimeError(f'{response.status_code}: {response.text[:500]}')
Path('page.png').write_bytes(response.content)
11. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off.
Use the API directly from Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for the request options. Failed bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and whether the request was billed. ScreenshotNeo also supports full-page and element captures, dark mode, device presets, retina scale, PDF settings, custom CSS and JavaScript, clicks, waits, blocking rules, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous jobs, bulk capture, usage reporting, and an MCP server for AI agents.
There are 1,000 free screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots, and every feature is available on every plan. Create a free ScreenshotNeo account.
12. Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| 401 or 403 from the provider | Missing, expired, or incorrectly formatted credentials | Load the key from the environment and follow that provider’s required Bearer, Basic, or SDK authentication format. |
| 429 Too Many Requests | Rate or concurrency limit | Reduce concurrency, honor retry headers when documented, and use bounded exponential backoff. |
| 5xx response | Transient provider or upstream failure | Retry a limited number of times, then record the request ID and response body for support. |
| Timeout | Slow target, JavaScript work, proxy delay, or an overly small timeout | Set separate connect/read timeouts, reduce unnecessary options, and retry only when appropriate. |
| HTML contains no expected data | Content is rendered client-side, blocked, or returned a challenge page | Inspect the raw response, enable the provider’s documented rendering option, or handle the challenge as an unsuccessful scrape. |
| JSON decode error | You received HTML, an empty body, or a provider error document | Check status and content type before calling .json(). |
| Screenshot file is unreadable | An API error was written as binary output | Check response.ok and save the error body separately. |
| Apify import or runtime error | Unsupported Python version or package mismatch | Use Python 3.11+ as documented and pin a compatible apify-client version. |
13. Performance, reliability, and cost considerations
- Latency: browser rendering, premium proxies, geographic routing, and large pages add work. Measure each option on representative URLs.
- Throughput: use bounded concurrency that stays within provider and target limits. Queue jobs when bursts exceed those limits.
- Retries: retry transient transport, 429, and selected 5xx failures; never retry authentication errors indefinitely.
- Caching: cache pages when freshness allows it. Include the URL and relevant options in the cache key.
- Cost: verify current quotas and prices directly with each provider. Rendering, proxy modes, extraction, and retries may affect usage.
- Observability: log provider, status, elapsed time, target hostname, attempt count, and request IDs while excluding API keys, cookies, authorization headers, and page secrets.
14. Production checklist
- List the exact pages and fields required.
- Confirm permitted use of each target.
- Choose an SDK or direct HTTP API that supports your Python version and output.
- Pin the package version and read its current documentation.
- Load credentials from runtime configuration.
- Start with the smallest request and inspect its raw response.
- Add finite timeouts, bounded retries, backoff, and rate limits.
- Separate provider failures from target-content failures.
- Redact secrets from logs and error reports.
- Recheck limits, pricing, and parameters before launch.
FAQ
Do all scraping APIs have the same Python interface?
No. Package names, method signatures, authentication, parameters, and response formats differ by provider.
Should I always use a provider SDK?
Use the maintained SDK when it matches your runtime and needs. Direct HTTP can be preferable when you need a small dependency surface or an endpoint the SDK does not expose.
When do I need JavaScript rendering?
Use it when the required content is created after page load. Verify by inspecting the returned HTML before enabling it by default.
Are retries guaranteed to make scraping reliable?
No. Retry behavior is client-specific and cannot fix invalid credentials, permanent blocks, changed page structure, or prohibited access.
Can ScreenshotNeo return scraped HTML?
ScreenshotNeo is designed for screenshots and PDFs. Use a content-extraction API when your output is HTML or structured fields.


