ScreenshotNeo

BlogHow-to

How to Integrate Scrapy with a Web Scraping API

Keep Scrapy’s spider and parsing workflow while routing downloads through a web scraping API. Configure Zyte API, handle binary responses, and troubleshoot common integration issues.

By the ScreenshotNeo team29 September 202611 min read

How to Integrate Scrapy with a Web Scraping API

To integrate Scrapy with a web scraping API, keep your spider’s normal Request and callback flow, and connect the API at the download layer. For Zyte API, the documented modern route is to install scrapy-zyte-api, set ZYTE_API_KEY, and enable scrapy_zyte_api.Addon in project settings. Your selectors and item parsing can usually stay the same.

This guide walks through a runnable setup, then covers request metadata, binary bodies, compatibility, operations, and common errors. Scrapy Cloud is optional hosting; it is not required to use a request-level API. See the [Scrapy request and response documentation](https://doc.scrapy.org/en/master/topics/request-response.html), [Zyte API integration guide](https://docs.zyte.com/zyte-api/usage/integrations.html), and [Zyte FAQ](https://docs.zyte.com/scrapy-cloud/usage/faq.html).

1. Understand the integration point

Scrapy spiders yield Request objects. The downloader obtains responses, and Scrapy passes each Response to the callback you specified. The callback parses the response and may yield items or more requests. A managed scraping API belongs between the spider’s request and its response, so the spider can often continue to treat the result as a regular Scrapy response.

A scraping API fits at the downloader boundary while Scrapy callbacks continue to parse responses.
A scraping API fits at the downloader boundary while Scrapy callbacks continue to parse responses.

That separation is the main practical benefit: parsing stays in Scrapy, while request handling is delegated to the provider integration. A provider may handle browser rendering or other fetch requirements, but its exact behavior, supported response types, authentication, and limits are provider-specific. Do not assume all APIs offer the same integration.

2. Check the project before installing

  1. Check the Python and Scrapy versions used in your environment.
  2. Inspect your existing ADDONS, downloader middleware, request handlers, and Twisted reactor setup.
  3. Choose an API with a documented Scrapy integration, and follow that provider’s current compatibility requirements.
  4. Keep the API key outside committed source code in production. The provider documents how to configure the key, but your application’s secret-storage mechanism depends on your deployment.

Zyte documents scrapy-zyte-api as requiring Python 3.8+ and Scrapy 2.0.1+. Check the [package integration documentation](https://docs.zyte.com/zyte-api/usage/integrations.html) for current requirements before upgrading or deploying.

3. Install and configure Zyte API

In the active virtual environment, install the package:

python -m pip install scrapy-zyte-api

Provide the key through your environment. In a POSIX shell for a local session:

export ZYTE_API_KEY="your-zyte-api-key"

Then merge the add-on into the settings module used by your project. Do not replace an existing ADDONS declaration if it already contains settings:

# settings.py
import os

ADDONS = {
    "scrapy_zyte_api.Addon": 500,
}

ZYTE_API_KEY = os.environ["ZYTE_API_KEY"]

The environment variable approach above reads the secret when settings load and fails clearly if it is missing. Adapt this to your deployment’s secret management. Never paste a live key into a checked-in settings file.

Keep a normal spider

A basic spider still yields Scrapy requests and parses Scrapy responses:

import scrapy


class ExampleSpider(scrapy.Spider):
    name = "example"
    start_urls = ["https://example.com/"]

    def parse(self, response):
        yield {
            "url": response.url,
            "title": response.css("title::text").get(),
            "status": response.status,
        }

        for href in response.css("a::attr(href)").getall():
            yield response.follow(href, callback=self.parse)

Run it with the project’s normal command, for example:

scrapy crawl example -O items.jsonl

In Zyte’s transparent integration mode, ordinary Scrapy requests to text resources such as HTML and JSON can pass through without changing how the spider constructs those requests. Confirm the provider’s current behavior for the request types and settings you use.

4. Pass callback data and provider options correctly

Scrapy distinguishes data passed to the callback from data intended for components such as middleware. Use cb_kwargs for your own callback values, and reserve meta for middleware or extension settings. Scrapy documents this distinction in its [Request API](https://doc.scrapy.org/en/master/topics/request-response.html).

import scrapy


class ProductSpider(scrapy.Spider):
    name = "products"

    def start_requests(self):
        yield scrapy.Request(
            "https://example.com/products/42",
            callback=self.parse_product,
            cb_kwargs={"product_id": "42"},
        )

    def parse_product(self, response, product_id):
        yield {
            "product_id": product_id,
            "name": response.css("h1::text").get(),
        }

Provider integrations may define request-level options through metadata, headers, or another documented mechanism. Follow that provider’s reference exactly. Avoid inventing metadata keys or copying settings from another provider; Scrapy’s generic meta dictionary is a transport mechanism, not a universal web-scraping API schema.

5. Handle HTML, JSON, and binary responses

Test representative content types explicitly. HTML is normally parsed through selectors, while JSON can be decoded from a response body. Binary responses need special care because providers may represent or transport them differently.

For Zyte’s integration, the examples recommend explicitly requesting httpResponseBody when you need a binary response body. The provider notes this because regular binary response handling may change in a future package version. This is Zyte-specific guidance; it is not a rule for every scraping API.

import scrapy


class BinarySpider(scrapy.Spider):
    name = "binary_example"

    def start_requests(self):
        yield scrapy.Request(
            "https://example.com/file.pdf",
            callback=self.parse_file,
            meta={
                "zyte_api": {
                    "httpResponseBody": True,
                }
            },
        )

    def parse_file(self, response):
        self.logger.info(
            "Received %d bytes from %s",
            len(response.body),
            response.url,
        )
        # Store or process response.body according to the application.

Confirm the response format and supported request options in the [Zyte API examples](https://docs.zyte.com/zyte-api/usage/integrations.html). For a JSON endpoint, check the HTTP status and content type before decoding, and handle invalid JSON as an expected failure case rather than assuming every response is valid.

6. Custom middleware and non-Zyte APIs

Scrapy provides middleware hooks around requests and responses. Spider middleware processes responses sent to spiders and requests or items coming from callbacks; custom spider middleware is enabled with SPIDER_MIDDLEWARES. Downloader integrations operate at a different part of the request path, and provider-specific instructions determine how to register them. See Scrapy’s [middleware documentation](https://docs.scrapy.org/en/latest/topics/spider-middleware.html) and the relevant provider’s current integration guide.

If no maintained package exists, you may need to adapt requests at the downloader boundary or call the provider’s HTTP API and convert its result into Scrapy response objects. That approach requires deliberate handling of status codes, headers, redirects, errors, retries, and response bodies. It is more than adding an API key to a spider: preserve Scrapy’s expected request-response behavior and test callbacks against the resulting responses.

7. Test the integration before a full crawl

  1. Run a small spider against one HTML page, one JSON resource, and one binary URL if your workload needs it.
  2. Check status codes, response bodies, content types, redirects, and the spider’s parsed output.
  3. Trigger a missing-key or invalid-key case in a safe environment and confirm the error is visible in logs.
  4. Check retries and failures using representative URLs, then verify that repeated failures do not create an unbounded retry loop.
  5. Measure request volume, response sizes, memory use, and crawl duration under realistic load.
  6. Review delay, concurrency, and rate-limit behavior before increasing throughput.

A crawl that completes is not enough to establish correctness. Compare item counts and key fields with a known small sample, and retain enough request context in logs to identify which URL and callback produced an error.

8. Compatibility, performance, and reliability

Reactor and async behavior

Zyte’s migration notes warn that projects using a non-asyncio Twisted reactor may need changes. Deferred and Future handling can also require attention during migration. If your project has custom async code, verify it in the same environment and startup configuration you will deploy. See the [Zyte migration guidance](https://docs.zyte.com/zyte-api/usage/migration.html).

Memory and response size

Zyte documents that API response bodies are Base64 encoded and can increase memory use by 33–37%. This is a vendor-documented implementation overhead, not a general property of every API integration or an independent benchmark. Large bodies and high concurrency can compound memory pressure, so monitor peak memory with representative payloads.

Delay, concurrency, and rate limits

Zyte documents that its integration respects Scrapy’s DOWNLOAD_DELAY, and discusses concurrency and rate-limit considerations. Recheck politeness settings after switching: behavior can differ from prior middleware, and higher concurrency is not automatically faster if it triggers limits, increases retries, or overwhelms memory. Start with the existing crawl settings, observe outcomes, and tune based on provider limits and the target’s requirements.

Reliability and retries

Record API-side failures separately from parsing failures where the integration makes that possible. A successful fetch can still produce a changed or incomplete page, while a valid page can fail your selectors. Track status, response size, callback exceptions, retry counts, and item validation so those cases are distinguishable. Use bounded retries and make item persistence safe to repeat if a request is retried.

Cost

API pricing varies by provider, plan, request type, and billing rules. The cited integration documentation does not establish a universal cost per request, so estimate using the provider’s current pricing and your measured request volume. Include retries, pagination, and browser-rendered requests in the estimate; do not assume one yielded Scrapy request always equals one billable provider operation.

9. Scrapy Cloud is optional

Zyte API processes requests through the Scrapy integration. Scrapy Cloud is a separate deployment and job-running service. Zyte explicitly says the products can be used independently. A local Scrapy project or another hosting platform can use the API integration without adopting Scrapy Cloud. If deploying on Scrapy Cloud, use the credential for the product you are configuring: the cloud tutorial distinguishes a Scrapy Cloud API key from a Zyte API key. See the [Scrapy Cloud FAQ](https://docs.zyte.com/scrapy-cloud/usage/faq.html) and [deployment tutorial](https://docs.zyte.com/scrapy-cloud/guides/tutorials/first-spider.html).

10. Troubleshooting common errors

Symptom Likely cause What to check or change
Settings error says the key is missing The environment variable is not present in the process that starts Scrapy. Set ZYTE_API_KEY in the shell, container, or deployment secret configuration, then restart the process.
Add-on does not appear to run Its entry was omitted, misspelled, or overwritten by a second settings definition. Merge scrapy_zyte_api.Addon into the existing ADDONS mapping and inspect the effective project settings.
Package installation or startup fails Python or Scrapy version is outside the documented package range, or dependency resolution selected an incompatible combination. Check the current package requirements, lockfile, and interpreter used by the command.
Binary response body is absent or unexpected The provider’s integration handles binary data differently from text responses. For Zyte, consult its binary example and request httpResponseBody where needed. Verify behavior for other providers independently.
Reactor startup or async errors The project uses a reactor or Deferred/Future pattern that differs from the integration’s assumptions. Review migration guidance and reproduce the failure in a minimal spider using the deployment’s reactor configuration.
Crawl slows down or memory rises Response bodies, API overhead, concurrency, or retries have increased resource use. Measure response sizes and peak memory, check Base64 overhead if using Zyte, then tune concurrency and retry limits.
Targets return errors or rate limits Request rate, provider limits, target behavior, or request options may not match the workload. Review provider response details, lower concurrency if needed, preserve appropriate delays, and use documented retry behavior.
Spider runs but items are empty The fetched page differs from the old response, selectors changed, or the callback is not receiving the expected content. Inspect a saved response, content type, status, and selector matches before changing the API configuration.

11. Or skip the browser setup

If the job is a one-off website screenshot rather than a Scrapy crawl, [ScreenshotNeo](https://screenshotneo.com) offers a single GET request for a PNG, JPEG, WebP, or PDF. Its API supports clean captures: cookie and consent banners are accepted and more than 60 known consent platforms, newsletter popups, and chat widgets are removed before the shot; each step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers say which outcome occurred. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to AI agents.

ScreenshotNeo can clear common consent banners and popups before generating a screenshot.
ScreenshotNeo can clear common consent banners and popups before generating a screenshot.

Here is a runnable cURL call; replace the target URL and use your API key. The [ScreenshotNeo API docs](https://screenshotneo.com/docs/) describe the available parameters and response behavior.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));

The Node example uses Bun’s file writer to keep it short; in Node.js, replace the final two lines with import { writeFile } from 'node:fs/promises'; and await writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));. ScreenshotNeo has 1,000 screenshots a month free with no card, and paid plans start at $5 for 3,000. See plan details and [create a free account](https://screenshotneo.com/account/sign-up/).

12. FAQ

Can I use a third-party API without rewriting my spider?

Often, yes, when the provider offers a Scrapy integration that returns responses compatible with the downloader flow. Parsing code may remain intact, but request options, binary handling, and failure behavior still require review.

Do I have to use Zyte API?

No. The example here uses Zyte because the supplied documentation describes its Scrapy add-on. Other APIs may use middleware, a package, or direct HTTP calls; follow their current documentation.

Is a screenshot API the same as a web scraping API?

No. A screenshot API returns a visual capture or PDF. A Scrapy integration fetches response data for a crawl and parsing workflow. Pick the tool that matches whether you need page images or structured data.

Does Scrapy Cloud provide the API key?

They are distinct products with distinct credentials. Use the key for the service being configured, as described in the provider’s deployment documentation.

Conclusion

Keep Scrapy responsible for request scheduling and parsing, and integrate the scraping API at the documented download boundary. Start with a small compatibility check, preserve existing settings, test each response type you need, and monitor reactor behavior, memory, delays, retries, and cost as the crawl grows.