How to Connect a Web Scraping API with n8n
Connect any web scraping API to n8n with HTTP Request, secure credentials, pagination, retries, and production-ready error handling.
The most flexible way to connect a web scraping API to n8n is the HTTP Request node. Configure the provider’s method and endpoint, store the API key in n8n credentials, add the target URL and scraper options, run one request, then map the returned records into downstream items. n8n’s HTTP Request node can query any REST API and can import a provider’s cURL example directly into its configuration.
This guide covers generic REST scraping APIs, pagination, authentication, response mapping, retries, rate limits, Apify’s managed integration, and a ScreenshotNeo option for website screenshots and PDFs.
1. What you need before building the workflow
- An n8n instance (Cloud or self-hosted).
- An API key or token from your scraping provider.
- The provider’s API reference, including the endpoint, HTTP method, authentication method, request fields, response schema, pagination model, rate limits, and retry guidance.
- A destination for the results, such as a database, spreadsheet, queue, webhook, or another API.
Do not assume that two providers use the same parameter names. One API may call the target url, another startUrls, and another may require a JSON array. Rendering, proxy, JavaScript, selector, and output-format options are provider-specific.
2. Connect a scraping API with n8n’s HTTP Request node
Step 1: Create a workflow
- Create a new workflow in n8n.
- Add a trigger. Manual Trigger is useful while building; use Schedule Trigger, a webhook, or an event trigger in production.
- Add an HTTP Request node after the trigger.
Step 2: Copy the provider’s request shape
Find the provider’s cURL example and import it into the HTTP Request node if the example is available. n8n can populate the method, URL, query parameters, headers, and body from a cURL command. This is usually safer than manually translating a request.
A generic request might look like this:
curl -X POST "https://api.example.com/v1/scrape" \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/products",
"render_js": true,
"output": "json"
}'
In n8n, the equivalent configuration is:
| HTTP Request field | Value |
|---|---|
| Method | POST |
| URL | https://api.example.com/v1/scrape |
| Authentication | The credential configured in the next step |
| Send Headers | Enable if the provider requires custom headers |
| Content-Type | application/json |
| Send Body | JSON |
| Body | The provider’s documented fields |
Step 3: Put the API key in n8n credentials
Keep secrets out of URL fields, expressions, Set nodes, and source-controlled workflow exports. In the HTTP Request node, use a predefined credential type when n8n provides one. Otherwise choose the generic authentication that matches the API:
- Header auth: add a header such as
Authorization: Bearer <token>orX-API-Key: <key>. - Basic auth: use the provider’s username and password or key pair.
- OAuth2: configure the authorization and token endpoints required by the provider.
- Custom auth: add the exact headers, query parameters, or body fields documented by the provider.
Use n8n’s credential store and restrict who can edit or view credentials. If a provider only accepts a query-string key, add it as a credential or protected expression according to your n8n deployment’s secret-management policy.
Step 4: Add the target URL and scraper options
Pass a URL from an earlier node with an expression. For example, if an incoming item contains targetUrl, use:
{{ $json.targetUrl }}
Typical provider options include:
| Option category | Examples | What to verify |
|---|---|---|
| Browser rendering | JavaScript, wait time, browser mode | Whether the provider executes scripts and how waiting is configured |
| Extraction | CSS selectors, XPath, schema, fields | Whether missing selectors fail the request or return empty values |
| Network | Proxy country, residential proxy, headers, cookies | Allowed values, additional charges, and legal restrictions |
| Output | HTML, Markdown, JSON, screenshots | Content type and response path containing records |
| Limits | Timeout, concurrency, maximum pages | Provider limits and n8n execution limits |
Step 5: Run one request before adding pagination
Execute the node with one known URL. Check all of these before building loops:
- HTTP status is successful.
- The response is JSON when you expect JSON.
- The records array or object path is correct.
- Empty results are distinguishable from an API error.
- Metadata such as a request ID, next-page URL, or usage count is preserved.
3. Turn the response into n8n items
n8n passes data between nodes as items. Scraping APIs often return one object containing an array of records, so split that array before writing to a database or spreadsheet.
Using Edit Fields or Split Out
If the response is:
{
"items": [
{"name": "First product", "price": 19.99},
{"name": "Second product", "price": 24.99}
],
"next": null
}
Add Split Out and set the field to split to items. Each product becomes an n8n item.
Using a Code node
const records = $json.items ?? [];
return records.map((record, index) => ({
json: {
...record,
source_url: $json.source_url ?? null,
record_index: index,
scraped_at: new Date().toISOString(),
},
}));
Preserve the source URL and provider request ID when available. They make deduplication, auditing, and replay easier.
4. Add pagination in n8n
The HTTP Request node supports pagination. Configure it only after confirming the provider’s single-page response and pagination fields.
Pagination by returned next URL
Choose Add Option → Pagination, then select Response Contains Next URL. Map the response field that contains the continuation URL. Stop when the field is absent or null.
Pagination by page number
Choose Update a Parameter in Each Request when the API expects a page parameter such as page. For one-based APIs, n8n’s documented expression pattern is:
{{ $pageCount + 1 }}
Set a maximum page count or another stopping condition. A practical request might include:
| Parameter | Example |
|---|---|
| Page | {{ $pageCount + 1 }} |
| Page size | The provider’s allowed maximum |
| Sort order | A stable field such as updated time plus ID |
Pagination by cursor or token
Some APIs return next_cursor or a token that must be sent in the next request. Map that response value into the next request’s query parameter or JSON body. Stop when the token is empty. Do not combine page numbers and cursors unless the provider explicitly supports it.
Pagination safeguards
- Set a maximum number of pages or records.
- Keep a stable sort order so records do not move between pages.
- Detect a repeated next URL or cursor to avoid an infinite loop.
- Deduplicate using the provider’s record ID or a hash of stable fields.
- Record the page number and source URL on every item.
- Honor the provider’s rate-limit and retry instructions.
5. Handle errors, retries, and empty results
Enable the HTTP Request node’s response options so your workflow can inspect non-2xx responses instead of silently treating them as normal data. Route failures to an error branch or an Error Trigger workflow.
| Symptom | Likely cause | Fix |
|---|---|---|
| 401 or 403 | Missing, expired, or incorrectly formatted credential | Check the provider’s required header or query parameter, rotate the key, and confirm the credential is selected on the node. |
| 400 | Wrong field name, data type, or body format | Compare the request with the provider’s current example; verify JSON versus form encoding and required fields. |
| 404 | Wrong endpoint or API version | Copy the current endpoint from the provider reference and check the account’s region or API version. |
| 408 or timeout | Slow target, JavaScript rendering, proxy delay, or too-short timeout | Increase the node timeout within provider limits, reduce page complexity, or use the provider’s asynchronous job endpoint. |
| 429 | Rate limit exceeded | Reduce concurrency, add provider-approved delay and retry behavior, and inspect any Retry-After value. |
| 5xx | Provider or upstream site failure | Retry transient errors with a cap, log the request ID, and send permanent failures to a review path. |
| 200 with no records | Selector mismatch, blocked page, empty page, or wrong response path | Inspect raw output, verify the target manually, and distinguish an empty result from a failed scrape. |
| HTML instead of JSON | Wrong Accept header, endpoint, or an intermediary error page | Check status, content type, redirects, and the provider’s required headers. |
| Duplicate records | Cursor repeated, unstable sorting, or workflow replay | Use a stable sort, detect repeated cursors, and upsert by a deterministic key. |
For transient failures, use a bounded retry strategy. A common pattern is 2–4 attempts with increasing delays, while avoiding retries for authentication and validation errors. The exact policy must follow the scraping provider’s documentation.
6. Schedule and scale the workflow
Batch targets before scraping
Represent each target as one n8n item and process batches with a controlled concurrency. This makes failures isolated and lets you resume without repeating every URL.
Control concurrency
High parallelism can trigger provider limits or target-site defenses. Use batching and delays, and leave headroom for retries. If the provider offers asynchronous jobs, submit jobs first and poll their status rather than holding one long synchronous execution.
Cache and deduplicate
Do not scrape unchanged pages unnecessarily. Store the last successful hash, timestamp, or provider record ID and skip work when the source has not changed. Keep raw responses only when you need replay or auditability; otherwise store normalized fields to reduce storage.
Monitor the workflow
- Count requested URLs, successful pages, empty pages, and failed pages.
- Track status codes, latency, provider request IDs, and retry counts.
- Alert on sustained 401, 403, 429, timeout, and 5xx responses.
- Keep a dead-letter path for URLs that require manual review.
- Use idempotent writes so an execution retry does not create duplicates.
7. Use Apify with n8n
Apify has an official n8n integration for running Actors, scraping a single URL, storing data, and triggering workflows from Actor or task events. Its API uses JSON requests and responses and supports Bearer authentication, so it can also be called through a generic HTTP Request node.
Use the native Apify node when
- You want reusable Actors and managed execution.
- You need Apify storage or Actor and task event triggers.
- Your team prefers a provider-specific node with mapped fields.
Use HTTP Request when
- The provider has no native n8n node.
- You need every provider parameter exposed directly.
- You are standardizing several providers behind one workflow pattern.
These choices are based on documented integration mechanics. Check each provider’s current API reference for pricing, browser capabilities, proxy behavior, limits, and retry rules before selecting it for production.
8. Test the same API outside n8n
Testing the request independently helps separate provider problems from workflow configuration.
cURL
curl -X POST "https://api.example.com/v1/scrape" \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{"url":"https://example.com","render_js":true}'
Python
import requests
response = requests.post(
"https://api.example.com/v1/scrape",
headers={
"Authorization": "Bearer YOUR_API_TOKEN",
"Content-Type": "application/json",
},
json={"url": "https://example.com", "render_js": True},
timeout=90,
)
response.raise_for_status()
print(response.json())
Node.js
const response = await fetch('https://api.example.com/v1/scrape', {
method: 'POST',
headers: {
Authorization: 'Bearer YOUR_API_TOKEN',
'Content-Type': 'application/json',
},
body: JSON.stringify({
url: 'https://example.com',
render_js: true,
}),
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
console.log(await response.json());
9. Or skip the browser setup
If your n8n workflow needs a clean website screenshot or PDF rather than extracted records, ScreenshotNeo provides a single GET request. Before capture, it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Only clean shots are billed: bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers. ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
Read the ScreenshotNeo API documentation for all options. In n8n, add an HTTP Request node with method GET, query parameters for access_key and url, and write the binary response to your desired destination.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
require('fs').writeFileSync('shot.webp', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF paper and margin settings, HTML/CSS rendering, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, Authorization, timezone, geolocation, transparent backgrounds, image resizing, configurable caching, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification.
There is a free plan with 1,000 screenshots per month and no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account and add the key to an n8n credential.
10. Cost, reliability, and security checklist
Cost
- Check whether the provider bills per request, successful page, browser minute, proxy request, extracted record, or storage unit.
- Estimate retries and pagination, not just the number of starting URLs.
- Use caching and deduplication where the provider permits it.
- Set workflow limits so a pagination bug cannot create an unbounded bill.
Reliability
- Use bounded retries and exponential or provider-recommended delays.
- Persist progress after each page or batch.
- Make downstream writes idempotent.
- Keep request IDs and raw error details for support cases.
- Separate permanent failures from transient failures.
Security
- Store keys in n8n credentials.
- Do not log Authorization headers or full URLs when they contain secrets.
- Validate incoming target URLs if a webhook lets users choose them.
- Restrict workflow editing and credential access.
- Review provider terms, robots policies, privacy requirements, and applicable law for the sites you scrape.
11. Frequently asked questions
Can n8n connect to any scraping API?
Yes, when the service exposes an HTTP API that n8n can call. The HTTP Request node supports REST methods, headers, query parameters, bodies, authentication, and pagination options.
Where should the API key go?
Use an n8n credential. Select a predefined credential when available, or configure generic Header, Basic, OAuth, or Custom authentication according to the provider’s documentation.
How do I scrape JavaScript-rendered pages?
Enable the provider’s documented browser or JavaScript-rendering option and configure a selector wait, delay, or network-idle wait when supported. The exact field names differ by provider.
How do I avoid duplicate pages?
Use a stable sort, detect repeated cursors or next URLs, store a deterministic record key, and make destination writes idempotent.
Should I use Apify’s n8n node or HTTP Request?
Use the native node for reusable Actors, storage, and event-driven workflows. Use HTTP Request when you need a provider endpoint or parameter that is not represented by a native node.
Can n8n save binary screenshots?
Yes. Configure the HTTP Request node to return a file or binary response and pass that binary property to storage, email, or another node. Confirm the provider’s content type and file extension.


