How to Scrape Twitch Data with an API
Use Twitch Helix and OAuth to collect users, streams, videos, and events with pagination, rate-limit handling, and reliable code examples.
Use Twitch’s official Helix API instead of scraping rendered Twitch pages. Register an application, obtain the token type required by your endpoint, send the token with the matching Client-Id, and follow the endpoint’s documented pagination and rate limits. This approach gives you documented fields and authorization behavior, but it does not promise every Twitch record or an exhaustive historical archive.
Twitch describes its API as providing “the tools and data used to develop Twitch integrations.” Read the Twitch API concepts, getting-started guide, and the endpoint reference before building a collector.
What “scraping Twitch” means here
In this guide, scraping means retrieving Twitch data through Helix API endpoints. You can request users, channels, live streams, games, videos, clips, followers, and other documented resources. The API returns structured JSON and requires OAuth authentication.
Do not assume that an endpoint exposes every item visible on Twitch, that a result is a complete all-time archive, or that pages form a consistent snapshot while you are reading them. For example, Twitch documents that “Get Videos by Game” returns about 500 videos at most. Check each endpoint’s limits and authorization requirements in the API reference.
Authentication: app token or user token?
Every integration starts with an application registration. Keep the client secret, access tokens, and refresh tokens as protected credentials; never put a client secret in browser-side JavaScript.
| Token | Use it for | What to check |
|---|---|---|
| App access token | Non-sensitive resources that do not require a user’s permission. EventSub webhook API calls require an app access token. | Use the client-credentials flow and confirm the endpoint accepts an app token. |
| User access token | Resources that require a user’s permission or specific scopes. | Send the scopes requested by the endpoint and obtain consent through the appropriate OAuth flow. |
The accepted token type and required scopes are endpoint-specific. Follow Twitch’s authentication documentation and OAuth token flow guidance.
Step 1: Register a Twitch application
- Open the Twitch developer console and register an application.
- Record the client ID and keep the client secret in a server-side secret store or environment variable.
- Choose the endpoint you need and read its authorization section before requesting a token.
- Decide whether your job needs a one-time snapshot (API polling) or ongoing updates (EventSub).
Step 2: Get an app access token
The following cURL request uses Twitch’s client-credentials grant. Replace the placeholders and keep the response token private.
curl -X POST 'https://id.twitch.tv/oauth2/token' \
-H 'Content-Type: application/x-www-form-urlencoded' \
-d 'client_id=YOUR_CLIENT_ID' \
-d 'client_secret=YOUR_CLIENT_SECRET' \
-d 'grant_type=client_credentials'
The JSON response contains an access_token. Use it as a bearer token together with the same application’s client ID.
Step 3: Make a focused Helix request
Here is the equivalent “get users” request from the official starter flow. Use a login or ID that exists, and URL-encode query parameters when constructing requests programmatically.
curl 'https://api.twitch.tv/helix/users?login=twitch' \
-H 'Authorization: Bearer YOUR_ACCESS_TOKEN' \
-H 'Client-Id: YOUR_CLIENT_ID'
Python
import os
import requests
client_id = os.environ["TWITCH_CLIENT_ID"]
token = os.environ["TWITCH_ACCESS_TOKEN"]
response = requests.get(
"https://api.twitch.tv/helix/users",
params={"login": "twitch"},
headers={
"Authorization": f"Bearer {token}",
"Client-Id": client_id,
},
timeout=30,
)
response.raise_for_status()
print(response.json())
Node.js
const clientId = process.env.TWITCH_CLIENT_ID;
const token = process.env.TWITCH_ACCESS_TOKEN;
const params = new URLSearchParams({ login: "twitch" });
const response = await fetch(`https://api.twitch.tv/helix/users?${params}`, {
headers: {
Authorization: `Bearer ${token}`,
"Client-Id": clientId,
},
});
if (!response.ok) {
throw new Error(`Twitch returned ${response.status}: ${await response.text()}`);
}
console.log(await response.json());
Choosing endpoints and parameters
Start with the resource you actually need, then copy its required query parameters and authorization rules from the reference. Common collection tasks include:
- Users: resolve logins or IDs to user records.
- Streams: find live channels and filter by game, language, or user.
- Games: resolve game IDs and names.
- Videos: list videos for a user or game, subject to the endpoint’s documented bounds.
- Clips: retrieve clips using the filters supported by the endpoint.
- Followers and other protected resources: check whether a user token and scopes are required.
Treat IDs as opaque strings. Do not cast them to integers or infer meaning from their values. Parse documented JSON fields, ignore unknown fields so additions remain compatible, and avoid depending on undocumented URL shapes or error-message wording.
Pagination with cursors
List endpoints use cursors rather than page numbers. Set first within the endpoint’s allowed range. If the response contains a pagination cursor, send it as after for the next request. before is supported only by some endpoints, and after and before cannot be used together.
import os
import requests
endpoint = "https://api.twitch.tv/helix/streams"
headers = {
"Authorization": f"Bearer {os.environ['TWITCH_ACCESS_TOKEN']}",
"Client-Id": os.environ["TWITCH_CLIENT_ID"],
}
cursor = None
seen = set()
while True:
params = {"first": 100, "language": "en"}
if cursor:
params["after"] = cursor
response = requests.get(endpoint, headers=headers, params=params, timeout=30)
response.raise_for_status()
payload = response.json()
for stream in payload.get("data", []):
stream_id = stream["id"]
if stream_id in seen:
continue
seen.add(stream_id)
print(stream_id, stream.get("user_login"), stream.get("viewer_count"))
next_cursor = payload.get("pagination", {}).get("cursor")
if not next_cursor or not payload.get("data"):
break
cursor = next_cursor
Lists are dynamic. A record can move, disappear, or appear twice while you page. Deduplicate by a stable documented ID, tolerate an empty page near the end, and do not claim snapshot consistency unless your own system provides it.
Polling versus EventSub
| Approach | Best for | Trade-offs |
|---|---|---|
| Polling Helix | A current-state read, scheduled snapshot, or backfill within endpoint limits. | Consumes rate-limit budget and can miss changes between polls. |
| EventSub Webhooks | Receiving events through an HTTPS callback. | Requires a reachable webhook and subscription lifecycle handling. |
| EventSub WebSockets | Long-lived connections from an application that can maintain a socket. | Requires reconnect and session management. |
| EventSub Conduits | Architectures using Twitch’s conduit transport. | Follow the subscription type’s transport support and operational requirements. |
Twitch recommends EventSub when you need updates such as a broadcaster going online, new followers or subscribers, cheers, or Channel Point redemptions. Delivery is at least once, so deduplicate by EventSub message ID and make notification processing idempotent. Validate incoming messages according to Twitch’s current EventSub and security guidance.
Rate limits and resilient collection
Twitch uses token buckets. The default request cost is one point unless an endpoint documents another cost. Limits are associated with the client ID/app; user-token requests have per-client, per-user limits. Read these response headers:
Ratelimit-Limit: the bucket limit.Ratelimit-Remaining: points left.Ratelimit-Reset: the reset time.
The 800 value shown in some examples is an example header, not a universal quota. On HTTP 429, wait until the reset time (with a small safety margin), then retry with exponential backoff. Also inspect endpoint-specific limits and costs.
import time
def get_with_backoff(session, url, *, headers, params, attempts=5):
for attempt in range(attempts):
response = session.get(url, headers=headers, params=params, timeout=30)
if response.status_code != 429:
response.raise_for_status()
return response
reset = response.headers.get("Ratelimit-Reset")
delay = max(1, int(reset) - int(time.time())) if reset and reset.isdigit() else 2 ** attempt
time.sleep(delay)
raise RuntimeError("Twitch rate limit did not recover after retries")
Dates, identifiers, and response compatibility
- API date-time values use RFC3339. EventSub timestamps can include nanosecond precision.
- Keep IDs as strings and preserve them exactly.
- Ignore unknown JSON fields and do not rely on field order.
- Do not build logic around the exact text of an error message.
- Store the request parameters, token type, response status, and retrieval time with each collection job so results can be audited.
Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| 401 Unauthorized | Expired, malformed, or wrong token. | Obtain a new token, send Authorization: Bearer TOKEN, and validate tokens using Twitch’s current instructions. |
| Client ID mismatch | The token was issued for a different application. | Send the client ID belonging to the application that issued the token. |
| 403 Forbidden | The endpoint needs a user token or a missing scope. | Read the endpoint authorization section and repeat the user-consent flow with the required scopes. |
| 400 Bad Request | Missing, invalid, or mutually exclusive query parameters. | Check parameter names, allowed ranges, date format, and whether after and before were sent together. |
| Empty data array | No matching records, a dynamic list changed, or you reached the end. | Check filters and stop pagination when the cursor is absent or the page is empty. |
| 429 Too Many Requests | Bucket exhausted. | Read Ratelimit-Reset, wait, reduce concurrency, and cache stable lookups. |
| Duplicate records | Dynamic results or at-least-once EventSub delivery. | Deduplicate by resource ID or EventSub message ID and make writes idempotent. |
Performance, reliability, and cost notes
- Request only the fields and filters the endpoint supports; smaller focused jobs are easier to retry.
- Use bounded concurrency so one worker cannot consume the entire application bucket.
- Cache rarely changing entities such as game and user lookups, while respecting freshness requirements.
- Persist cursors and the last successful page so a failed job can resume without starting over.
- Use EventSub for frequent state changes instead of polling every channel continuously.
- Record rate-limit headers and 429s as metrics. Alert on sustained depletion rather than a single response.
- Estimate cost from your own infrastructure, storage, and polling frequency; Twitch’s API documentation does not provide a universal monetary price for Helix requests.
Review the current Twitch Developer Services Agreement and applicable policies for your storage, redistribution, and commercial use case.
Or skip the browser setup
If your workflow also needs screenshots of Twitch pages, ScreenshotNeo provides a single GET request that returns a PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status.
See the ScreenshotNeo API documentation for all options. This call captures a Twitch URL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.twitch.tv -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://www.twitch.tv"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://www.twitch.tv' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also offers an MCP server for Claude, Cursor, and other MCP clients, so AI agents can call take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Can I collect Twitch data without OAuth?
No. Twitch integrations require an application and an OAuth access token. The required token type depends on the endpoint.
Is the API data a complete archive of Twitch?
No. Each endpoint has its own filters, bounds, and retention behavior. The videos-by-game endpoint, for example, documents a limit of about 500 videos.
Should I poll or subscribe to events?
Poll for a snapshot or scheduled read. Use EventSub when you need ongoing notifications and can handle at-least-once delivery.
Can I use the same pagination code for every endpoint?
Use the cursor pattern as a starting point, but verify each endpoint’s allowed page size and whether backward pagination is supported.
How should I handle a repeated EventSub notification?
Store processed message IDs and make the resulting operation idempotent before acknowledging or acting on the notification.


