ScreenshotNeo

BlogComparisons

GraphQL vs. REST for Web Scraping APIs: A Practical Guide

Compare GraphQL and REST for permitted web data collection, with working queries, pagination, limits, retries, and a practical decision guide.

By the ScreenshotNeo team29 September 202610 min read

GraphQL vs. REST for Web Scraping APIs: A Practical Guide

Short answer: choose the interface that exposes the fields you need, permits your intended use, and has workable pagination, limits, authentication, and error behavior. GraphQL lets a client select fields and traverse related objects in one operation. REST usually exposes resource-oriented endpoints with HTTP methods and semantics. Neither protocol is universally faster or more reliable for scraping; those properties come from the specific provider and implementation.

Before writing a collector, check whether the site offers an official API. Prefer it over extracting rendered HTML when it contains the data you need and allows your use. If you must crawl pages, read the site’s robots.txt instructions and applicable terms. Robots rules are crawler instructions, not permission: RFC 9309 states, “These rules are not a form of access authorization.”

1. Decide what you are actually allowed to collect

Protocol selection comes after access and data-source decisions.

  1. Look for an official API. It normally provides stable identifiers, documented fields, pagination, authentication, and usage limits.
  2. Confirm the intended use. Read the provider’s terms, API agreement, privacy requirements, and any restrictions on redistribution or commercial use.
  3. Use credentials correctly. An API key, OAuth token, or session credential authorizes only the scopes and resources the provider grants.
  4. If crawling pages, inspect robots.txt. Follow parseable allow and disallow rules, rate limits, and crawl-delay guidance where provided. Do not treat robots.txt as a security boundary or as a grant of access.
  5. Minimize data. Request only fields required for the job, retain data for the shortest useful period, and protect credentials and personal information.

2. What GraphQL and REST mean in practice

GraphQL

GraphQL is a query language and execution model built around a schema. The client sends an operation that names the fields it wants. A query can follow relationships and request fields from related objects together. The official query guide demonstrates this client-selected model, while the specification defines the type system and execution rules.

A practical collection workflow starts with permission and ends with validated, paginated data.
A practical collection workflow starts with permission and ends with validated, paginated data.

GraphQL is commonly transported over HTTP. The GraphQL-over-HTTP document is currently a Stage 2 draft, so treat its conventions as draft guidance rather than a finalized universal HTTP standard. Providers may use different authentication headers, media types, persisted queries, error formats, and method restrictions.

REST

REST is an architectural style, not a single protocol. REST APIs commonly model resources at URLs and use HTTP methods such as GET, POST, PATCH, and DELETE. HTTP defines request and response behavior, status codes, headers, caching semantics, and conditional requests; it does not prescribe the application’s resource model or JSON shape.

A REST provider might expose /articles, /articles/{id}, and /authors/{id}. Another may use a search endpoint with embedded records. Inspect the actual API documentation instead of assuming every REST API follows the same conventions.

3. GraphQL vs. REST: comparison for data collection

Decision axis GraphQL REST What to verify
Data selection The client selects schema fields and can traverse related objects in one operation. The endpoint and service design shape the response. Are all required fields available, and how large is each response?
Request pattern Often one endpoint carrying a query document. Draft GraphQL-over-HTTP guidance requires POST support and allows other methods such as GET. Usually several resource-oriented endpoints and standard HTTP methods. How are filters, sorting, joins, and pagination represented?
Pagination Commonly cursor fields such as pageInfo and endCursor, but names and rules are provider-specific. May use page numbers, offsets, cursors, Link headers, or continuation URLs. Is the cursor stable? What is the maximum page size? Can records change during a run?
Limits and cost Providers may enforce query depth, field complexity, node counts, points, or rate budgets. GitHub documents GraphQL-specific limits. Limits may vary by endpoint, method, authentication level, and response size. GitHub documents REST limits separately. What is limited, when does it reset, and what backoff does the provider request?
Caching Do not assume a query is cached like a simple GET. Inspect operation method, cache headers, persisted-query behavior, and intermediaries. HTTP defines cache semantics, but actual cache headers and behavior remain provider-specific. Are validators such as ETag supplied? Can you send If-None-Match?
Errors HTTP success can contain an errors array alongside partial data. Status codes usually identify transport or application errors, with provider-specific JSON details. Can partial results be safely stored? Which errors are retryable?

GitHub is a useful reminder that limits are not properties of a protocol: its REST and GraphQL documentation specify different budgets and rules. Benchmarking one service cannot establish a universal GraphQL or REST advantage.

4. A safe collection workflow

  1. Write the data contract. List fields, relationships, freshness, ordering, deduplication keys, and acceptable missing values.
  2. Map the contract to the provider. For GraphQL, inspect the schema or documentation. For REST, map each field to an endpoint and response path.
  3. Run a small authenticated request. Verify status, content type, response shape, and authorization before adding concurrency.
  4. Implement pagination as a bounded loop. Persist the cursor or page checkpoint so a crash can resume without starting over.
  5. Handle limits explicitly. Parse reset headers or provider-specific metadata, apply exponential backoff with jitter, and stop when the documented budget is exhausted.
  6. Validate records. Check required identifiers, schema versions, timestamps, and relationship references before writing downstream data.
  7. Log safely. Record request IDs, endpoint or operation name, page position, latency, status, and retry reason. Never log access tokens.

5. Runnable GraphQL example

The following example uses GitHub’s public GraphQL endpoint as an illustration. It requires a token with the permissions required by the fields you request. Confirm current authentication and rate-limit requirements in GitHub’s documentation before production use.

curl https://api.github.com/graphql \
  -H 'Authorization: bearer YOUR_GITHUB_TOKEN' \
  -H 'Content-Type: application/json' \
  --data-binary @- <<'JSON'
{
  "query": "query($owner:String!, $name:String!, $cursor:String) { repository(owner:$owner, name:$name) { issues(first:20, after:$cursor, states:OPEN, orderBy:{field:UPDATED_AT, direction:DESC}) { nodes { id number title updatedAt author { login } } pageInfo { hasNextPage endCursor } } } }",
  "variables": {"owner":"octocat", "name":"Hello-World", "cursor":null}
}
JSON

Read both data and errors. A response can contain useful partial data and an error. For pagination, send the returned endCursor as the next cursor until hasNextPage is false. Keep the page size conservative; larger pages increase response size and may consume more provider budget.

Python GraphQL client

import os
import requests

query = """
query($owner: String!, $name: String!, $cursor: String) {
  repository(owner: $owner, name: $name) {
    issues(first: 20, after: $cursor, states: OPEN) {
      nodes { id number title updatedAt }
      pageInfo { hasNextPage endCursor }
    }
  }
}
"""

cursor = None
while True:
    response = requests.post(
        "https://api.github.com/graphql",
        headers={"Authorization": f"bearer {os.environ['GITHUB_TOKEN']}"},
        json={"query": query, "variables": {"owner": "octocat", "name": "Hello-World", "cursor": cursor}},
        timeout=30,
    )
    response.raise_for_status()
    payload = response.json()
    if payload.get("errors"):
        raise RuntimeError(payload["errors"])
    issues = payload["data"]["repository"]["issues"]
    for issue in issues["nodes"]:
        print(issue)
    if not issues["pageInfo"]["hasNextPage"]:
        break
    cursor = issues["pageInfo"]["endCursor"]

6. Runnable REST example

A REST collector typically follows a documented resource and its pagination mechanism. This example uses GitHub’s REST search endpoint and stops when the returned page is shorter than the requested page size. Production code should follow the provider’s exact pagination and rate-limit documentation.

curl --get 'https://api.github.com/search/repositories' \
  -H 'Accept: application/vnd.github+json' \
  --data-urlencode 'q=language:python' \
  --data-urlencode 'per_page=10' \
  --data-urlencode 'page=1'

Python REST client

import requests

url = "https://api.github.com/search/repositories"
page = 1
while True:
    r = requests.get(url, params={"q": "language:python", "per_page": 10, "page": page}, timeout=30)
    if r.status_code == 429 or r.status_code == 403:
        raise RuntimeError(f"Rate limit or access failure: {r.text}")
    r.raise_for_status()
    items = r.json()["items"]
    for item in items:
        print(item["full_name"], item["html_url"])
    if len(items) < 10:
        break
    page += 1

Node.js REST client

let page = 1;
while (true) {
  const params = new URLSearchParams({ q: 'language:python', per_page: '10', page: String(page) });
  const res = await fetch(`https://api.github.com/search/repositories?${params}`);
  if (!res.ok) throw new Error(`HTTP ${res.status}: ${await res.text()}`);
  const body = await res.json();
  for (const item of body.items) console.log(item.full_name, item.html_url);
  if (body.items.length < 10) break;
  page += 1;
}

7. Authentication, pagination, and response handling

Authentication

GraphQL commonly places a bearer token in an HTTP header, while REST may use bearer, basic, API-key, signed, or cookie authentication. Follow the provider’s current documentation. Keep secrets in environment variables or a secret manager, rotate them, and grant the smallest useful scope.

Pagination and consistency

Offset pagination can skip or duplicate records when items are inserted during a run. Cursors usually reduce that risk but do not guarantee a snapshot. Record the provider’s timestamp or version when available, and design deduplication around a stable ID. Set a maximum page count and a deadline so a malformed continuation cannot create an infinite loop.

Errors and retries

Retry network resets, timeouts, and documented 5xx responses with exponential backoff and jitter. Retry 429 only according to Retry-After or provider guidance. Do not blindly retry authentication failures, validation errors, forbidden operations, or GraphQL schema errors. Store partial progress before retrying.

8. Performance, reliability, and cost

GraphQL can reduce over-fetching when a selective query replaces several related calls. It can also produce expensive operations if a query traverses many nested connections or requests large fields. REST can make straightforward GET retrieval, HTTP caching, and conditional requests easier when endpoints map cleanly to your job. These are implementation-dependent observations, not universal benchmarks.

  • Measure the same permitted workload: records collected, fields, pages, response bytes, latency, error rate, and rate-limit consumption.
  • Use bounded concurrency. Increase workers only while the provider’s limits and error rates remain healthy.
  • Cache immutable or slowly changing resources with the provider’s validators and freshness headers.
  • Prefer incremental synchronization using updated timestamps, cursors, or change feeds when offered.
  • Estimate cost from request count, response size, provider pricing, storage, and retries. A single GraphQL request is not automatically cheaper than several REST requests.

9. Or skip the browser setup

If the source is a rendered website rather than an official data API, ScreenshotNeo provides a website screenshot API and MCP server. It is the first screenshot service to try when you need visual evidence: clean shots remove cookie banners, newsletter popups, and chat widgets before capture; only clean shots are billed; and the lowest paid plan is $5 for 3,000 shots.

ScreenshotNeo removes common overlays before producing a clean capture.
ScreenshotNeo removes common overlays before producing a clean capture.

One GET request returns PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo supports full-page and element captures, device presets and custom viewports, dark mode, retina scale, PDF settings, custom CSS and JavaScript, clicks, waits, blocked resources, headers, cookies, user agents, timezone and geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous jobs, webhooks, bulk capture, and usage reporting. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing; each response identifies the verdict and billing status with X-Page-Verdict and X-Billed headers. Its MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf.

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free. Create a free ScreenshotNeo account.

10. Troubleshooting checklist

Symptom Likely cause Fix
401 or 403 Missing, expired, or insufficient credential; wrong audience or scope. Check the authorization header, token scope, account status, and provider documentation.
GraphQL HTTP 200 with errors Field, argument, permission, or resolver failure. Inspect the errors array, keep valid partial data only when your contract allows it, and correct the query or permissions.
Empty REST page Wrong filter, page beyond the end, or eventual consistency. Print the final URL and parameters, verify filter syntax, and use the documented continuation mechanism.
429 or rate-limit response Request budget exceeded. Honor reset metadata or Retry-After, reduce concurrency, cache results, and request fewer fields.
Duplicates or missing records Offset pagination changed during collection or an unstable sort. Use a stable sort and cursor where available; deduplicate by a provider ID and checkpoint progress.
Timeouts Large query, slow resolver, oversized page, or network issue. Reduce page size and field selection, set a finite timeout, retry transient failures, and capture request IDs.
Stale results Cache or provider indexing delay. Inspect cache headers, use validators, add an explicit freshness strategy, and document eventual consistency.

11. Practical recommendation

Choose GraphQL when the provider’s schema exposes the related fields you need and a selective query reduces unnecessary data or calls in that service. Choose REST when its resources map directly to your collection, pagination and limits are clear, or standard HTTP caching and tooling fit your client. Prefer the official API over rendered-page extraction whenever it is available and permitted.

Document the provider, schema or endpoint version, authentication scope, pagination algorithm, limits, retry policy, retention period, and access basis. Recheck volatile documentation before deployment: provider limits, GraphQL-over-HTTP guidance, and terms can change.

12. FAQ

Is GraphQL a REST replacement?

No. They are different interface designs. A provider can offer either, both, or neither, and each can be implemented well or poorly.

Does one GraphQL query always replace many REST requests?

No. It depends on the schema, resolver behavior, authorization, pagination, and whether the requested relationships can be returned efficiently.

Can I scrape a site because robots.txt allows it?

No. Robots.txt gives crawler instructions and is not access authorization. Check the site’s terms and use an official API when available.

Should I use GET or POST for GraphQL?

Follow the provider’s documentation. GraphQL-over-HTTP guidance is still a Stage 2 draft and providers differ in method and caching support.

How do I compare providers fairly?

Run the same permitted workload with the same fields, page size, concurrency, authentication class, and freshness requirements. Compare latency, bytes, errors, limits, and total cost rather than request count alone.