ScreenshotNeo

BlogHow-to

How to Scrape Reviews and Q&A Data

A platform-aware guide to collecting reviews and Q&A data with APIs, authorized crawling, pagination, provenance, and safe storage.

By the ScreenshotNeo team29 September 20268 min read

How to Scrape Reviews and Q&A Data

Short answer: start with the review or Q&A platform’s official API, export, or licensed feed. Crawl HTML only when the current terms, robots.txt instructions, and applicable law permit automated collection. Keep source IDs, timestamps, request provenance, pagination state, and the platform’s display and retention rules with every record. A few visible pages or API excerpts are not a complete dataset.

This guide shows a defensible workflow for collecting reviews and questions-and-answers, with Python, cURL, and Node.js examples. It also explains why “scrape” is platform-specific: Yelp says its site may not be scraped, Google Maps terms prohibit copying reviews outside its services, while documented APIs provide narrower, governed access.

1. Define the data boundary before writing code

Write a one-page collection specification:

  • Targets: product, business, ASIN, place ID, or URL; include variants and locales.
  • Fields: stable source ID, rating scale, review or question text, answer text, author label if permitted, created and updated timestamps, language, verified-purchase flag, source URL, and fetch time.
  • Window: start and end dates, plus whether edited or deleted records must be tracked.
  • Use: internal analysis, support search, display, redistribution, or model training. Each use can have different permission and retention rules.
  • Personal data: avoid reviewer identifiers unless necessary. Define access controls and deletion handling.

Record what “complete” means. An endpoint returning three excerpts, a ranked page, or a weekly insight feed must be reported as a limited sample or summary, never as all reviews.

2. Check permission and choose the narrowest adequate route

Read the current terms, API documentation, data license, and robots.txt before collecting. RFC 9309 describes robots.txt as a crawler instruction protocol; following it does not itself grant a license. Yelp’s support policy says third-party software may not scrape or copy Yelp site content, even though Yelp documents a Places API reviews endpoint that returns up to three excerpts per business (Yelp API documentation; Yelp support policy). Google Maps Platform terms prohibit scraping or exporting Maps content for use outside its services, and Places policies specify attribution, direct source access, and storage limits (terms; Places policies).

Route Use when Resolve first
Official API It covers the records and your intended use Eligibility, fields, quotas, regions, freshness, attribution, storage and cost
Licensed feed or partner You need broader coverage or commercial rights Sources, permitted combinations, retention, display, redistribution and training rights
Direct crawl No adequate authorized route exists and the site permits it Terms, robots rules, rate, identification, privacy, copyright, database and jurisdiction rules

Amazon illustrates why an API is not an unrestricted dump. Its Customer Feedback API is for eligible sellers and vendors, lists US, UK, France, Italy, Germany, Spain and Japan, refreshes weekly, is English-only, and exposes review-topic insights at ASIN or browse-node level. Check the current API documentation and role requirements before designing around it.

3. Build a conservative collector

Python HTML collector (only where crawling is authorized)

The example below demonstrates identification, robots checking, bounded retries, and provenance. Adapt selectors to the permitted site; do not use it to bypass access controls, CAPTCHAs, logins, or rate limits.

A governed collection flow preserves source IDs, pagination and provenance.
A governed collection flow preserves source IDs, pagination and provenance.
import json, time
from datetime import datetime, timezone
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
from urllib.robotparser import RobotFileParser

START = "https://example.com/reviews"
UA = "ReviewResearchBot/1.0 (+https://example.com/contact)"

def allowed(url):
    p = RobotFileParser(urljoin(url, "/robots.txt"))
    p.read()
    return p.can_fetch(UA, url)

s = requests.Session()
s.headers.update({"User-Agent": UA, "Accept": "text/html"})
rows = []
url = START
for page in range(1, 6):
    if not allowed(url):
        raise RuntimeError(f"robots.txt disallows {url}")
    for attempt in range(4):
        r = s.get(url, timeout=30)
        if r.status_code == 429:
            wait = int(r.headers.get("Retry-After", "10"))
            time.sleep(min(wait, 120)); continue
        r.raise_for_status(); break
    soup = BeautifulSoup(r.text, "html.parser")
    for card in soup.select("article.review"):
        text = card.select_one("[data-review-text]")
        if not text: continue
        rows.append({
            "source_id": card.get("data-review-id"),
            "text_raw": text.get_text(" ", strip=True),
            "rating": card.get("data-rating"),
            "source_url": url,
            "fetched_at": datetime.now(timezone.utc).isoformat(),
            "page": page
        })
    nxt = soup.select_one("a[rel=next]")
    if not nxt: break
    url = urljoin(url, nxt["href"])
    time.sleep(2)

with open("reviews.jsonl", "w", encoding="utf-8") as f:
    for row in rows: f.write(json.dumps(row, ensure_ascii=False) + "\n")

Use a documented API in preference to selectors. Save the request URL or endpoint identifier, query parameters, response status, API version, and a hash of the raw response when terms allow it. Never defeat a bot check or rotate identities to evade a block.

cURL pagination pattern for an authorized API

curl --fail-with-body --retry 3 --retry-delay 2 \
  -H "Authorization: Bearer $API_TOKEN" \
  --get "https://api.example.com/v1/reviews" \
  --data-urlencode "business_id=abc123" \
  --data-urlencode "limit=100" \
  --data-urlencode "cursor=$NEXT_CURSOR"

Read the documented next cursor, stop when it is absent, and log the cursor that produced each page. Do not assume page numbers are stable while reviews are being added or removed.

Node.js fetch with backoff

const sleep = ms => new Promise(r => setTimeout(r, ms));
async function getPage(cursor) {
  const u = new URL('https://api.example.com/v1/reviews');
  u.searchParams.set('business_id', 'abc123');
  u.searchParams.set('limit', '100');
  if (cursor) u.searchParams.set('cursor', cursor);
  for (let n = 0; n < 4; n++) {
    const res = await fetch(u, {headers: {Authorization: `Bearer ${process.env.API_TOKEN}`} });
    if (res.status === 429) { await sleep(Math.min(120000, 2000 ** (n + 1))); continue; }
    if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
    return res.json();
  }
  throw new Error('rate limit persisted after retries');
}

4. Normalize, deduplicate, and preserve meaning

Keep raw and derived representations separate. A practical record has source, source_id, entity_id, kind (review, question, answer), rating, language, created_at, updated_at, text_raw, text_normalized, source_url, fetched_at, api_version, and provenance. Store translation, redaction, sentiment labels, and other transformations as separate fields with method and timestamp.

Prefer the platform ID for deduplication. If no stable ID exists, combine canonical URL, author label, timestamp and a normalized-text hash, then flag collisions for review. Keep a change log when text, rating, answer or moderation status changes. Do not silently merge two products, locations, or language variants.

5. Measure coverage and quality

  • Count requested, successful, denied, rate-limited, empty and malformed pages.
  • Record the API maximum result set, ranking or selection rule, and missing cursors.
  • Compare collected counts with source totals where the source exposes them.
  • Stratify by date, language, rating, product variant and locale to reveal ranking bias.
  • Mark deleted, inaccessible and untranslated records instead of dropping them without explanation.

Yelp’s three-excerpt limit and Amazon’s weekly, English-only insight feed are coverage constraints that belong in your dataset documentation. A dashboard should show the collection window and source route beside every aggregate.

6. Retention, display, and review integrity

Apply the source’s storage and attribution requirements to raw and derived data. Google Places policies require author attribution and direct access to source reviews and restrict caching beyond stated exceptions. Amazon’s community guidance says to post only your own content or content you have permission to use; a connected person answering product Q&A must disclose that connection (community guidelines; promotional-content guidance).

FTC staff guidance recommends reasonable authenticity processes, equal treatment of positive and negative reviews, and no editing that changes a review’s message. Its Consumer Reviews and Testimonials Rule took effect October 21, 2024; the guidance is not project-specific legal advice (platform guide; rule Q&A). Preserve original text, label summaries as summaries, and disclose sampling and affiliations.

7. Reliability, performance, and cost

Use bounded concurrency rather than one request per record. Honor documented quotas, exponential backoff with jitter, connection reuse, and checkpointing after each page. Cache only when the license permits it; otherwise retain identifiers and refetch on schedule. For long jobs, queue cursors and make writes idempotent so a retry cannot duplicate records.

Estimate cost as requests or records multiplied by the provider’s rate, plus storage, translation and processing. API quotas, role eligibility and refresh cadence can matter more than raw request price. For a crawl, budget bandwidth and parsing CPU, then add review time for selector changes and access denials. Stop on repeated 401, 403, CAPTCHA or legal-notice responses; escalating retries increases load and can breach terms.

8. Troubleshooting checklist

Symptom Likely cause Fix
403 or legal notice Access is disallowed or credentials lack scope Stop; recheck terms and API eligibility. Use an approved endpoint or licensed feed.
429 responses Quota or rate exceeded Reduce concurrency, honor Retry-After, add jitter, request a documented quota increase.
Empty pages Wrong locale, filters, cursor, or a limited endpoint Log the exact request; verify IDs and documented coverage; do not treat empty as end-of-data until the API says so.
Duplicate reviews Offset pagination shifted during updates or retries replayed a page Use stable IDs, cursor pagination and idempotent upserts; retain fetch logs.
Missing older records Ranking, retention window or endpoint cap Document the cap and use an export or licensed source if full history is required.
Changed HTML selectors Site redesign or client-rendered content Prefer the official API; otherwise version selectors, add fixtures and stop safely on schema drift.
Text meaning changed Cleaning, translation or truncation Keep raw text, transformation metadata and an audit trail; never rewrite sentiment.

9. Or skip the browser setup

If your workflow only needs a clean image or PDF of a review or Q&A page, ScreenshotNeo provides a single GET request. It accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each step can be disabled. Only clean shots are billed: bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers.

Consent elements and overlays can be handled before a clean capture.
Consent elements and overlays can be handled before a clean capture.

See the ScreenshotNeo API docs for all options. A minimal capture:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

For review evidence, relevant options include full-page capture with lazy images loaded, a CSS element selector, custom CSS or JavaScript, click-before-capture, hide selectors, waits for a selector, delay or network idle, custom headers/cookies/user agent/Authorization, timezone and geolocation, image format and resizing, a chosen cache TTL, signed links, async jobs with signed webhooks, bulk capture of 100 URLs per call, PDF paper size/margins/landscape/page ranges, and an MCP server with take_screenshot, get_page_info and capture_pdf for AI agents. Use authorization only where you have the right to view the page and avoid capturing personal data you do not need.

ScreenshotNeo has 1,000 free screenshots per month with no card; paid plans start at $5 for 3,000. Every feature is on every plan. Create a free ScreenshotNeo account.

10. FAQ

Is scraping reviews always illegal?

No universal answer exists. Permission depends on the platform, purpose, data, location and downstream use. Check current terms, API rules, robots instructions and applicable law before collecting.

Can an API response be stored forever?

Not automatically. Providers may limit caching, retention, attribution and display. Record those rules with the dataset and expire or refresh records accordingly.

How do I prove what I collected?

Keep request timestamps, source IDs or URLs, cursors, API version, response hashes where allowed, parser version, transformations and failure logs.

Should I publish reviewer names?

Only when necessary and permitted. Minimize personal data, follow attribution requirements, and provide deletion or access controls appropriate to your use.