ScreenshotNeo

BlogHow-to

How to Scrape G2 Reviews With JavaScript

Learn the JavaScript mechanics for review-page parsing, pagination, and validation—while using G2’s permitted API or written consent for automated access.

By the ScreenshotNeo team29 September 20269 min read

How to Scrape G2 Reviews With JavaScript

There are two separate questions behind “How do I scrape G2 reviews with JavaScript?” The first is technical: how do you request a page, verify that review content arrived, parse repeated cards, and follow pagination? The second is permission: whether you may automate that process against G2.

G2’s Terms of Use, last updated July 9, 2026, prohibit automated, programmatic, or mechanical extraction of site content—including reviews, ratings, reviewer metadata, rankings, and product information—without G2’s express prior written consent. The same terms prohibit bypassing access controls, bot detection, CAPTCHAs, robots.txt directives, IP blocking, and other protections. Treat the code below as an implementation pattern for a source you are authorized to collect from, or use it only after obtaining G2’s written permission for your project.

For legitimate programmatic access to G2 data, start with the official G2 API documentation. G2 describes that API as providing programmatic access to product, category, and review data. Confirm eligibility, pricing, licensing, rate limits, and permitted reuse with G2 before building a production integration.

JavaScript review scraper architecture

An authorized collector normally has six stages:

A reliable collector validates each page before parsing and stores provenance with every record.
A reliable collector validates each page before parsing and stores provenance with every record.
  1. Build a URL. Keep the product slug, page number, locale, and any approved filters explicit.
  2. Request the document. Use a normal HTTP client first when the authorized source returns review HTML in the response.
  3. Validate the response. HTTP 200 alone does not prove that reviews were returned. Check the final URL, content type, expected markers, and the number of matching cards.
  4. Parse repeated records. Extract stable fields such as title, rating, review text, public role or segment, posting date, product name, and average rating.
  5. Follow pagination. Discover the next approved page URL, stop when it is absent, and deduplicate records.
  6. Persist with provenance. Store the source URL, retrieval time, parser version, and raw response hash so results can be audited or reprocessed.

Selectors and page behavior can change. A parser should fail visibly when it finds zero cards or a different document shape instead of silently producing an empty dataset.

Before you write code: permission and data-use checks

Use this checklist before automating any G2 page:

  • Obtain G2’s express prior written consent for automated extraction, unless your agreement explicitly grants the required access.
  • Ask whether reviews, reviewer identities, ratings, and derived fields may be stored, transformed, displayed, or redistributed.
  • Confirm an allowed request rate, pagination limit, retention period, and authentication method.
  • Prefer G2’s official API when it meets your use case. The API’s documented availability does not by itself establish that every project has access or redistribution rights.
  • Do not use proxy rotation, CAPTCHA solving, stealth settings, undocumented private endpoints, or other methods intended to defeat access controls.

Complete Node.js example with Cheerio

The following example demonstrates the mechanics against an authorized reviews URL. Replace the URL and selectors with values supplied or approved by the site owner. It uses Node.js 20+, native fetch, and Cheerio.

import * as cheerio from 'cheerio';
import crypto from 'node:crypto';

const startUrl = process.env.REVIEWS_URL;
if (!startUrl) throw new Error('Set REVIEWS_URL to an authorized reviews URL');

const maxPages = Number(process.env.MAX_PAGES || 10);
const pauseMs = Number(process.env.PAUSE_MS || 1000);

const sleep = ms => new Promise(resolve => setTimeout(resolve, ms));
const absolute = (href, base) => new URL(href, base).toString();

async function fetchHtml(url) {
  const response = await fetch(url, {
    headers: {
      'accept': 'text/html,application/xhtml+xml',
      'user-agent': 'AuthorizedResearchBot/1.0 (contact: research@example.com)'
    },
    signal: AbortSignal.timeout(30_000)
  });

  const contentType = response.headers.get('content-type') || '';
  const html = await response.text();
  if (!response.ok) throw new Error(`HTTP ${response.status} for ${url}`);
  if (!contentType.includes('text/html')) {
    throw new Error(`Expected HTML but received ${contentType}`);
  }
  if (!html.toLowerCase().includes('review')) {
    throw new Error(`Response does not contain an expected review marker: ${url}`);
  }
  return { html, finalUrl: response.url };
}

function parsePage(html, pageUrl) {
  const $ = cheerio.load(html);
  const reviews = [];

  // Use selectors approved for your authorized source. These are examples.
  $('.review-card').each((_, element) => {
    const card = $(element);
    const ratingText = card.find('.review-rating').attr('aria-label') ||
      card.find('.review-rating').text().trim();
    reviews.push({
      title: card.find('.review-title').text().trim() || null,
      rating: ratingText || null,
      reviewText: card.find('.review-text').text().trim() || null,
      roleOrSegment: card.find('.reviewer-role').text().trim() || null,
      postedAt: card.find('time').attr('datetime') || card.find('time').text().trim() || null,
      productName: card.find('.product-name').text().trim() || null,
      averageRating: card.find('.average-rating').text().trim() || null,
      sourceUrl: pageUrl
    });
  });

  const nextHref = $('a[rel="next"]').attr('href') ||
    $('a.next-page').attr('href') || null;
  return { reviews, nextUrl: nextHref ? absolute(nextHref, pageUrl) : null };
}

async function crawl() {
  const output = [];
  const seenUrls = new Set();
  const seenRecords = new Set();
  let url = startUrl;

  for (let page = 1; page <= maxPages && url; page++) {
    if (seenUrls.has(url)) throw new Error(`Pagination loop detected at ${url}`);
    seenUrls.add(url);

    const { html, finalUrl } = await fetchHtml(url);
    const parsed = parsePage(html, finalUrl);
    if (parsed.reviews.length === 0) {
      throw new Error(`No review cards parsed on page ${page}: ${finalUrl}`);
    }

    for (const review of parsed.reviews) {
      const key = crypto.createHash('sha256')
        .update(JSON.stringify(review)).digest('hex');
      if (!seenRecords.has(key)) {
        seenRecords.add(key);
        output.push({ ...review, page });
      }
    }

    url = parsed.nextUrl;
    if (url) await sleep(pauseMs);
  }

  return output;
}

const reviews = await crawl();
console.log(JSON.stringify({ count: reviews.length, reviews }, null, 2));

Install Cheerio with npm install cheerio, then run:

REVIEWS_URL='https://authorized.example/reviews' MAX_PAGES=5 node scrape.mjs > reviews.json

What each validation step protects against

Check Failure it catches Recommended response
Status and final URL Redirect to login, consent, or an error page Stop and record the redirect chain
Content type JSON, PDF, or a block page returned instead of HTML Route to the correct parser or stop
Content marker HTTP 200 shell with no review data Save a diagnostic sample and investigate
Card count Selector drift or an empty page Alert, do not emit an empty success result
Pagination loop check Repeated next links or canonical URL changes Stop after a repeated URL

How do I handle pagination across G2 review pages?

Pagination is usually represented by a next link or a page-number query such as ?page=2. In an authorized integration, prefer the canonical next link supplied by the page. If your approved contract specifies page numbers, construct them with URL rather than string concatenation:

const pageUrl = new URL('https://authorized.example/reviews');
pageUrl.searchParams.set('page', String(pageNumber));
const response = await fetch(pageUrl);

Stop when there is no next link, the API or agreement’s maximum is reached, or a page returns no new record identifiers. Keep a set of visited URLs and a set of stable record keys. A delay between requests helps you stay within an approved rate limit; it is not a way to bypass one. Save the page URL for every record so a later correction can trace the source.

How do I parse ratings and review text safely?

Text often contains whitespace, hidden accessibility labels, and localized formats. Keep the original string as well as a normalized value. For example:

function cleanText(value) {
  return value?.replace(/\\s+/g, ' ').trim() || null;
}

function parseNumericRating(value) {
  const match = String(value || '').replace(',', '.').match(/\\d+(?:\\.\\d+)?/);
  return match ? Number(match[0]) : null;
}

const rawRating = card.find('[aria-label*="star"], .review-rating').first().attr('aria-label')
  || card.find('.review-rating').text();
const record = {
  ratingRaw: cleanText(rawRating),
  rating: parseNumericRating(rawRating),
  reviewText: cleanText(card.find('.review-text').text())
};

Do not infer missing ratings, reviewer identities, sentiment, or product attributes. Preserve null and document any derived field separately from source fields.

When plain HTTP is not enough

A plain request can return a document that is different from what a browser displays: a loading shell, a consent page, a login response, or HTML without hydrated review data. Inspect the authorized response body and compare it with the permitted browser view. If the site owner provides a supported rendered-access method, use that method under the written agreement. Do not inspect or call undocumented endpoints to evade G2’s restrictions.

For a compliant production project, the official G2 API is generally easier to reason about than scraping page markup. Ask G2 for the endpoint, authentication, fields, pagination model, quotas, retention rules, and redistribution terms that apply to your account.

Reliability, performance, and cost planning

  • Bound the crawl. Set a maximum page count and an overall deadline. A bad next link should not run forever.
  • Use bounded retries. Retry transient 429 or 5xx responses only when your agreement permits it, with exponential backoff and a maximum attempt count. Never retry a permission or authentication error.
  • Separate fetching from parsing. Store raw authorized responses or hashes so selector changes can be replayed without re-requesting content.
  • Control concurrency. Sequential requests are simplest. If parallelism is approved, use a small queue and honor the documented rate limit.
  • Measure useful work. Track pages requested, pages parsed, cards found, duplicates, retries, bytes, and elapsed time. A high HTTP success rate with zero cards is a parser failure.
  • Budget API usage. Confirm whether G2 charges per request, per record, or by plan. The research available here does not establish API pricing or universal eligibility.
  • Protect data. Reviews can contain personal information. Restrict logs, encrypt stored data, define retention, and follow the data-use terms in your authorization.

Troubleshooting common errors

Symptom Likely cause Fix
403 or 401 Missing or invalid authorization Use the approved credential or contact G2; do not attempt to bypass the response.
429 Rate limit exceeded Stop, follow the documented retry-after value, and reduce approved concurrency.
200 but zero reviews Login, consent, bot-check, empty shell, or selector drift Log final URL and a redacted response sample; verify access and selectors with the owner.
Parser works, then breaks Markup or class names changed Use selectors supplied by the source, add fixture tests, and alert on zero-card pages.
Duplicate reviews Overlapping pages or changing sort order Deduplicate with a stable authorized identifier, or a documented hash fallback.
Pagination repeats Canonical link or query normalization loop Normalize URLs and stop when a URL has already been visited.
Encoding looks wrong Incorrect charset handling Use the response charset and retain the raw value for diagnosis.

Or skip the browser setup

If your goal is to capture the page itself for documentation or a permitted review workflow, ScreenshotNeo provides a one-call website screenshot API. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Only clean shots are billed: bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers.

See the ScreenshotNeo API docs for all options. cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also offers an MCP server so Claude, Cursor, and other MCP clients can call take_screenshot, get_page_info, and capture_pdf. Plans include 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Can I scrape G2 reviews if the pages are publicly visible?

Public visibility does not remove G2’s automated-extraction restriction. Obtain express prior written consent or use an authorized API agreement.

ScreenshotNeo removes common consent and overlay elements before producing a clean capture.
ScreenshotNeo removes common consent and overlay elements before producing a clean capture.

Does G2’s API guarantee access to every review field?

No. The documentation describes programmatic product, category, and review access, but field availability, eligibility, pricing, and reuse rights must be confirmed for your project.

Should I use Playwright for G2?

Only if G2 authorizes rendered browser access and provides an approved method. A browser does not override the Terms of Use or access controls.

How can I tell whether an empty result is real?

Check the status, final URL, content type, expected markers, selector count, and pagination state. Store a diagnostic response and stop when validation fails.