ScreenshotNeo

BlogHow-to

How to Scrape Glassdoor Data with Browser Automation

Glassdoor’s terms require express written permission for automated scraping. Learn how to verify authorization and use Playwright safely on a site you control.

By the ScreenshotNeo team4 October 20268 min read

Direct answer: Do not scrape Glassdoor with browser automation unless you have Glassdoor’s express written permission for the specific access and use. Glassdoor’s Terms of Use prohibit using software or automated agents to scrape, strip, or mine service data without that permission. Ordinary use is described as personal and non-commercial unless Glassdoor separately agrees to commercial use. A page being publicly viewable does not itself authorize automated collection.

This guide shows how to check the authorization path and how to write a small Playwright data-extraction example against a page you own or are authorized to automate. It does not provide Glassdoor selectors, a Glassdoor scraper, or instructions for bypassing access controls. Read Glassdoor’s Terms of Use and confirm the currently applicable terms for your jurisdiction before collecting data.

1. Establish permission before choosing a tool

Before writing code, define the use case well enough that Glassdoor can give a meaningful answer. Identify the fields you need—such as company ratings, reviews, salary information, or job listings—along with the companies, geography, date range, expected volume, collection frequency, storage period, and intended analysis. Say whether you plan to publish, redistribute, or use the results commercially.

Request express written authorization that covers those details. Ask whether the intended fields and method are allowed, what limits apply, and whether a separate commercial agreement is needed. Stop if permission is denied or unclear; do not infer authorization from public visibility or technical access.

Check for an authorized data route

Glassdoor’s terms refer to APIs or formal embedding features where Glassdoor makes them available. A surfaced API terms document does not establish that a public API is currently available, that any reader can obtain access, that reviews are included, or that a particular use is allowed. Ask Glassdoor directly for current documentation and written confirmation of eligibility, fields, geography, volume limits, retention, redistribution rights, and commercial terms.

If more than one authorized route is offered, compare them on those criteria. Do not assume any route has particular coverage, freshness, limits, or price until Glassdoor confirms it.

2. Build a permission-safe browser automation example

The following Node.js example uses Playwright to read a small HTML page created inside the script. It demonstrates locating repeated records and extracting their visible fields without sending requests to Glassdoor or any third-party site. Run this pattern against a site you control, a test fixture, or a target that has explicitly authorized your automation.

Install and run

  1. Install Node.js and create a project: npm init -y.
  2. Install Playwright: npm install playwright.
  3. Save the script below as extract.mjs, then run node extract.mjs.
import { chromium } from 'playwright';

const html = `
  <!doctype html>
  <html lang="en">
    <meta charset="utf-8">
    <title>Authorized demo directory</title>
    <main>
      <article aria-label="Company record">
        <h2>Northwind Studio</h2>
        <p>Rating: 4.2</p>
      </article>
      <article aria-label="Company record">
        <h2>Contoso Works</h2>
        <p>Rating: 3.8</p>
      </article>
    </main>
  </html>
`;

const browser = await chromium.launch({ headless: true });
try {
  const page = await browser.newPage();
  await page.setContent(html);

  const records = await page.getByRole('article', {
    name: 'Company record'
  }).evaluateAll((articles) => articles.map((article) => ({
    name: article.querySelector('h2')?.textContent?.trim() ?? null,
    ratingText: article.querySelector('p')?.textContent?.trim() ?? null
  })));

  console.log(JSON.stringify(records, null, 2));
} finally {
  await browser.close();
}

The output is a JSON array of the two records. The page is deliberately an in-memory fixture: replace it only with an authorized page and adapt the fields to that page’s permitted data contract. For a real authorized site, prefer accessible roles, labels, and visible text where they express the interface. Playwright recommends user-facing locators and explains that locators re-resolve elements and provide auto-waiting; selectors tied to the DOM’s internal structure can break when it changes. See the Playwright locator guide.

Adaptation checklist for an authorized site

  • Confirm the written permission covers the page, fields, frequency, and downstream use.
  • Use stable, user-facing locators or a test ID contract provided by the site owner; avoid long CSS or XPath chains based on incidental page structure.
  • Wait for the specific record region you need rather than sleeping for an arbitrary long interval. Set a reasonable timeout and handle missing or delayed content.
  • Extract only approved fields. Keep collection date, source, geography, field definitions, and permission reference with the resulting data.
  • Use a small, approved request rate and stop if the site returns an access denial, a bot check, or another signal that the activity is not permitted.

3. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. It captures a permitted page as PNG, JPEG, WebP, or PDF; it is not a Glassdoor data API and does not grant permission to access or collect Glassdoor content. For a page you are authorized to capture, one GET request can return an image. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
import { writeFile } from 'node:fs/promises';

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. These capture features do not bypass Glassdoor’s terms or replace authorization for data collection.

Sign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.

4. Validate and protect authorized results

Permission to collect data does not make user-generated content verified fact. Glassdoor does not guarantee that service content is accurate, current, suitable, reliable, or high quality. Preserve provenance and communicate those limits when presenting analysis.

  • Store the collection date, source page, geography, field definitions, and permission basis with each batch.
  • Keep raw values distinct from normalized or derived values so later corrections can be traced.
  • Record missing or ambiguous fields as missing; do not silently infer them.
  • Apply the retention and redistribution limits in the written authorization and any applicable agreement.
  • Reconfirm permission and terms when the purpose, data fields, scale, or publication plan changes.

5. Reliability, performance, and cost

Browser automation depends on a live page and its changing interface. Authorized collection can fail when content loads late, a locator matches multiple elements, a field is absent, or the page layout changes. Playwright’s auto-waiting helps with actionability, but it cannot guarantee that a site will keep the same content or interface. Use bounded timeouts, validate required fields, log failures without storing unnecessary personal data, and review changes against the authorized scope.

Do not estimate Glassdoor collection cost, volume capacity, or performance from browser mechanics alone. Confirm any authorized limits and commercial terms directly with Glassdoor. For your own approved workflow, account for browser runtime, retries, review, storage, and maintenance; more concurrency is not automatically appropriate and must remain within the permission granted.

6. Troubleshooting

Symptom Likely cause What to do
Permission is not explicit or the reply is ambiguous The intended automated access or use has not been authorized clearly. Pause collection. Ask Glassdoor to confirm the exact fields, purpose, frequency, retention, and reuse in writing.
You cannot confirm API access or review coverage The existence of API terms does not establish availability, eligibility, or which fields are offered. Request current API or licensing documentation and written confirmation directly from Glassdoor.
Playwright reports a locator timeout The target is absent, delayed, hidden, or the locator no longer matches. On an authorized page, inspect the page state and use an appropriate user-facing locator with a bounded timeout. Do not treat a timeout on Glassdoor as a reason to evade a block.
A locator matches more than one element The locator is too broad for the page. Narrow it using a meaningful accessible name or a parent region and verify the expected count before extracting.
Extracted fields are empty or inconsistent The page may render later, the data may be missing, or the page structure may have changed. Validate required fields, preserve nulls, record the source and collection time, and check the authorized page contract.
A bot check, access denial, or CAPTCHA appears The service is restricting the request or the permitted access path does not cover it. Stop automated access and contact Glassdoor about authorization. Do not bypass the control.
Data looks current but later changes Service content may change and Glassdoor does not guarantee its accuracy or currency. Keep timestamps and provenance, avoid presenting submissions as verified facts, and refresh only if the authorization permits it.

FAQ

Can I scrape Glassdoor reviews that are visible without signing in?

Visible access does not by itself grant permission for automated scraping. Glassdoor’s terms require express written permission for scraping, stripping, or mining service data.

Does Glassdoor have a public API for reviews?

The research available for this guide does not establish current public API availability, review coverage, or eligibility. Ask Glassdoor for current documentation and written confirmation for your use case.

Does using Playwright make collection permitted?

No. Playwright is a browser automation tool. Its documentation explains how to build more resilient automation; it does not authorize access to a particular service.

Can ScreenshotNeo collect Glassdoor data?

ScreenshotNeo captures page images or PDFs; it is not a Glassdoor data API. Use it only for pages you are authorized to capture, and obtain Glassdoor’s express written permission before any automated access covered by its terms.

Sources