ScreenshotNeo

BlogHow-to

How to Scrape Indeed Jobs and Company Profiles

Learn how to collect Indeed job and employer data through authorized APIs and integrations, with practical schemas, controls, and implementation guidance.

By the ScreenshotNeo team29 September 202610 min read

How to Scrape Indeed Jobs and Company Profiles

Short answer: do not begin by sending automated requests to Indeed HTML pages. First define the fields and purpose, then check whether an approved Indeed API or partner integration covers your use case. Indeed’s current legal and developer documents restrict scraping, permanent copies of user or job-seeker content, algorithmic harvesting queries, and attempts to bypass limits or security controls. If your project qualifies, build against the documented program, preserve employer identity accurately, and design privacy, retention, and deletion controls from the start.

This guide explains how to collect job and company data responsibly, how to choose between an authorized Indeed route and employer-supplied data, and how to build a reliable ingestion pipeline without depending on undocumented page markup.

1. Start with authorization and data governance

Indeed says API users are bound by its Terms, Developer Terms, additional guidelines, and documentation. It may monitor and limit API calls and can restrict or terminate access. Review the current Indeed Terms and the relevant Developer Agreement before writing a collector.

The Developer Agreement prohibits scraping or building permanent copies of End User or Job Seeker content except where required for an approved integration or expressly permitted by documentation. It also prohibits algorithmic queries that replace human input and attempts to bypass limits or security measures. Those rules affect a job board, recruiting search engine, analytics warehouse, and internal reporting tool alike.

Define the exact fields and purpose

Write a short data specification before requesting access. Separate public job-ad fields from employer identifiers, application events, and job-seeker or profile information. For each field, record:

  • Purpose: search, recruiting workflow, analytics, employer reporting, or another approved use.
  • Source: an Indeed API, an employer or ATS, or another documented partner route.
  • Retention: transient display, cache with an expiry, or permanent storage if expressly permitted.
  • Owner: the team responsible for corrections, deletion requests, and access control.

A project that only needs a current search result may require a different integration from one that needs application events or employer data. Do not assume that a field visible in a browser is available for automated collection or long-term storage.

2. Choose an approved Indeed route

Indeed documents several APIs and integrations, including Employer Data, Job Sync, Indeed Apply, and Sponsored Jobs routes. Access, eligibility, limits, billing, and permitted purpose vary by program and by the current documentation. Start in the Indeed Developer Portal and identify the program that matches your role.

An approved API pipeline validates identity and privacy before storing job records.
An approved API pipeline validates identity and privacy before storing job records.
Route Typical data or workflow Questions to resolve
Employer Data Employer and ATS identity data used to link jobs to companies Which identifiers are required? How are mismatches corrected? What sharing disclosures apply?
Job Sync Employer or ATS job-feed synchronization Feed format, update cadence, removal behavior, validation, and retention rules
Indeed Apply Application events and application submission workflows HTTPS endpoints, personal-data handling, event security, and user-request processes
Sponsored Jobs Sponsored-job management and related reporting Eligibility, active sponsorship requirements, limits, and current charges

Indeed’s documentation can change. Record the version or date of the terms you reviewed, the approval received, the scopes granted, the request limits, and the deletion requirements. Indeed states that it may set and enforce API limits, disable access, or discontinue an API or feature.

3. Design a durable job and company schema

A stable internal model keeps your application independent from presentation changes and makes quality checks explicit. Keep source identifiers alongside your own IDs so updates and deletions can be reconciled.

Example job record

{
  "source": "indeed",
  "source_job_id": "documented-id",
  "employer_id": "internal-employer-id",
  "ats_id": "documented-ats-id",
  "title": "Example title",
  "location": {
    "raw": "Example location",
    "country": "US",
    "region": ""
  },
  "description": "Content returned by the approved integration",
  "employment_type": "",
  "posted_at": "",
  "updated_at": "",
  "status": "active",
  "source_url": "",
  "retrieved_at": "2026-09-29T00:00:00Z"
}

Example employer record

{
  "source": "indeed",
  "internal_employer_id": "employer-123",
  "ats_id": "documented-ats-id",
  "legal_name": "Example Company",
  "display_name": "Example Company",
  "website": "",
  "locations": [],
  "description": "",
  "source_updated_at": "",
  "last_verified_at": "2026-09-29T00:00:00Z",
  "status": "active"
}

Use a unique ATS and employer identifier pair when the Employer Data guidance requires it. Indeed warns against mismatching companies and asks clients to disclose sharing with data providers. Add a uniqueness constraint, an audit trail for identifier changes, and a review queue for uncertain matches.

4. Implement an authorized ingestion client

The examples below show the shape of a production client without inventing an undocumented endpoint. Replace the placeholder URL and authentication details with the endpoint, credentials, headers, and request schema supplied by your approved Indeed program.

Python: bounded requests with retries

import time
import requests

API_URL = "https://approved-indeed-endpoint.example/v1/jobs"
TOKEN = "YOUR_APPROVED_TOKEN"

session = requests.Session()
session.headers.update({
    "Authorization": f"Bearer {TOKEN}",
    "Accept": "application/json",
})

params = {"updated_since": "2026-09-28T00:00:00Z", "limit": 100}
for attempt in range(4):
    response = session.get(API_URL, params=params, timeout=30)
    if response.status_code == 200:
        payload = response.json()
        break
    if response.status_code in (429, 500, 502, 503, 504):
        wait = min(30, 2 ** attempt)
        time.sleep(wait)
        continue
    response.raise_for_status()
else:
    raise RuntimeError("Approved API remained unavailable after retries")

for job in payload.get("jobs", []):
    # Validate identifiers, normalize fields, and upsert by source_job_id.
    print(job.get("source_job_id"))

Node.js: pagination and backoff

const API_URL = 'https://approved-indeed-endpoint.example/v1/jobs';
const token = process.env.INDEED_APPROVED_TOKEN;

async function getPage(cursor) {
  const url = new URL(API_URL);
  url.searchParams.set('limit', '100');
  if (cursor) url.searchParams.set('cursor', cursor);

  for (let attempt = 0; attempt < 4; attempt++) {
    const res = await fetch(url, {
      headers: { Authorization: `Bearer ${token}`, Accept: 'application/json' }
    });
    if (res.ok) return res.json();
    if (![429, 500, 502, 503, 504].includes(res.status)) {
      throw new Error(`Approved API error ${res.status}`);
    }
    await new Promise(resolve => setTimeout(resolve, Math.min(30000, 2 ** attempt * 1000)));
  }
  throw new Error('Approved API remained unavailable after retries');
}

let cursor;
do {
  const page = await getPage(cursor);
  for (const job of page.jobs ?? []) {
    // Validate employer and ATS IDs before writing the record.
    console.log(job.source_job_id);
  }
  cursor = page.next_cursor;
} while (cursor);

cURL: inspect one approved response

curl --fail-with-body \
  -H "Authorization: Bearer YOUR_APPROVED_TOKEN" \
  -H "Accept: application/json" \
  "https://approved-indeed-endpoint.example/v1/jobs?limit=10"

Use the actual URL and parameters in your program documentation. Keep tokens in a secret manager, never in source control or browser code. Log request IDs, status codes, latency, and page cursors, but redact authorization headers and unnecessary personal data.

5. Validate identity, freshness, and quality

Employer identity is a core data-quality problem. A wrong company association can mislead candidates and damage search quality. Build these checks into every import:

  1. Require the documented ATS and employer identifiers.
  2. Reject records with an unknown or conflicting identifier pair.
  3. Run duplicate detection on source IDs, normalized names, domains, and locations.
  4. Keep job-to-employer links versioned so corrections are auditable.
  5. Mark jobs inactive when the approved feed says they were removed or closed.
  6. Track retrieved_at and source update timestamps separately.
  7. Send ambiguous matches to a human review queue rather than guessing.

Indeed’s Employer Data guidance says misleading or stale information can lead to removal. Schedule freshness checks according to the approved feed’s documented cadence, and expose a “last updated” value to users when appropriate.

6. Privacy and security controls

Publish a privacy notice that accurately explains what your application collects, how it handles and shares data, and the role of Indeed or other processors. Protect personal information under applicable law, honor data requests, and restrict staff and service access to the minimum necessary scope.

  • Use HTTPS for Indeed Apply endpoints and webhook receivers.
  • Encrypt stored credentials and rotate them according to your security policy.
  • Separate job metadata from applicant or profile data with separate permissions.
  • Define deletion and correction workflows before production launch.
  • Do not copy profile or job-seeker content into a permanent database unless the approved documentation permits it.
  • Keep an audit log of exports, administrative access, and deletion events.

7. If you are considering direct HTML scraping

Direct page collection is not a safe default. A browser-visible page does not grant permission to automate collection, store a permanent copy, or reproduce an Indeed experience. Do not work around login walls, bot controls, rate limits, CAPTCHAs, or other security devices. Do not issue algorithmic harvesting queries intended to replace human input.

If no authorized route covers your purpose, pause the build and reassess the product requirement. Options may include obtaining data directly from participating employers or ATS customers, using a documented Indeed integration, or reducing the fields and retention period. This is also the point to obtain legal advice for your specific jurisdiction and business model. Indeed’s Terms FAQ explicitly says its FAQs are for convenience, are not exhaustive or legal advice, and do not replace the binding Terms.

8. Performance, reliability, and cost planning

Performance

  • Request only fields and date ranges required by the approved scope.
  • Use documented pagination and incremental updates instead of repeatedly loading the full dataset.
  • Process pages in a bounded worker pool that stays within published limits.
  • Normalize and validate asynchronously so ingestion workers remain responsive.
  • Cache immutable or slowly changing employer attributes only when retention rules allow it.

Reliability

  • Retry only transient failures such as 429 and documented server errors.
  • Honor Retry-After when supplied and add exponential backoff with jitter.
  • Make writes idempotent using the source identifier and a version or update timestamp.
  • Persist cursors so a process restart resumes without duplicating records.
  • Alert on authentication failures, schema changes, unusual empty pages, and rising mismatch rates.

Cost and access risk

Budget for API charges where the selected program bills usage, plus storage, processing, monitoring, and review time. Sponsored Jobs access can depend on active monthly sponsorship spend and may incur charges described in current documentation. Also budget for approval and maintenance work: Indeed may change limits, disable access, or discontinue a feature, so keep a documented fallback such as employer-supplied feeds.

9. Or skip the browser setup

If your workflow needs screenshots of job pages or company pages for an authorized internal review, documentation, or visual QA process, ScreenshotNeo provides a single GET request that returns a PNG, JPEG, WebP, or PDF. It is separate from Indeed data access: you still need permission for the pages and must follow Indeed’s terms.

Consent banners and distracting overlays can be removed before an authorized screenshot.
Consent banners and distracting overlays can be removed before an authorized screenshot.

ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and each response reports the result with X-Page-Verdict and X-Billed headers. Its MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf.

See the ScreenshotNeo API documentation for options such as full-page capture, CSS-selector element capture, custom headers and cookies, wait conditions, blocking selected resources, PDF paper sizes and margins, caching TTLs, signed links, asynchronous jobs, bulk capture, and usage reporting.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

There are 1,000 screenshots each month on the free plan with no card. Paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.

10. Troubleshooting checklist

Symptom Likely cause Fix
401 or 403 from an Indeed endpoint Missing approval, expired token, wrong scope, or disabled access Check the approved program, rotate credentials, and contact the program owner. Do not attempt to bypass the response.
429 responses Published rate or request limit exceeded Slow the worker pool, honor retry guidance, and request a higher documented limit if available.
Duplicate companies Name-based matching without stable identifiers Use the documented ATS/employer pair and route ambiguous records to review.
Stale or closed jobs remain visible No removal sync or failed incremental job Process status changes, monitor cursors, and run a documented reconciliation job.
Applicant data appears in logs Verbose payload logging Redact payloads, reduce log fields, rotate exposed secrets, and review access.
Screenshot contains a popup The page element was not removed or the selector was incomplete Use ScreenshotNeo hide selectors or custom JavaScript and wait for the page to settle.

11. FAQ

Is there an Indeed API for job listings?

Indeed documents APIs and integrations, including Job Sync and Employer Data. Availability and eligibility depend on your use case and the current developer documentation; there is no universal assumption that every project can obtain the same fields.

Can I build a job board from Indeed listings?

Only if the approved program and documentation permit that purpose, display, retention, and sharing model. Directly copying listings into a permanent competing database can conflict with the Developer Agreement.

Can I scrape Indeed company pages?

A public page is not automatic permission for automated collection or permanent storage. Check the current Terms and use an approved API, partner integration, or employer-supplied data instead.

How should I handle an employer identity mismatch?

Stop the write, preserve the source identifiers, and send the record to review. Never attach a job to the closest company name by guesswork.

Where can I read the binding rules?

Start with the Indeed Terms, the Developer Agreement, and the Employer Data API Guidelines. The Terms FAQ is only a convenience summary.