ScreenshotNeo

BlogHow-to

How to Scrape Job Postings with an AI Job Board Scraper

Build a job-posting pipeline that starts with permission, parses predictable fields with rules, and uses AI only for supported extraction.

By the ScreenshotNeo team29 September 202612 min read

How to Scrape Job Postings with an AI Job Board Scraper

An AI job board scraper should collect only postings you are authorized to use, turn permitted records into a stable schema, and use a model to structure the parts that rules cannot reliably classify. Start with an official API, an approved partner feed, an employer-authorized first-party career page, or a licensed dataset. A page being visible in a browser is not blanket permission to crawl, store, or redistribute its contents.

This guide shows how to design the pipeline, normalize job data, add evidence-backed AI extraction, and operate it responsibly. It does not provide a method to bypass a job board’s access controls or terms. Platform policies and API availability can change; check the current terms and obtain the required authorization before you collect data.

1. Decide whether you may collect the data

Before writing a crawler, record the source, the permission that covers it, which fields you may collect, the allowed geography, rate limits, retention period, permitted uses, and who handles deletion requests. Keep a copy of the applicable agreement or written authorization. If one of those points is unclear, resolve it with the source or your legal reviewer before ingesting records.

Prefer, in order, a suitable official API, an approved partner feed, a publisher integration, or an employer’s first-party careers page with explicit authorization. An API is not automatically unrestricted: access can be limited by agreement, scope, quotas, storage rights, or product purpose. Keep the source identifier and retrieval time with every record so you can trace it back and remove it later.

  • LinkedIn: LinkedIn says it does not allow third-party software that scrapes, changes the appearance of, or automates activity on its website. Its Job Posting API is restricted to approved developers, and Microsoft’s current overview says it is not accepting new partnerships for that API and points applicants to Apply Connect. An unaffiliated scraper should not crawl LinkedIn; pursue an approved route or use a different authorized source. See LinkedIn’s automated activity policy and the Job Posting API overview.
  • Indeed: Indeed documents APIs and partner integrations, but access is conditional on the relevant agreement and documentation. Its agreement restricts how API data may be queried, retained, and used. Request the appropriate access and scope, and get clarification in writing before storing or redistributing records. Read the Indeed Partner Developer Agreement and Indeed developer documentation.

These requirements are product constraints as much as legal questions. If you cannot establish permission and a permitted use, do not run a browser automation job against that source.

2. Design the collection-to-AI pipeline

Keep acquisition, parsing, AI extraction, validation, and publishing as separate stages. That makes it possible to stop collection without losing the ability to review existing records, to reprocess permitted data when your schema changes, and to show why a field has a particular value.

Separate authorized collection, deterministic normalization, AI extraction, and human review so each step can be audited.
Separate authorized collection, deterministic normalization, AI extraction, and human review so each step can be audited.
  1. Collect permitted records. Call the approved API or read the authorized feed. Apply its pagination, quota, and rate-limit instructions. Save the source ID or canonical URL and retrieval timestamp. Do not collect candidate profiles or account data unless your agreement explicitly permits it.
  2. Keep the permitted source payload. Retain raw content only if the source agreement allows it and only for the approved retention period. Otherwise keep the fields you are allowed to retain, plus a source reference and the minimum evidence needed to audit extraction.
  3. Normalize with deterministic code. Standardize whitespace, URLs, currencies, dates, locations, and enumerated values before calling a model. Rules are cheaper and easier to debug for predictable transformations.
  4. Use AI for genuinely ambiguous fields. Ask it to extract skills, classify seniority, or map a location phrase to a known category. Require evidence text from the permitted description. Treat model output as a suggestion until it passes validation.
  5. Validate and route exceptions. Reject a record without a source reference or employer. Flag contradictory salary values, invalid dates, or an evidence span that does not support the extracted value. Send low-confidence or legally sensitive cases to a reviewer.
  6. Deduplicate and expire. Prefer the source’s stable ID. Otherwise use a normalized canonical URL with employer, title, location, and posting date as supporting keys. Re-check freshness and expire records as the source agreement requires.
  7. Monitor and honor deletion. Track quota use, parser failures, schema drift, duplicate rate, extraction confidence, and deletion turnaround. Pause ingestion if the source changes its terms or access status.

3. Define a stable job schema

Use one canonical internal format even if sources name fields differently. Keep the original compensation text as well as any parsed range; a numeric salary without its currency, period, or source context can be misleading. Unknown values should be null, not guessed.

{
  "source": "authorized_feed_name",
  "source_id": "feed-record-123",
  "canonical_url": "https://careers.example.com/jobs/123",
  "retrieved_at": "2026-09-29T12:00:00Z",
  "title": "Senior Data Engineer",
  "employer": "Example Co",
  "location_text": "Remote, United States",
  "remote_status": "remote",
  "employment_type": "full_time",
  "compensation_text": "$140,000-$170,000 per year",
  "salary_min": 140000,
  "salary_max": 170000,
  "salary_currency": "USD",
  "salary_period": "year",
  "skills": ["Python", "SQL"],
  "seniority": "senior",
  "posted_at": "2026-09-20",
  "application_url": "https://careers.example.com/jobs/123/apply",
  "expires_at": null,
  "extraction": {
    "model": "model-name-and-version",
    "prompt_version": "job-fields-v1",
    "confidence": 0.92,
    "evidence": {
      "skills": ["Build data services in Python and SQL"]
    }
  }
}

Choose controlled values for fields such as remote status and employment type, and document their meanings. Keep the source’s wording in separate fields when a controlled category loses useful nuance. The schema should also make provenance and expiration visible instead of burying them in logs.

4. Parse and validate an authorized feed in Python

The following runnable example reads a local JSON file containing records from a feed you are authorized to use. It normalizes a few fields, flags incomplete entries, and writes a clean JSON Lines file. It deliberately does not fetch job-board pages. Save it as normalize_jobs.py and run python normalize_jobs.py input.json output.jsonl. The input is an array of objects with fields such as id, url, title, company, location, description, and posted_at.

import json
import re
import sys
from datetime import datetime, timezone
from urllib.parse import urlsplit, urlunsplit


def clean_text(value):
    if not isinstance(value, str):
        return None
    return re.sub(r"\s+", " ", value).strip() or None


def canonical_url(value):
    value = clean_text(value)
    if not value:
        return None
    parts = urlsplit(value)
    if parts.scheme not in ("http", "https") or not parts.netloc:
        return None
    # Drop fragments; retain query parameters because the source may use them
    # to identify a particular posting.
    return urlunsplit((parts.scheme.lower(), parts.netloc.lower(),
                       parts.path.rstrip("/"), parts.query, ""))


def normalize(record, source_name):
    url = canonical_url(record.get("url"))
    title = clean_text(record.get("title"))
    employer = clean_text(record.get("company"))
    if not url or not title or not employer:
        return None, "missing canonical URL, title, or employer"

    posted = clean_text(record.get("posted_at"))
    # Preserve unparseable source text for review; never make up a date.
    posted_iso = None
    if posted:
        try:
            posted_iso = datetime.fromisoformat(
                posted.replace("Z", "+00:00")
            ).date().isoformat()
        except ValueError:
            pass

    return ({
        "source": source_name,
        "source_id": clean_text(record.get("id")),
        "canonical_url": url,
        "retrieved_at": datetime.now(timezone.utc).isoformat(),
        "title": title,
        "employer": employer,
        "location_text": clean_text(record.get("location")),
        "compensation_text": clean_text(record.get("salary")),
        "description": clean_text(record.get("description")),
        "posted_at": posted_iso,
        "application_url": canonical_url(record.get("application_url")),
        "extraction": None
    }, None)


def main():
    if len(sys.argv) != 3:
        raise SystemExit("Usage: python normalize_jobs.py input.json output.jsonl")
    with open(sys.argv[1], encoding="utf-8") as f:
        records = json.load(f)
    if not isinstance(records, list):
        raise SystemExit("Input must be a JSON array")

    accepted, rejected = 0, 0
    with open(sys.argv[2], "w", encoding="utf-8") as out:
        for index, item in enumerate(records):
            if not isinstance(item, dict):
                rejected += 1
                print(f"row {index}: not an object", file=sys.stderr)
                continue
            job, error = normalize(item, "authorized_feed")
            if error:
                rejected += 1
                print(f"row {index}: {error}", file=sys.stderr)
                continue
            out.write(json.dumps(job, ensure_ascii=False) + "\\n")
            accepted += 1
    print(f"accepted={accepted} rejected={rejected}", file=sys.stderr)


if __name__ == "__main__":
    main()

In production, use the actual source identifier and its timestamp where available, rather than replacing them with locally generated values. Add schema validation, structured logs, and tests for the source’s documented response format. The example keeps a description field only to illustrate an authorized local feed; omit it or delete it according to the source’s storage and retention terms.

5. Add AI extraction with evidence and review

Give the model only the permitted fields needed for the task. Ask for a constrained JSON object, provide the allowed categories, and explicitly require null when the text does not support a value. Require one or more short evidence spans for extracted skills and classifications. Store the model name and version, prompt version, confidence or review status, and the source spans permitted by your retention rules.

For example, an extraction request can ask for:

{
  "skills": [
    {"name": "Python", "evidence": "Build services with Python"}
  ],
  "seniority": {
    "value": "senior",
    "evidence": "Senior Data Engineer",
    "confidence": 0.91
  },
  "remote_status": {
    "value": "remote",
    "evidence": "This position is remote within the United States",
    "confidence": 0.96
  }
}

Validate the response against a schema before saving it. Reject unexpected keys and category values. Check that each cited span occurs in the allowed source text, while recognizing that a matching span alone does not prove the interpretation is correct. Send low-confidence, conflicting, or consequential classifications to a person. Preserve a clean distinction between extraction for search and decisions about people: do not use the scraper to infer protected traits or silently rank candidates.

Indeed’s AI and Automated Employment Decision Tools FAQ lists prohibited uses and places responsibility for job-posting content on employers. Keep the system focused on organizing job-ad text, and obtain appropriate legal review before connecting it to hiring decisions or candidate evaluation.

6. Deduplication, expiry, and monitoring

Use a stable source ID as the primary key whenever possible. If there is no source ID, canonicalize the URL and combine it with employer and title; use location and posting date as additional evidence. Similarity matching can suggest possible duplicates, but do not automatically merge records from separate authorized sources if that would erase provenance, attribution, or different expiration terms.

Deduplication and expiration should preserve source provenance and carry deletion through every store.
Deduplication and expiration should preserve source provenance and carry deletion through every store.

Expiration should be an explicit state transition, not a permanent deletion by default. Depending on the agreement, remove a record, mark it expired, or retain only a minimal tombstone to prevent accidental re-import. Re-check freshness on a schedule permitted by the source, and make deletion requests propagate through raw storage, search indexes, caches, and derived AI outputs.

Useful operational metrics include fetch success and quota use, records accepted or rejected, parser failures by source version, duplicate candidates, stale postings, AI validation failures, confidence distribution, and deletion turnaround. Alert on sudden shifts. A page layout change, API field rename, expired credential, rate-limit response, and a real drop in postings can look similar unless you track these signals separately.

7. Costs, performance, and reliability

Keep collection close to the source’s allowed cadence. Fetching only changed records, honoring pagination and conditional requests where supported, and avoiding repeated processing of unchanged text reduce both load and compute expense. Apply exponential backoff for transient errors only within the source’s limits; stop on authorization failures, revoked credentials, or a policy change instead of retrying aggressively.

AI calls usually cost more than simple normalization and can add latency. Use deterministic parsers for dates, URLs, obvious salary formats, and stable enumerations; reserve model calls for fields that need interpretation. Cache extraction results by a content hash only while the source terms allow retention, and include the model and prompt version in the key so changed instructions trigger deliberate reprocessing. Batch where the model interface permits it, but keep per-record outcomes and evidence so one bad response does not poison a whole batch.

Reliability comes from idempotent ingestion, bounded queues, retry rules, schema validation, and a dead-letter path for records that repeatedly fail. Store checkpoints so a restart does not duplicate a page of results. Use least-privilege credentials, encrypted secrets, restricted staff access, and audit logs. Maintain a kill switch per source. On a terms update, revoked API scope, unexpected CAPTCHA, or changed access policy, pause that source and review before resuming.

Budget for the approved source’s fees or partner requirements, storage and deletion work, parser maintenance, model calls, and human review. There is no universal price or accuracy figure: quotas, access, fields, and model behavior depend on the particular source and implementation. Measure your own permitted workload before committing to a broad ingestion schedule.

8. Troubleshooting

Symptom Likely cause What to do
401 or 403 response Missing, expired, or insufficiently scoped credentials; access may not be approved. Check the official integration instructions and credential scope. Do not switch to browser automation to evade the denial.
429 or quota errors Request frequency or volume exceeded the permitted limit. Honor the documented retry delay, reduce concurrency, checkpoint progress, and request a quota change through the approved channel if needed.
Empty results Wrong query, pagination cursor, geography, or API scope; alternatively, no matching postings exist. Log request parameters without secrets, inspect the documented response and pagination state, and verify access with a small permitted query.
Parser suddenly rejects many rows Schema drift, changed field names, a new date format, or malformed feed data. Quarantine samples, compare them with the documented schema, version the parser, and alert before silently dropping records.
Salary or location looks wrong Ambiguous ranges, currencies, pay periods, multiple locations, or unsupported AI inference. Keep the source wording, require evidence, validate units, and mark ambiguous fields unknown for review.
Duplicates remain Source IDs are missing or URL parameters differ across references. Normalize only parameters known to be tracking fields; combine URL with employer/title and retain each source provenance.
Model returns invalid JSON or unsupported skills Open-ended prompt, output truncation, or extraction without sufficient context. Use schema-constrained output where available, cap input, validate evidence spans, retry a bounded number of times, then route to review.
Records return after deletion Re-ingestion, cache, index, or derived output was not included in deletion handling. Use a deletion ledger or tombstone where permitted, propagate deletion to every store, and verify it against the retention policy.

9. Or skip the browser setup

For a visual snapshot of a page you are authorized to inspect, ScreenshotNeo is a website screenshot API and MCP server. It can help document the appearance of a permitted career page; it does not grant permission to collect, store, or republish job-board data, and a screenshot is not a replacement for an approved jobs API or feed. One GET request returns an image or PDF. See the ScreenshotNeo API docs.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before a capture; those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes screenshot, page-info, and PDF tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. All features are on every plan.

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

10. Frequently asked questions

Can AI scrape a job board by itself?

A model can interpret content you are authorized to process, but it does not create permission to access a site or remove contractual, privacy, and retention limits. Keep source access and model extraction as separate steps.

No. Use a qualified reviewer to resolve unclear rights or permitted-use questions. A model can help extract facts from a posting, but its output is not legal authorization.

What if there is no API for my source?

Look for an approved partner feed, a publisher integration, a licensed data provider, or written employer authorization for a first-party page. If none grants the access and use you need, choose another source.

How do I keep extracted fields explainable?

Store source provenance, model and prompt versions, validation status, and permitted evidence spans. Make unknown and disputed values visible, and provide a review path rather than silently filling gaps.

Can I use the resulting data to rank candidates?

This guide covers job-posting organization, not candidate evaluation. Do not infer protected traits or connect extraction to hiring decisions without a separately documented, legally reviewed process.