Scrape Job Listings at Scale and Extract Insights with AI
Build a permission-aware pipeline for collecting job postings, extracting structured fields with AI, and analyzing results with clear provenance.

To scrape job listings at scale, first decide which sources you are authorized to access and how you may use the resulting data. Prefer licensed feeds or approved APIs, preserve each record’s source and retrieval time, normalize fields without guessing, and use AI to extract facts into a validated schema. An API integration for employers or ATS partners is not automatically a general-purpose job-search feed.
This guide describes a source-by-source workflow for public or licensed job-posting data. It does not assume that any job board permits unrestricted bulk collection. The right source depends on your geography, occupations, scale, intended use, and rights to retain or redistribute records.
1. Define what you are allowed to collect
Write down the purpose before choosing a scraper or model. Decide which occupations and geographies you need, which fields matter, how often records must be refreshed, how long you will keep them, and whether results will remain internal or be published or redistributed. Then review the current terms, API requirements, and data license for every source.
Access is a source-by-source decision. LinkedIn’s Job Posting API terms describe application vetting, approval, an access request with a specified use case, and continuing compliance responsibilities. Those terms apply to that API; they are not a universal rule for every job board. Check the [LinkedIn Job Posting API Terms](https://www.linkedin.com/legal/l/job-posting-api-terms) and confirm that your approved scope matches your actual use.
The cited LinkedIn and Indeed interfaces are posting integrations, not open search feeds. LinkedIn documents an API for members to post jobs from an ATS. Indeed describes Job Sync as an ATS-partner GraphQL API for creating, upserting, expiring, and checking the status of postings. Neither description establishes a general API for downloading all listings. See [LinkedIn’s API overview](https://learn.microsoft.com/en-us/linkedin/talent/job-postings/api/sync-job-postings?view=li-lts-2026-04) and [Indeed Job Sync API](https://docs.indeed.com/job-sync-api).
Rights to retrieve data do not automatically grant rights to retain it, publish it, or redistribute derived records. LinkedIn’s terms constrain use to the approved integration and describe data-type and onward-transfer restrictions. Indeed’s API guidelines discuss matching job data to the client’s career-site data and near-real-time updates in its ATS integration context. Those expectations should not be generalized into a freshness SLA for research datasets. Review [LinkedIn’s terms](https://www.linkedin.com/legal/l/job-posting-api-terms) and [Indeed’s API guidelines](https://docs.indeed.com/legal-terms/additional-api-terms-and-guidelines).
Robots.txt can communicate crawler preferences, but it is not permission to use copyrighted, contractual, or database-protected material. Crawler settings can also vary by bot and purpose; [OpenAI’s crawler documentation](https://developers.openai.com/api/docs/bots) is an example of that distinction, not a job-board policy.
Source approval checklist
- Identify the source owner, approved interface, and applicable license or terms.
- Confirm that your intended collection method and cadence are permitted.
- Record permitted fields, retention limits, redistribution limits, and attribution requirements.
- Keep evidence of approval and the version/date of the terms you reviewed.
- If permission is unclear, resolve it before collecting at scale.
2. Acquire through an allowed interface
Use a licensed feed or approved API where one is available for your purpose. Keep credentials in a secret manager or environment variable, not in source code. Record the source name, interface and version, retrieval timestamp, request or batch identifier, and any applicable source limits. Build retries around documented limits and transient failures rather than guessing a rate that might violate the source’s rules.

For example, LinkedIn’s cited API overview labels version 202604 as the latest version represented by that page and documents an application maximum of 100,000 requests per UTC day. It also lists promoted-job throttles of 2,000 records per minute and 60,000 records per day. These are limits for the documented integration, can change, and are not recommendations for scraping. Re-check the [current LinkedIn API overview](https://learn.microsoft.com/en-us/linkedin/talent/job-postings/api/sync-job-postings?view=li-lts-2026-04) before implementing that integration.
For an authorized feed that returns JSON, a small Python ingestion pattern can preserve provenance while writing newline-delimited JSON. Adapt the endpoint, authentication, pagination, and response field names to the feed’s official documentation; the example does not imply that any particular job board exposes this endpoint.
import json
import os
from datetime import datetime, timezone
from pathlib import Path
import requests
FEED_URL = os.environ["AUTHORIZED_FEED_URL"]
TOKEN = os.environ["AUTHORIZED_FEED_TOKEN"]
OUTPUT = Path("raw-postings.jsonl")
response = requests.get(
FEED_URL,
headers={"Authorization": f"Bearer {TOKEN}"},
timeout=(10, 60),
)
response.raise_for_status()
payload = response.json()
retrieved_at = datetime.now(timezone.utc).isoformat()
# Map this to the response shape documented by your licensed feed.
items = payload["jobs"]
with OUTPUT.open("a", encoding="utf-8") as out:
for item in items:
record = {
"source": "authorized-feed",
"source_listing_id": item.get("id"),
"retrieved_at": retrieved_at,
"raw": item,
}
out.write(json.dumps(record, ensure_ascii=False) + "\n")
For paginated APIs, follow the provider’s documented cursor or next-page mechanism. Persist the cursor after each successfully stored page so a restart does not lose progress. Handle 429 responses according to the documented retry guidance; for transient server errors, use bounded exponential backoff with jitter. Do not retry authorization failures indefinitely.
3. Preserve provenance and normalize records
Keep a raw, permitted representation or an authorized reference to the source record alongside normalized output. That lets you audit extraction, corrections, and deduplication decisions. Avoid collecting fields your project does not need, and honor source retention rules when deciding how long raw content can remain stored.
A practical normalized record often includes:
- Provenance: source, source listing ID, retrieval timestamp, source version, canonical URL, and raw-field references.
- Posting: title, employer, location, employment type, posting date, and description or permitted text reference.
- Compensation: original salary text, plus a numeric range and currency only when explicitly supported.
- Analysis: extracted skills, job family, remote or hybrid wording, evidence snippets, and extraction version.
- Lifecycle: first seen, last checked, duplicate group, and expiration or refresh status.
Normalize cautiously. Preserve the original employer and location strings as well as normalized forms. Do not turn “competitive salary” into a numeric estimate, infer a currency from geography without marking the assumption, or convert a vague date into a precise one. Use null or an ambiguity status when the source does not provide a reliable value.
4. Extract fields with AI and validate them
Give the model the posting text you are allowed to process and a schema that defines the fields, types, and allowed values. Request explicit unknown values and short evidence snippets for important extracted facts. Structured output can constrain formatting, but it cannot certify that a returned salary, skill, or location is actually present in the source. OpenAI’s [Structured Outputs documentation](https://developers.openai.com/api/docs/guides/structured-outputs) describes supported JSON Schema features, including strings, numbers, booleans, integers, objects, arrays, enums, anyOf, and selected string formats.

Here is a runnable extraction example using the OpenAI Python SDK and the Responses API. Set OPENAI_API_KEY in the environment and pass only text you are permitted to process. The schema is deliberately small; extend it to match your analysis and validate each field against source evidence.
import json
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
posting_text = """Paste an authorized job posting here."""
schema = {
"type": "object",
"properties": {
"title": {"type": ["string", "null"]},
"employer": {"type": ["string", "null"]},
"location": {"type": ["string", "null"]},
"employment_type": {
"type": ["string", "null"],
"enum": ["full-time", "part-time", "contract", "internship", "temporary", None]
},
"salary_text": {"type": ["string", "null"]},
"skills": {"type": "array", "items": {"type": "string"}},
"remote_wording": {"type": ["string", "null"]},
"evidence": {"type": "array", "items": {"type": "string"}}
},
"required": ["title", "employer", "location", "employment_type", "salary_text", "skills", "remote_wording", "evidence"],
"additionalProperties": False
}
result = client.responses.create(
model="gpt-4o-2024-08-06",
input=[
{"role": "system", "content": "Extract only facts stated in the posting. Use null for unknown values. Evidence must be short verbatim snippets from the supplied text."},
{"role": "user", "content": posting_text},
],
text={"format": {"type": "json_schema", "name": "job_posting", "strict": True, "schema": schema}},
)
record = json.loads(result.output_text)
print(json.dumps(record, indent=2, ensure_ascii=False))
In production, validate semantics after parsing: check that salary bounds are numeric and ordered if you derive them, dates parse under an explicit policy, enums match your taxonomy, and evidence snippets occur in the supplied text. Record the model and prompt/schema version for reproducibility. Route uncertain or high-impact cases to review rather than silently converting a plausible answer into a fact.
5. Deduplicate and keep records fresh
Start with stable source listing IDs when available. Across sources, compare normalized employer, title, location, and posting text, but preserve the basis for each merge. Near-identical descriptions can represent different locations or requisitions; an aggressive similarity threshold can erase real openings. Keep separate source records and attach a shared duplicate-group identifier if you need a consolidated view.
Refresh, expire, and delete records according to source terms and the purpose of the analysis. Track first-seen and last-checked timestamps. A missing listing may mean it expired, moved, or became temporarily unavailable; distinguish “not found on refresh” from a confirmed closure unless the source provides a reliable status.
6. Analyze with transparent limits
Once records are normalized, useful analyses include skill frequency, explicitly listed compensation ranges, location distribution, remote or hybrid wording, and changes over time. Report the source universe, geography, time window, missing-field rate, deduplication method, and sampling limits with every result. Job postings show demand in the observed sources; they do not measure the entire labor market or actual hiring outcomes.
Compare candidate feeds or integrations using source coverage and authorization model, occupation and geography coverage, field completeness, refresh cadence, historical depth and retention rights, API limits and reliability, duplicate and missing-field rates, and total integration and inference costs. Employer ATS posting workflows and licensed research feeds serve different purposes. The reviewed documentation does not establish a vendor ranking or comparative extraction-accuracy result.
7. Performance, reliability, and cost
Separate acquisition, normalization, extraction, and analysis into stages. Store an ingestion checkpoint and make writes idempotent using a source ID plus source name where permitted. This supports safe restarts and limits duplicate processing. Batch model requests where the API and task allow it, but keep enough per-record status to retry or inspect one failed item without rerunning an entire collection.
Measure the cost drivers directly: source access or licensing, storage and retention, model input and output volume, human review, and operational maintenance. Trim irrelevant text only when doing so preserves the evidence needed for the fields you extract. Cache extraction results keyed by a content hash and schema/prompt version when permitted; a changed posting or schema should trigger reprocessing. No general cost or accuracy figure applies across sources and models, so estimate from a representative authorized sample and monitor actual usage.
Reliability comes from treating source records and model output differently. Keep the original permitted values, attach extraction provenance, validate deterministic fields in code, and make uncertainty visible. Track failure categories such as authentication, throttling, malformed source records, model errors, schema validation failures, and human-review outcomes. Alert on changes in volume or missing-field rates, since an upstream format change can otherwise look like a real labor-market trend.
8. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| 401 or 403 from a source | Missing, expired, or unapproved credentials; requested use outside the approved scope. | Check the integration’s credential setup and access approval. Do not rotate through accounts or evade access controls. |
| 429 or throttling | Documented rate or quota limit reached. | Follow the provider’s retry headers or guidance, reduce concurrency, and resume from a saved cursor. |
| Records disappear between runs | Pagination state was not persisted, or source lifecycle changes were treated as permanent deletion. | Store page cursors and timestamps. Model unavailable, expired, and deleted states separately. |
| Salary or location looks plausible but is wrong | The model inferred missing information or normalization was too aggressive. | Retain the original text, require evidence, use null for unknown values, and validate against the source. |
| JSON parsing or schema validation fails | Response format, SDK version, or schema constraints do not match the API’s supported structured-output subset. | Check the current Structured Outputs documentation, simplify unsupported schema constructs, and handle refusal or incomplete responses explicitly. |
| Duplicate openings were merged | Similarity logic ignored requisition IDs, location, or employment details. | Raise the merge threshold, preserve source records, and retain the match features and rule for audit. |
| Analysis shifts sharply after a deployment | Source coverage, extraction prompt, schema, or deduplication logic changed. | Version these components and annotate reports with the change date; compare missing rates and a reviewed sample. |
Or skip the browser setup
If your authorized workflow needs a visual record of an individual public page, [ScreenshotNeo](https://screenshotneo.com) can return a screenshot or PDF with one GET request. It is a capture tool, not a bulk job-listings feed or a substitute for source permission. See the [ScreenshotNeo API documentation](https://screenshotneo.com/docs/) for parameters and formats.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Every feature is available on every plan.
Sign up for ScreenshotNeo and get 1,000 free screenshots a month with no card.
FAQ
Can I analyze job postings from LinkedIn or Indeed?
Only through an access path and use that the relevant terms and approvals allow. The cited LinkedIn and Indeed APIs describe posting workflows for approved integrations, not general feeds of all search listings.
Does structured output make the extracted data true?
No. It constrains response shape. Validate each important value against the source text and preserve evidence and provenance.
Can I publish aggregate trends from collected postings?
That depends on the source terms, license, jurisdiction, and whether the aggregation or underlying records may be published. Confirm those rights before release.
How often should I refresh records?
Use the cadence allowed by the source and needed for your analysis. Record timestamps and report the observation window; do not imply a universal freshness guarantee.


