ScreenshotNeo

BlogHow-to

How to Extract Teaching Jobs From School Websites Automatically

Build a reliable teaching-job index by checking official feeds first, collecting public careers pages where needed, and tracking each listing’s source and freshness.

By the ScreenshotNeo team4 October 20269 min read

To extract teaching jobs from school websites automatically, first look for an official RSS feed, JSON feed, API, or statewide educator job board. Use that structured source when it covers the roles and locations you need. If a school or district offers only a public careers page, build and validate a site-specific collector that revisits the page on a schedule. Store the original listing URL and when you last saw each posting; a listing disappearing from a page does not confirm that the job was filled.

This guide uses United States examples because the available sources are strongest there. Statewide boards, available feeds, and their coverage vary by location and can change.

1. Inventory official sources first

Start with the district’s careers page, any recruiting vendor linked from it, and your state education agency’s educator jobs page. Check for RSS, JSON, an API, or a downloadable feed before parsing page HTML.

  • The Illinois Education Job Bank describes RSS posting feeds and says a subscriber ID is required.
  • SchoolFront documents RSS and JSON feeds for internal and external postings. Its documentation says an administrator enables feeds in recruiting configuration and shares the resulting URL with a webmaster or technical team.
  • USAJOBS developer documentation covers search, RSS, exports, and REST services. USAJOBS is a general federal job system, not a district-specific source.

Also check whether a state board already aggregates district vacancies. For example, the Indiana Department of Education describes a statewide board that aggregates district postings and says it upgraded the board in January 2026. The Missouri Department of Elementary and Secondary Education says its statewide educator job board launched in January 2025. Oklahoma also describes statewide educator job-board services through its education agency. These examples do not establish coverage for every state, role, or vacancy.

Before depending on a source, record its geographic and role coverage, fields, pagination, update cadence, access requirements, and whether it includes temporary, substitute, or non-teaching jobs. A feed usually means less page-specific parsing code because it provides an explicit integration route; the cited sources do not quantify maintenance savings.

2. Choose a collection method

Method Use it when Check
Official feed or API The district or recruiting platform offers one and its coverage meets your needs. Fields, pagination, authentication, cadence, and role coverage.
State educator job board It covers the target geography and job categories. Which districts and vacancies are included, and how to reach the original listing.
Site-specific page collector No suitable structured source is offered and you have checked the site’s published terms and access instructions. Markup changes, JavaScript rendering, duplicate records, access controls, and freshness.

Do not assume one scraper will work reliably on every district site. Recruiting platforms and page structures differ, and direct collection needs per-site validation. A research study illustrates that repeated collection is feasible in a defined setting: it collected vacancies from 242 of Washington’s 295 district websites, scanning twice weekly, usually Mondays and Fridays, from early December 2021 through December 2022. That result describes the study’s sites and method, not a universal coverage or accuracy guarantee.

3. Build a site-specific collector

For a public careers page without a feed, the basic pipeline is: request or render the page, identify individual posting links, extract fields, normalize them, compare them with prior observations, and save the observation time. The selectors and parsing rules below are necessarily site-specific. Replace the example selectors after inspecting the target page and checking its access instructions.

Python example with Beautiful Soup

This runnable example fetches a page that exposes posting links in server-rendered HTML. Install dependencies with python -m pip install requests beautifulsoup4, set CAREERS_URL to the authorized public careers page, then run it. It emits candidate records as JSON; validate the selector and fields against the actual site before scheduling it.

import json
import os
from datetime import datetime, timezone
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup

CAREERS_URL = os.environ["CAREERS_URL"]
# Replace this with a selector verified on the target careers page.
POSTING_LINK_SELECTOR = "a.job-link"

response = requests.get(
    CAREERS_URL,
    headers={"User-Agent": "TeachingJobIndexer/1.0 contact: jobs@example.org"},
    timeout=(10, 30),
)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
observed_at = datetime.now(timezone.utc).isoformat()

records = []
for link in soup.select(POSTING_LINK_SELECTOR):
    href = link.get("href")
    title = " ".join(link.get_text(" ", strip=True).split())
    if not href or not title:
        continue
    records.append({
        "title": title,
        "employer": "Example School District",  # Set from verified page context.
        "location": None,  # Extract only if the page provides it reliably.
        "source_url": urljoin(response.url, href),
        "first_seen_at": observed_at,  # Preserve the original value on later scans.
        "last_seen_at": observed_at,
        "status_observation": "listed_on_employer_site",
    })

print(json.dumps(records, ensure_ascii=False, indent=2))

The placeholder selector, employer, and optional location must be adapted to the target. For a JavaScript-rendered page, a plain HTTP request may return an empty shell rather than the rendered listings. Use an authorized feed or API where available; otherwise choose a browser-rendering approach appropriate to that site and test it against real page behavior. There is no universal parser supplied by the sources for every district platform.

Schedule and persist observations

  1. Run each collector at a measured interval appropriate to the source. The research example used twice-weekly scans, but it does not prescribe a crawl rate for other sites.
  2. Upsert each posting using a stable key, such as a vendor posting ID when present. If there is no ID, derive a key from the employer and canonical posting URL; keep the URL itself for verification.
  3. Set first_seen_at only on initial observation and refresh last_seen_at each time the listing is seen.
  4. When a posting is absent, mark it not_observed_since or possibly_removed. Do not label it filled unless an authoritative source confirms that status.
  5. Retain enough history to diagnose page changes and explain why a record’s observed status changed.

4. Normalize and deduplicate records

Different systems use different title and location formats. Normalize fields before combining sources, but retain the original values so a reader can trace each interpretation.

Field What to store
Employer District or school name, plus the original source label.
School and location School when named; city, state, or other location only when supplied reliably.
Title and category Original title and a normalized category such as teacher, paraeducator, administrator, or principal. Keep subject-area specialties where useful.
Source Original vacancy URL, source type, and any stable vendor posting ID.
Dates Source-posted date if available, first-seen time, last-seen time, and last successful scan time.
Observation status For example, listed on employer site, possibly removed, or confirmed closed—with the evidence source for confirmation.

For research that compares vacancies with licensure or employment data, align role and subject categories with the relevant state’s classifications where practical. Do not discard the source title: categorization is an interpretation, and retaining both makes corrections possible.

Deduplicate first by an authoritative posting ID when available. Otherwise compare canonical URLs and normalized employer/title/location combinations, while avoiding merges based only on similar titles: one district may have several openings for the same role.

5. Handle freshness and uncertainty honestly

A scan tells you what the source showed at a particular time; it does not reveal the employer’s full hiring state. A job can remain online after it is filled, be removed before hiring is complete, cover multiple hires, or fail to reflect internal transfers. The Washington study discusses these limitations and cautions that vacancy postings can be an imperfect proxy for hiring needs.

  • Display “Last observed on [date]” and link to the employer’s original listing.
  • Use “listed on employer site” for a posting currently observed. Use “not observed since [date]” after it disappears.
  • Only say “filled” or “closed” when the employer or its recruiting platform provides an authoritative confirmation.
  • Keep failed scans separate from successful scans. A timeout or blocked request must not make every previously seen listing appear removed.

6. Validate access and operations

The reviewed sources do not establish a single permission rule for every school website. Check the site’s terms and published access instructions, use an offered feed or API when available, and do not treat public visibility alone as permission for any collection pattern. A legal conclusion for a specific site or jurisdiction requires separate review.

Keep collection bounded and observable: set request timeouts, log the source and scan result, record parser failures, and alert when a source suddenly returns no listings or a different page shape. Store only fields needed for the index, protect any feed credentials, and avoid collecting applicant information from forms or private areas.

7. Troubleshooting

Symptom Likely cause Fix
No postings found The selector is a placeholder or the page renders listings with JavaScript. Inspect the page structure, update the site-specific selector, and determine whether an official feed/API exists. If rendering is required, use a suitable authorized browser-rendering method.
HTTP 403 or access denied The host rejected the request or access is restricted. Check published access instructions and use the official feed or contact the site owner for an approved route. Do not try to bypass access controls.
HTTP 429 or repeated timeouts The source is rate-limiting requests, under load, or unreachable. Reduce scan frequency, add bounded retries with backoff, and treat the run as incomplete rather than marking postings missing.
Duplicate jobs Tracking parameters, alternate URL forms, or multiple source cards refer to one posting. Canonicalize URLs carefully, prefer vendor IDs, and review possible matches before merging distinct openings.
Old jobs appear open The listing remains online after its hiring status changed. Display observation dates and the source link; only mark closed when the source confirms it.
Many jobs suddenly disappear The source changed layout, pagination, access behavior, or the collector failed. Check scan health and sample the page manually. Do not convert a failed or anomalous scan into removals.

8. Performance, reliability, and cost

Structured feeds reduce parsing work when their coverage and fields fit your use case, but validate update cadence and completeness before relying on them. For page collection, prioritize scan health, stable keys, timeouts, and retained observations over aggressive polling. The reviewed material gives no universal crawl rate, performance benchmark, or cost estimate. Your cost depends on the number of sources, whether browser rendering is needed, how often each is checked, and how much history you retain.

Keep collection modular by source: one adapter for each feed or site, a shared normalization layer, and a common storage schema. This confines layout changes to the affected adapter. Track successful scans separately from listing counts so a site returning zero jobs can be distinguished from a broken collector.

Or skip the browser setup

If a public careers page needs browser rendering to inspect, ScreenshotNeo can return a screenshot or PDF with one GET request. It is a screenshot API and MCP server; a screenshot can help inspect a page, but it does not replace a structured feed or extract job records into your database by itself.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Sign up for 1,000 free screenshots a month, with no card.

FAQ

Is there one API for all school district vacancies?

The sources here do not establish one universal API. Check each state board, district, and recruiting platform for its own feed or integration.

Can a missing posting be treated as filled?

No. Absence from a scan means it was not observed there at that time. Seek confirmation from the employer before reporting that it is filled.

Should I collect job descriptions too?

Only if your use case needs them and the source permits the collection. Keep the employer’s original listing link so readers can verify details and apply there.