ScreenshotNeo

BlogHow-to

How to Track Government Contracts With Web Scraping

Track federal opportunities and awards with permitted APIs, scheduled refreshes, normalization, deduplication, alerts, and audit trails.

By the ScreenshotNeo team29 September 202610 min read

How to Track Government Contracts With Web Scraping

Direct answer: Build a government-contract tracker on official data interfaces instead of scraping HTML pages. Use SAM.gov Contract Opportunities for current federal notices, SAM.gov Contract Awards for award records, and USAspending.gov for spending history. Schedule permitted API or data-file retrievals, normalize the records, detect changes, and alert on new notices, amendments, deadlines, and awards. SAM.gov’s Terms of Use prohibit automated data gathering and web-scraping tools, so do not mine pages or use Login.gov credentials with a bot.

This guide shows an implementation pattern that is useful for a small internal tracker and can be expanded into a production pipeline. It covers opportunity discovery, award monitoring, data modeling, refresh schedules, deduplication, compliance, reliability, cost, troubleshooting, and a browser-free way to capture source pages for review.

1. Choose the right official source

Government-contract tracking involves several datasets. Keep them separate because they answer different questions.

Question Primary source Useful fields
What can I bid on? SAM.gov Contract Opportunities Notice identifier, title, agency, posted date, response deadline, set-aside and eligibility fields, source URL
Who won an award? SAM.gov Contract Awards and the GSA Contract Awards API Awardee, agency, award type, obligations, dates, related identifiers
How much did an agency or recipient receive? USAspending.gov public API Recipients, agencies, geographic breakdowns, award and spending history

SAM.gov describes its contracting service as a centralized source for opportunities, awards, and subcontract reports. The USAspending API allows the public to access comprehensive U.S. government spending data, including recipients and agency or geographic breakdowns. Use the official documentation for the exact endpoint, authentication, filters, pagination, and download options available to your account.

2. Define the tracking questions and scope

Write the questions your tracker must answer before writing code. A narrow scope reduces noise and makes alerts actionable.

  • Opportunities: Which new notices match our keywords, agencies, NAICS codes, PSC codes, set-aside types, or geographic constraints?
  • Deadlines: Which response dates changed, and which notices are approaching their deadline?
  • Recompetes: Which expiring or recently awarded work resembles a tracked contract?
  • Awards: Which vendors won awards from a tracked agency, and what obligations or award types were recorded?
  • Audit: What did the source return, when was it retrieved, and what changed since the previous response?

Store the search definition with each run. A future reviewer should be able to see the keyword set, agency filters, date window, endpoint, and retrieval timestamp that produced an alert.

3. Use a compliant discovery architecture

A practical tracker has six layers.

A reliable tracker separates discovery, normalization, change detection, and alerts.
A reliable tracker separates discovery, normalization, change detection, and alerts.
  1. Discovery: Query the permitted SAM.gov opportunity interface or download. Capture the notice identifier, title, agency, dates, eligibility fields, and source URL.
  2. Awards: Query SAM.gov Contract Awards and USAspending separately. Preserve award identifiers and recipient information instead of joining records only by title.
  3. Refresh: Poll on a schedule based on deadline urgency. Store the source endpoint, response status, retrieval time, and a hash or version marker.
  4. Normalization: Map agency names, NAICS and PSC codes, identifiers, dates, and dollar fields into stable internal columns while retaining the original source value.
  5. Alerts: Notify on new notices, changed deadlines, amendments, sources-sought notices, and awards involving tracked agencies, vendors, or codes.
  6. Compliance: Use public API access, authorized keys, extracts, or downloads. Do not automate a logged-in SAM.gov session or mine restricted, sensitive, or FOUO data without authorization.

SAM.gov’s Terms of Use state that automated data gathering and web-scraping tools are prohibited and that detected accounts can be denied access through Login.gov. Treat that as an engineering requirement: your collector should call an allowed interface and retain evidence of what it called.

4. Design a durable internal schema

Keep opportunities and awards in different tables. A notice may receive amendments, expire without an award, or lead to multiple related records. An award may have no directly matching opportunity in your local data.

Opportunity fields

  • source, source_url, and notice_id
  • title, agency_raw, and normalized agency_id
  • posted_at, response_deadline, and any archive or close date
  • naics_codes, psc_codes, set-aside, eligibility, and place-of-performance fields when present
  • retrieved_at, source_status, content_hash, and first_seen_at

Award fields

  • award_id, related identifiers, recipient_name, and normalized recipient ID when supplied
  • agency_raw, award_type, obligations, and relevant start or end dates
  • retrieved_at, endpoint, response status, and raw-response reference

Never overwrite the original value when normalizing. Keep both agency_raw and agency_normalized; agency naming can change and the raw value is part of your audit trail.

5. Implement scheduled retrieval

The following Python pattern is intentionally endpoint-agnostic. Set CONTRACTS_ENDPOINT to the authorized endpoint documented by SAM.gov, GSA, or USAspending for your use case. The script records response metadata, computes a content hash, and writes a raw snapshot that a later normalization job can process.

import hashlib
import json
import os
from datetime import datetime, timezone
from pathlib import Path

import requests

endpoint = os.environ['CONTRACTS_ENDPOINT']
api_key = os.environ.get('CONTRACTS_API_KEY')
params = {
    'keyword': os.environ.get('CONTRACT_KEYWORD', 'cybersecurity'),
    'page': 1,
    'limit': 100,
}
headers = {'Accept': 'application/json'}
if api_key:
    headers['X-Api-Key'] = api_key

retrieved_at = datetime.now(timezone.utc).isoformat()
response = requests.get(endpoint, params=params, headers=headers, timeout=60)
response.raise_for_status()
raw = response.content
content_hash = hashlib.sha256(raw).hexdigest()

record = {
    'retrieved_at': retrieved_at,
    'endpoint': response.url,
    'status': response.status_code,
    'content_hash': content_hash,
    'payload': response.json(),
}
Path('snapshots').mkdir(exist_ok=True)
Path(f'snapshots/{content_hash}.json').write_text(
    json.dumps(record, indent=2), encoding='utf-8'
)
print(json.dumps({k: record[k] for k in record if k != 'payload'}, indent=2))

Consult the USAspending endpoint documentation and the SAM.gov or GSA documentation for the endpoint-specific parameter names. Do not assume that a filter accepted by one source works on another.

6. Add pagination, deduplication, and change detection

API responses are usually paginated. The documented GSA Contract Awards API behavior allows up to 100 records per page, while the default is 10. Read the response’s pagination metadata when available and stop when there is no next page. Persist the complete query and page number so a run can be reproduced.

Deduplicate with a source-native identifier first. If an identifier is missing, use a cautious compound key such as normalized title, agency, posted date, and deadline, then mark the match as probabilistic for review. Hash the canonicalized record to detect amendments. A changed deadline should create an event even if the notice title is unchanged.

def canonical_hash(item):
    selected = {
        'id': item.get('notice_id') or item.get('award_id'),
        'title': item.get('title'),
        'agency': item.get('agency'),
        'posted': item.get('posted_at'),
        'deadline': item.get('response_deadline'),
    }
    encoded = json.dumps(selected, sort_keys=True, separators=(',', ':')).encode()
    return hashlib.sha256(encoded).hexdigest()

for item in records:
    key = item.get('notice_id') or item.get('award_id')
    version = canonical_hash(item)
    previous = database.lookup_version(key)
    if previous != version:
        database.save_version(key, version, item)
        alerts.emit('contract_record_changed', item)

7. Schedule refreshes around deadlines

There is no single correct polling interval. Use urgency and source behavior to choose one:

Use case Starting schedule Reason
Long-range market research Daily Lower request volume and sufficient for broad trend tracking
Active bid pipeline Several times per day Faster visibility into amendments and deadline changes
Near-deadline monitoring Shorter interval approved by the interface limits Reduces update lag while respecting quotas and terms

Record both source time fields and your retrieval time. An alert should say “deadline changed in the source record” and include when your system observed it. Do not claim that a tracker is real time unless the source promises that behavior.

8. Build useful alerts

Start with alerts that require a decision:

  • New notice matching a saved keyword, agency, NAICS, or PSC filter
  • Response deadline added, removed, or changed
  • Amendment or sources-sought notice detected
  • Award posted for a tracked agency, recipient, or code
  • Source errors, authentication failures, or a refresh that exceeds its expected lag

Include the notice or award identifier, title, agency, changed fields, source URL, retrieval timestamp, and the query that matched it. Deduplicate notifications by event hash so a temporary retry does not send the same alert repeatedly.

9. When a browser is useful—and when it is not

Use a browser only for permitted human review, authenticated workflows that explicitly allow automation, or pages that your own organization controls. A browser does not make prohibited scraping compliant. It also adds failure modes: consent banners, newsletter popups, chat widgets, bot checks, JavaScript timing, and lazy-loaded content can obscure what a reviewer sees.

Clean visual captures remove common consent and widget distractions before review.
Clean visual captures remove common consent and widget distractions before review.

For source evidence, save the official response and URL first. If you need a visual snapshot of a public page for an internal review, capture it after the data retrieval and label it with the retrieval time. Do not use a screenshot as a substitute for the official record.

10. Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. One GET request can capture a clean PNG, JPEG, WebP, or PDF. Before capture, it accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response reports the result through X-Page-Verdict and X-Billed headers.

See the ScreenshotNeo API documentation for all options. This cURL example captures the public SAM.gov opportunities page:

curl -G 'https://api.screenshotneo.com/v1/shot' \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://sam.gov/opportunities \
  -o sam-opportunities.webp

Python:

import requests

r = requests.get(
    'https://api.screenshotneo.com/v1/shot',
    params={'access_key': 'YOUR_API_KEY', 'url': 'https://sam.gov/opportunities'},
    timeout=90,
)
r.raise_for_status()
open('sam-opportunities.webp', 'wb').write(r.content)

Node.js:

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://sam.gov/opportunities'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
require('node:fs').writeFileSync('sam-opportunities.webp', data);

ScreenshotNeo also supports full-page capture with lazy images loaded, CSS-element capture, device presets, custom viewports, retina scale, PDF paper and margin controls, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, Authorization, timezone, geolocation, transparent backgrounds, resizing, configurable caching, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

There is a free tier of 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account to capture review pages without setting up a browser.

11. Troubleshooting common failures

Symptom Likely cause Fix
HTTP 401 or 403 Missing, expired, or unauthorized API key Check the source’s authentication rules, rotate the key, and keep it in a secret store rather than source code.
Repeated duplicate alerts No stable identifier or event hash Use the source-native ID, store a canonical hash, and make alert delivery idempotent.
Missing amendments Refresh interval is too slow or only the original notice is stored Poll more often within documented limits and persist every version.
Inconsistent agency names Raw labels vary between records or sources Keep the raw value and map it to a maintained internal agency table.
Empty result set Filter syntax differs by API or the date window is too narrow Validate one filter at a time against the official documentation and log the final request.
Timeouts or partial pages Large result set, transient network error, or unhandled pagination Reduce page size, retry with bounded exponential backoff, and checkpoint each page.
Screenshot shows a popup The page needs a consent or widget cleanup step Use ScreenshotNeo’s cleanup controls or capture the underlying official data response instead.
Screenshot is billed unexpectedly The response was a clean successful capture Inspect X-Page-Verdict and X-Billed; cache hits and failed loads are not billed.

12. Performance, reliability, and cost

Measure coverage and update lag rather than inventing an accuracy percentage. Report which sources, agencies, date ranges, and record types are included, plus the time between a source update and your observation.

For reliability, use bounded retries, request timeouts, pagination checkpoints, raw-response retention, and a dead-letter queue for records that fail normalization. Alert when a scheduled run produces an abnormal zero-result response or when the source returns repeated errors. Keep source URLs and retrieval timestamps so an analyst can reproduce the finding.

Control cost by narrowing queries, requesting only needed fields when supported, caching unchanged pages, and separating frequent deadline checks from slower historical backfills. USAspending’s documented page-size limit of 100 records can reduce round trips when your query legitimately returns large pages. ScreenshotNeo caching lets you choose a TTL; use it for repeated visual reviews, while official API responses remain the authority for contract data.

13. Compliance checklist

  • Use public APIs, authorized API access, extracts, or downloads.
  • Review SAM.gov terms and the terms for every source you call.
  • Never use Login.gov credentials for automated data mining.
  • Store only data your organization is permitted to retain and display.
  • Keep raw responses, source URLs, timestamps, query definitions, and change history.
  • Separate opportunity records from award and spending records.
  • Document coverage, refresh cadence, known gaps, and update lag.

14. FAQ

Can I scrape SAM.gov for new solicitations?

SAM.gov’s Terms of Use prohibit automated data gathering and web-scraping tools. Use its permitted APIs, data files, downloads, or other authorized interfaces instead.

Should opportunities and awards share one table?

No. They have different identifiers, lifecycles, and meanings. Link them when a reliable relationship exists, but preserve each source record.

How often should a tracker refresh?

Match the schedule to deadline urgency and documented source limits. Daily is reasonable for market research; active bids need more frequent checks.

Is a screenshot evidence of an award?

No. A screenshot is a visual reference. Retain the official API response or download as the authoritative record and use the screenshot only to help a reviewer.

How do I prove what changed?

Store each raw response, retrieval timestamp, source URL, and canonical content hash. Compare versions and emit an event for changed deadlines, amendments, or award fields.