How to Build an Ad Monitoring Tool
Build a monitor for public ad archives or authorized accounts with source-aware collection, dated snapshots, change alerts, and honest coverage reporting.

To build an ad monitoring tool, choose a documented source that fits your use, collect records on a schedule, preserve each source response with its query and observation time, normalize records for search, and compare successive observations to flag meaningful changes. This guide starts with public ad archive monitoring for competitor research, focused on Meta and Google archive sources. For reporting on your own campaigns, use the platform’s authorized account API and follow its separate access rules. Specify target countries and ad categories before building: source coverage and available fields depend on them.
An archive monitor reports what its sources returned. It cannot prove that an ad was never shown when it is absent from an archive. Show collection success, geography, query, source, and freshness wherever users inspect results.
1. Pick sources that fit the monitoring job
Start by comparing sources on access policy, geography and category coverage, fields, historical window, update cadence, and connector maintenance. Treat a source without a verified automation interface as a manual or user-assisted workflow; do not evade access controls with brittle scraping.
| Source | Useful for | Important limits |
|---|---|---|
| Meta Ad Library API | Research queries for supported ad categories and geographies | Requires Facebook account and developer setup. Political or issue-ad access requires identity and location confirmation. Fields vary by category and region. |
| Meta Ad Library Report / public library | People who need to search or inspect ads manually | Do not assume manual search is an automation endpoint. Meta points general searches for currently running ads to the Ad Library. |
| Google Ads Transparency Center | Search ads associated with verified advertisers; inspect region, last run date, and format where exposed | Coverage is tied to verified advertisers and the Center’s current scope. Do not treat it as a complete archive of every ad. |
| Google Ads API | Authorized campaign reporting and monitoring that benefits the account user | It is not a general public competitor-ad archive API. Developer-service policies prohibit specified scraping and proxy access. Check current access setup and policy before implementation. |
| Independent observation | Research questions where an archive does not expose the needed observations | Requires an authorized, compliant collection method and may not scale. Label it separately from archive data. |
Meta’s Ad Library API documentation describes developer registration, policy agreement, app creation, and Graph API queries. Common fields include Library ID, creative content, associated Page name and ID, delivery dates, and placement. Political or social-issue ads add spend and impression ranges and demographic reach. UK and EU ads have estimated impression and targeting or reach details; advertiser and payer information is identified for EU ads. These are source-specific fields, not universal columns. Meta says demographic estimates use multiple factors, including age and gender information users provide in profiles. See the Meta Ad Library API documentation.
Google’s Ads Transparency Center is distinct from Google Ads API. Google’s 2023 announcement described search by verified advertiser, region, last date run, and format. Advertiser verification information may include the advertiser’s name or organization and location, with creatives and served dates or locations publicly available. The Google announcement is a dated description, not a current completeness guarantee. For own-account reporting, consult Google Ads API policies and the developer token guide; Google’s access workflow changes, so recheck it during implementation.
2. Define scope and a source-aware data model
Write down platforms, advertiser or Page identities, keywords, country or region, ad categories, polling interval, and retention period. Keep political and social-issue collection configuration separate because access requirements may differ. For every field, record whether it is source-provided, estimated, a range, or unavailable. Missing is not zero.
Keep both the source-native record and a normalized view. A practical relational schema could begin with these tables:
CREATE TABLE watch_queries (
id INTEGER PRIMARY KEY,
source TEXT NOT NULL,
query_json TEXT NOT NULL,
enabled INTEGER NOT NULL DEFAULT 1,
interval_seconds INTEGER NOT NULL,
geography TEXT,
category TEXT
);
CREATE TABLE observations (
id INTEGER PRIMARY KEY,
watch_id INTEGER NOT NULL REFERENCES watch_queries(id),
source TEXT NOT NULL,
native_ad_id TEXT,
advertiser_name TEXT,
advertiser_id TEXT,
geography TEXT,
creative_text TEXT,
media_reference TEXT,
delivery_start TEXT,
delivery_end TEXT,
status TEXT,
placement TEXT,
observed_at TEXT NOT NULL,
source_updated_at TEXT,
raw_json TEXT NOT NULL,
parser_version TEXT NOT NULL,
response_complete INTEGER NOT NULL,
UNIQUE(source, native_ad_id, observed_at)
);
This stores a source-native ID, query association, observation time, raw response, and parser version. Add optional fields such as spend range, impression range, reach estimate, payer, or format only with explicit provenance and scope. When a provider does not supply a field, store null and an availability reason, such as not_returned_for_region. Do not put personal data into the database unless it is needed and authorized.
3. Implement a collector adapter
Give each documented source its own adapter. The adapter handles authorization, pagination, source-specific response parsing, and rate limits, then yields source records plus completeness metadata. This runnable Python skeleton shows the pipeline boundary without inventing provider endpoints or credentials. Connect fetch_documented_page only to the selected source’s documented API and parameters after completing its access setup.

from dataclasses import dataclass
from datetime import datetime, timezone
from typing import Any, Iterable
import json
@dataclass
class PageResult:
records: list[dict[str, Any]]
next_cursor: str | None
complete: bool
raw_response: dict[str, Any]
# Implement this function using the source's official documentation.
def fetch_documented_page(query: dict[str, Any], cursor: str | None) -> PageResult:
raise NotImplementedError("Add the authorized source adapter")
def collect(query: dict[str, Any], save_observation) -> None:
cursor = None
seen_cursors: set[str] = set()
while True:
page = fetch_documented_page(query, cursor)
observed_at = datetime.now(timezone.utc).isoformat()
for record in page.records:
save_observation({
"source": query["source"],
"query": query,
"observed_at": observed_at,
"raw_record": record,
"raw_page": page.raw_response,
"parser_version": "1",
"response_complete": page.complete,
})
if not page.next_cursor:
break
if page.next_cursor in seen_cursors:
raise RuntimeError("Pagination cursor repeated; stop to avoid a loop")
seen_cursors.add(page.next_cursor)
cursor = page.next_cursor
# Schedule collect(query, save_observation) with your job runner.
# Keep a separate run record for started_at, finished_at, outcome and error.
The adapter should retry transient network errors with bounded exponential backoff and jitter, but should not retry rejected authorization as if it were a transient outage. Save a run record even when a query returns no ads. This lets the interface distinguish “successful empty result” from “connector failed.” Use idempotent writes so re-running a failed window does not duplicate observations.
4. Normalize, deduplicate, and retain history
Use the source plus native ID as the preferred identity. If no stable ID exists, do not silently merge records on creative text alone; identical creative may be served by different advertisers or in different regions. Preserve the raw record, query, region, source name, retrieval time, parser version, and permitted source reference so a user can audit a change.
Keep source dates separate from your own clock. For example, delivery_start is a provider-reported date if exposed, while observed_at means your collector saw the record then. Do not label the first observation as the ad’s true first-seen date outside your system.
For small, low-frequency projects, CSV can be enough if you need simple inspection and export. The Carter Center’s 2021 political advertising monitoring toolkit recommends CSV for small collections and SQL or NoSQL at larger volumes. A relational database works well when users need joins, filters, and auditable history; a document store may fit source responses with varying shapes. Decide based on expected volume, query patterns, concurrent users, history, and export needs rather than starting with infrastructure you cannot yet justify. See the Carter Center toolkit.
5. Detect changes and send useful alerts
Compare each successful observation with the latest prior observation for the same source-native ID and relevant query scope. Alert on events users can act on: a newly observed record, a record no longer returned, a creative text or media reference change, or a delivery status/date change. “No longer returned” is not proof an ad stopped running; phrase it as an observation about the source.

def classify(previous: dict | None, current: dict) -> list[str]:
if previous is None:
return ["newly_observed"]
changes = []
for field in ("creative_text", "media_reference", "status",
"delivery_start", "delivery_end", "placement"):
if previous.get(field) != current.get(field):
changes.append(field + "_changed")
return changes
# For disappearance, compare only after a complete successful query run.
# A failed or partial page must never mark every missing record inactive.
def mark_not_returned(previous_ids: set[str], current_ids: set[str], complete: bool):
if not complete:
return []
return sorted(previous_ids - current_ids)
Deduplicate notifications by event key, such as source, ad ID, change type, and new observation timestamp. Let users choose thresholds and repeat policy. Include the source link or identifier, query, geography, observation time, and a short before/after summary. Google Cloud Monitoring’s alert policies provide one documented model for conditions, notification channels, and repeat notifications; use an equivalent feature in your chosen monitoring system if you are not on Google Cloud. See Google Cloud alerting policy documentation.
6. Schedule collection and operate it reliably
- Choose a cadence the source allows. Polling more often increases API and storage load but may not improve freshness if the archive itself updates slowly. Follow documented limits and back off on throttling.
- Track each run. Record start and end time, query, pages fetched, result count, completion, retry count, and error class.
- Expose freshness and coverage. Show last successful sync, region and category scope, pagination completeness, and stale feeds. Do not treat a partial result as a complete search.
- Rerun failed windows safely. Use idempotent observation keys and bounded retries. Preserve failed run metadata so operators can distinguish gaps from empty results.
- Watch schemas and credentials. Alert on rejected authorization, changed response fields, unexpected empty responses, and pagination failures. Keep credentials in a secret store, rotate them under the provider’s rules, and avoid logging tokens.
- Set retention intentionally. Retain only the raw response and creative references permitted for your use, and document deletion and export behavior.
Performance is mostly shaped by query count, pagination, polling cadence, and media handling. Fetch only fields needed for the use case, cap concurrency per provider, and separate lightweight metadata polling from expensive media retrieval where permitted. Index source/native ID, advertiser, geography, observed time, and watch ID. Store large media only when authorized and needed; otherwise keep a reference and a content hash for change detection. Cost includes connector maintenance, scheduled compute, database storage, alert delivery, and operator time. Estimate using watches × runs per day × pages per run × retained days, then revisit actual volume before scaling.
7. Show creatives without confusing screenshots with archive data
When a user needs a visual reference for a public page or an authorized page they can access, a screenshot can supplement an ad record. It does not establish that the ad was served, who saw it, or how much it spent. Store the capture URL, capture time, viewport, and relation to the source record separately. Respect access permissions and the source’s terms.
For a do-it-yourself visual capture, use a browser automation tool against pages you are authorized to access, and retain only the necessary image or PDF. If the source provides a library record or creative asset directly, preserve that source reference as the evidence instead of implying a screenshot is platform data.
Or skip the browser setup
For a permitted public page or authorized account page, ScreenshotNeo can return an image with one request. Its clean-shot flow accepts cookie or consent banners and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, with page verdict and billed status in response headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. ScreenshotNeo is a website screenshot API and MCP server by Yorker Media; see ScreenshotNeo and its API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
One thousand screenshots a month are free with no card; paid plans start at $5 for 3,000. Sign up free for ScreenshotNeo.
8. Troubleshoot common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Authorization rejected | Missing account confirmation, expired credential, app setup incomplete, or access level mismatch | Check the selected source’s official setup and policy pages; record the failure and do not classify it as an empty result. |
| Political or issue records unavailable | Identity/location confirmation or category-specific access is required | Complete the documented process and keep that scope distinct from general ads. |
| Ads missing from results | Geography/category mismatch, unsupported archive coverage, pagination bug, or source update delay | Display query scope and completeness; inspect pagination and compare using the source’s own interface where appropriate. Absence is not proof of non-delivery. |
| Repeated records | Pagination overlap, unstable identity, or retries inserted duplicate rows | Upsert by source and native ID, retain observation snapshots, and make writes idempotent. |
| Everything appears inactive | A partial or failed run was treated as a complete query | Only infer “not returned” after a successful, complete run; retain run status and cursor completeness. |
| Spend/reach shown as zero | Missing or inapplicable provider field converted to a numeric default | Store null with an availability reason; label estimates and ranges as such. |
| Google Ads API used for competitor archive collection | Account campaign API was mistaken for public transparency search | Use the Transparency Center for supported public advertiser lookup, or an authorized account API for the account user’s permitted reporting. Review API policies before release. |
| Screenshot is blank or obstructed | Consent overlay, dynamic page, bot check, unsupported page state, or failed load | Verify the page is accessible and authorized, then configure waits or capture settings. Treat the screenshot as a visual supplement, not an archive record. |
9. Validate coverage without claiming completeness
Coverage should be a visible property of every report: source, geography, category, query, collection window, last successful sync, and whether pagination completed. Separate a provider’s delivery dates from your observation history. Add a “data unavailable” reason when a field is not offered. These practices make the tool useful even when sources differ or change.
Independent monitoring research shows why this caution matters. Silva and colleagues’ 2020 Facebook Ads Monitor study used a browser plugin with more than 2,000 volunteers in Brazil and evaluated a classifier against 10,000 manually labelled ads. The authors reported some detected political ads absent from Facebook’s Ad Library in that study context. Those figures describe that research, not current platform-wide completeness or a general performance benchmark. See the Facebook Ads Monitor paper.
FAQ
Can an ad monitoring tool prove an ad never ran?
No. A missing record may reflect archive scope, query limits, geography, update timing, or collection failure. Report what the source returned and when.
Should I start with a database?
Not necessarily. A small collection with simple export needs can begin in CSV. Move to SQL or NoSQL when query, history, volume, or concurrency needs justify it.
Can I use Google Ads API to track competitors?
Do not assume so. It is an authorized campaign API with its own policies, not a general public competitor archive. Check Google’s current policy and use an appropriate transparency source for public lookup.
What makes a change alert trustworthy?
It identifies the source and scope, cites the observed time, compares stable identities, and only reports disappearance after a complete successful collection run.