Google News API for Brand and Media Monitoring
Google News has no documented brand-monitoring API. Use its RSS search feed carefully, with deduplication, limits, and reliable fallback workflows.
Direct answer: Google News does not offer a documented public API for brand monitoring. You can query its RSS search feed for a low-cost stream of candidate stories, but treat that feed as an undocumented discovery aid rather than a complete clipping service or an API with guaranteed limits, freshness, pagination, or coverage.
The observed search-feed pattern is:
https://news.google.com/rss/search?q=<url-encoded-query>&hl=en-US&gl=US&ceid=US:en
Build monitoring around explicit searches, XML parsing, link normalization, deduplication, retrieval timestamps, and human review of ambiguous matches. The guide commonly used for this pattern reports roughly 100 items per feed and no pagination; that is a third-party observation, not an official Google quota.
1. What “Google News API” means in practice
Searches for a Google News API often refer to the RSS search feed. Google has not published a reference specification for that endpoint, so its URL shape and response behavior can change. Google’s official Publisher Center documentation describes feeds for publisher content sections and explains that Google News discovers content through several methods, including web crawling and publisher-submitted content. That documentation does not define an external brand-monitoring API or promise that an RSS search captures every relevant article.
Use the feed when you need:
- a small, inexpensive stream of recent candidate articles;
- simple polling from a cron job or serverless function;
- brand, product, executive, competitor, or topic discovery;
- a starting point for a review queue.
Do not describe it as a complete archive, a guaranteed alerting system, or a supported Google API contract.
2. Construct a Google News search feed
Basic query
https://news.google.com/rss/search?q=Acme&hl=en-US&gl=US&ceid=US:en
Always URL-encode the query. Include known spelling variants and disambiguating terms:
"Acme Corp" OR "Acme Cloud" OR @acme -jobs
"Acme" (security OR breach OR outage)
("Jane Doe" OR "J. Doe") Acme
| Parameter | Purpose | Example |
|---|---|---|
q |
Search expression | %22Acme%20Corp%22%20OR%20%22Acme%20Cloud%22 |
hl |
Language and interface locale in the observed pattern | en-US |
gl |
Country edition in the observed pattern | US |
ceid |
Google News edition identifier in the observed pattern | US:en |
Test each exact query and edition in a feed reader or HTTP client before making it an operational dependency. Keep separate feeds for materially different concepts instead of one very broad query.
3. Fetch the feed with cURL
This saves the XML response for inspection:
curl --fail --show-error --location \
-G 'https://news.google.com/rss/search' \
--data-urlencode 'q="Acme Corp" OR "Acme Cloud"' \
--data-urlencode 'hl=en-US' \
--data-urlencode 'gl=US' \
--data-urlencode 'ceid=US:en' \
-o acme-news.xml
For production polling, check the HTTP status, content type, response size, and whether the XML contains the expected channel and item elements.
4. Parse results in Python
The following example uses only Python’s standard library. It fetches a feed, extracts common RSS fields, records retrieval time, and deduplicates links.
#!/usr/bin/env python3
import html
import time
import urllib.parse
import urllib.request
import xml.etree.ElementTree as ET
QUERY = '"Acme Corp" OR "Acme Cloud"'
params = {
"q": QUERY,
"hl": "en-US",
"gl": "US",
"ceid": "US:en",
}
url = "https://news.google.com/rss/search?" + urllib.parse.urlencode(params)
request = urllib.request.Request(
url,
headers={"User-Agent": "brand-monitor/1.0 (+https://example.com/contact)"},
)
with urllib.request.urlopen(request, timeout=30) as response:
xml_bytes = response.read()
root = ET.fromstring(xml_bytes)
seen = set()
retrieved_at = int(time.time())
for item in root.findall("./channel/item"):
title = (item.findtext("title") or "").strip()
link = (item.findtext("link") or "").strip()
published = (item.findtext("pubDate") or "").strip()
source = item.find("source")
source_name = (source.text or "").strip() if source is not None else ""
if not link or link in seen:
continue
seen.add(link)
print({
"title": html.unescape(title),
"link": link,
"source": source_name,
"published": published,
"retrieved_at": retrieved_at,
})
For a long-running service, store these records in a database rather than printing them. Preserve the original link and the retrieval timestamp so later changes in feed behavior are traceable.
5. Parse results in Node.js
Node.js has fetch in current releases but no built-in XML parser. Install a small parser first:
npm install fast-xml-parser
import { XMLParser } from "fast-xml-parser";
const query = '"Acme Corp" OR "Acme Cloud"';
const params = new URLSearchParams({
q: query,
hl: "en-US",
gl: "US",
ceid: "US:en"
});
const response = await fetch(`https://news.google.com/rss/search?${params}`, {
headers: { "user-agent": "brand-monitor/1.0" }
});
if (!response.ok) {
throw new Error(`Google News returned ${response.status}`);
}
const xml = await response.text();
const parser = new XMLParser({ ignoreAttributes: false });
const parsed = parser.parse(xml);
const rawItems = parsed?.rss?.channel?.item ?? [];
const items = (Array.isArray(rawItems) ? rawItems : [rawItems])
.map(item => ({
title: item.title ?? "",
link: item.link ?? "",
source: typeof item.source === "object" ? item.source["#text"] : (item.source ?? ""),
published: item.pubDate ?? "",
retrieved_at: new Date().toISOString()
}))
.filter(item => item.link);
const unique = [...new Map(items.map(item => [item.link, item])).values()];
console.log(JSON.stringify(unique, null, 2));
6. Build a dependable monitoring workflow
- Define each concept. Create separate queries for the brand, product names, executives, competitors, and high-value risk terms.
- Poll on a schedule. Record the exact query, edition, request time, HTTP status, and response size.
- Parse defensively. Treat missing fields, a single item, malformed XML, and namespace changes as normal failure cases.
- Normalize links. Remove harmless tracking parameters where your policy allows, normalize casing only where safe, and retain the original URL.
- Deduplicate. Use the normalized URL plus title and source as fallback keys. The same story can appear through several feeds or syndicated outlets.
- Classify matches. Route ambiguous names and unrelated uses to human review. A brand string alone is not proof of relevance.
- Persist evidence. Store title, link, source label when present, publication date, retrieval time, query, edition, and processing status.
- Alert after filtering. Send notifications only after deduplication and relevance checks to prevent noisy alerts.
Because the observed feed may contain roughly 100 items with no pagination, split busy monitoring into narrower queries or time windows and deduplicate the combined results. This can improve collection depth but cannot prove that every article was found.
7. Coverage, freshness, and completeness limits
- No completeness guarantee: Google News uses multiple discovery methods, so a search feed is not the full Google News corpus.
- Undocumented refresh behavior: Do not assume a fixed polling interval, publication order, or alert latency.
- Limited retrieval depth: The roughly 100-item cap and lack of pagination reported by the technical guide make the feed unsuitable as a historical archive.
- Edition effects: Language and country parameters can change which results appear. Monitor the editions relevant to your audience.
- Query ambiguity: Common brand names produce unrelated results. Add product, geography, industry, or executive terms.
- Syndication: One event may produce many near-duplicate URLs. Deduplicate before counting mentions.
8. Google Alerts and structured alternatives
Google Alerts is an adjacent Google tool. Google’s Data Portability schema lists NEWS as a possible source and RSS as a possible delivery method for Alerts subscriptions. That supports RSS delivery, but does not establish monitoring completeness, a fixed delivery frequency, or historical retention.
GDELT Cloud is a separate structured news-data service with APIs for stories, events, entities, and tone or share-of-voice analysis. Evaluate it against your requirements; its coverage and scoring are provider-defined and it should not be described as equivalent to Google News results.
| Decision factor | Questions to answer |
|---|---|
| Coverage | Which outlets, countries, languages, and source types are included? |
| Freshness | How quickly do new stories become available? |
| Querying | Are Boolean operators, phrase searches, exclusions, and language filters supported? |
| History | Is there pagination, retention, export, or replay? |
| Metadata | Are canonical URLs, authors, source labels, dates, and duplicate identifiers available? |
| Analytics | Do you need sentiment, entities, events, or share-of-voice metrics? |
| Operations | What authentication, quotas, pricing, rate limits, and support exist? |
9. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| HTTP 400 or an empty feed | Malformed query or missing URL encoding | Build the URL with URLSearchParams, urlencode, or --data-urlencode; test a simple quoted term first. |
| Unexpected language or geography | Incorrect hl, gl, or ceid |
Use a consistent locale and edition, then compare results across the editions you actually need. |
| Only recent items appear | Feed is a current-results stream with limited depth | Poll regularly, persist every successful response, and use narrower queries or time windows. Do not assume backfill exists. |
| Duplicate stories | Syndication or multiple monitored queries | Normalize URLs and combine URL, title, source, and publication time for deduplication. |
| Relevant brand mentions are missing | Query spelling, edition, ranking, or Google News discovery limits | Add variants and disambiguators, monitor multiple editions, and use a second data source for important coverage. |
| XML parser failure | Malformed response, unexpected HTML, or schema change | Check status and content type, log a bounded response sample, retry with backoff, and alert when the feed shape changes. |
| Too many irrelevant matches | Ambiguous brand name | Add product, sector, location, or exclusion terms and send uncertain matches to review. |
| Polling overload | Too many feeds or overly frequent requests | Use a scheduler, exponential backoff, bounded concurrency, caching, and a sensible polling interval. |
10. Performance, reliability, and cost
- Performance: Keep requests small by splitting concepts deliberately. Parse and deduplicate incrementally, and use bounded concurrency rather than launching every feed at once.
- Reliability: Set connection and read timeouts, retry transient failures with exponential backoff, record status and latency, and preserve the last successful result. A missing feed response should not erase prior mentions.
- Cost: RSS retrieval itself is lightweight, but production cost comes from polling infrastructure, storage, notification delivery, human review, and any second data provider.
- Observability: Track per-query item counts, unique-link counts, empty responses, parse failures, HTTP errors, and time since the last successful poll.
- Data handling: Store only the metadata needed for monitoring and link back to the publisher article. Respect applicable publisher terms and privacy requirements.
11. Or skip the browser setup
RSS gives you article links. If your workflow also needs clean screenshots of those pages for a review queue, evidence record, or report, ScreenshotNeo can capture a URL with one request. It is a screenshot API and MCP server, not a Google News feed.
Use the ScreenshotNeo API documentation for the complete option list. Basic call:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Before the capture, ScreenshotNeo accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
12. FAQ
Is there an official Google News API for brand monitoring?
There is no documented public API reference for the Google News search RSS endpoint. Treat the feed as an observed, undocumented access pattern.
Can I retrieve all historical mentions?
No. The reported feed behavior is roughly 100 items with no pagination, and Google does not promise that the feed represents every News result. Poll continuously and store results if you need your own history.
Does Google Alerts replace the RSS search feed?
It can be an adjacent alerting option. Its documented subscription schema supports NEWS and RSS delivery, but it does not establish a completeness or retention guarantee for brand monitoring.
Should I use GDELT Cloud instead?
Use requirements to decide. GDELT Cloud is designed for structured news analysis and offers story, event, entity, and tone or share-of-voice capabilities; its coverage and metrics differ from Google News.
How often should I poll?
Choose an interval based on acceptable alert delay, feed count, and operational limits. Start conservatively, measure freshness and failures, and use backoff during errors.


