How to Create an Aggregator Website: Pull Many Sources into One
Build an aggregator that imports permissioned feeds and APIs, deduplicates items, preserves attribution, and publishes useful links.

A content aggregator collects structured items from multiple sources, makes them consistent, removes duplicates, and presents them with clear links back to their publishers. Build it around permissioned RSS or Atom feeds and documented APIs. Fetch sources on a schedule, store normalized records, show short excerpts and attribution, and link readers to the original pages. A WordPress block or plugin can prove the idea; a custom ingestion service gives you control over ranking, deduplication, search, alerts, and sources beyond feeds.
This guide walks through source selection, a runnable Python feed aggregator, storage and rendering decisions, WordPress options, crawler and reuse rules, operations, and common failures. If your aggregator also needs screenshots of source pages, ScreenshotNeo is one option to add that capability without running a browser yourself.
1. Decide what one aggregated item means
Before you connect sources, define the thing your site publishes. It might be a headline with a two-sentence excerpt, a job listing, an event, a product, or a research paper. Your source format and display rules depend on that decision.
Write down these policies before onboarding publishers:
- Identity: What makes two records the same item: a feed GUID, canonical URL, API ID, or another stable key?
- Attribution: Which source name, author, publication date, and original link appear on every card?
- Reuse: Are you displaying a title, a short excerpt, an image, or the full source content? Start with excerpts and links unless the publisher or API terms grant broader rights.
- Freshness: How often does each source need to be checked?
- Removal: How can a source owner report an item that should be removed or corrected?
These are product and editorial decisions as much as engineering choices. Keeping them explicit makes it possible to explain where an item came from and what permission applies to it.
2. Choose sources and record their terms
Prefer an official RSS or Atom feed, followed by a documented API. WordPress sites commonly expose feeds in formats such as RSS 2.0 and Atom; its REST API provides structured JSON for applications. See the [WordPress feed documentation](https://developer.wordpress.org/advanced-administration/wordpress/feeds/) and [REST API introduction](https://developer.wordpress.org/rest-api/).
For each source, create a registry record containing:
- Source owner and display name
- Feed or API URL, source identifier, and terms URL
- Authentication method and documented rate limits, if any
- Expected polling interval and the source’s robots.txt URL
- Contact or process for corrections and removal requests
Do not assume that finding a feed gives you permission to republish its contents. Feed availability is a technical interface; the publisher’s terms and applicable content rights govern what you may reuse. WordPress documents feed customization, including limiting syndicated information and adding machine-readable copyright statements in its [feed guidance](https://developer.wordpress.org/advanced-administration/wordpress/feeds/). API terms also apply to content retrieved through an API; see [WordPress.com API terms](https://developer.wordpress.com/docs/api/).
3. Build a normalized item model
Each source uses slightly different field names and conventions. Convert incoming data into one shape before ranking or rendering it. A useful starting record is:

{
"source_id": "example-blog",
"source_name": "Example Blog",
"canonical_url": "https://example.com/posts/one",
"title": "An example item",
"author": "A. Writer",
"published_at": "2026-09-28T12:00:00Z",
"excerpt": "A short description supplied by the source...",
"image_url": null,
"feed_guid": "post-123",
"fetched_at": "2026-09-29T09:00:00Z",
"terms_url": "https://example.com/terms"
}
Prefer the feed’s stable GUID or the canonical URL as the item identity. If neither exists, use a carefully normalized title plus source and publication time as a fallback. Keep a hash of normalized text as a second check for repeated entries. Do not deduplicate on title alone: unrelated sources can publish similarly titled items.
Preserve the original URL even if you also store a normalized version for deduplication. Redirects, tracking parameters, and URL changes can make identity tricky; retain the fetched URL and canonical destination so you can audit and repair records rather than silently losing provenance.
4. Fetch feeds incrementally with Python
The example below is a small command-line prototype. It fetches RSS or Atom feeds, normalizes entries, deduplicates them by GUID or link, and prints JSON records. It intentionally stores short feed summaries and source links; a production site should persist the records and render them with attribution. Install the dependency with python -m pip install feedparser, save the code as aggregate.py, and run python aggregate.py.
from datetime import datetime, timezone
import hashlib
import json
import feedparser
SOURCES = [
{"id": "python-blog", "name": "Python Blog", "url": "https://blog.python.org/feeds/posts/default"},
{"id": "wordpress-news", "name": "WordPress News", "url": "https://wordpress.org/news/feed/"},
]
def utc_timestamp(entry, field):
value = entry.get(field)
if not value:
return None
try:
return datetime(*value[:6], tzinfo=timezone.utc).isoformat()
except (TypeError, ValueError):
return None
def clean_excerpt(entry):
summary = entry.get("summary", "")
# Feed summaries may contain HTML. Keep this prototype conservative;
# production rendering must sanitize HTML with an allowlist sanitizer.
return " ".join(summary.split())[:500]
def main():
seen = set()
items = []
fetched_at = datetime.now(timezone.utc).isoformat()
for source in SOURCES:
parsed = feedparser.parse(source["url"])
if parsed.bozo and not parsed.entries:
print(json.dumps({"warning": "feed_parse_failed", "source": source["id"]}))
continue
for entry in parsed.entries:
link = entry.get("link")
if not link:
continue
guid = entry.get("id") or entry.get("guid")
identity = guid or link
digest = hashlib.sha256(identity.encode("utf-8")).hexdigest()
if digest in seen:
continue
seen.add(digest)
items.append({
"source_id": source["id"],
"source_name": source["name"],
"canonical_url": link,
"title": entry.get("title", "Untitled"),
"author": entry.get("author"),
"published_at": utc_timestamp(entry, "published_parsed"),
"excerpt": clean_excerpt(entry),
"feed_guid": guid,
"fetched_at": fetched_at,
})
print(json.dumps(items, ensure_ascii=False, indent=2))
if __name__ == "__main__":
main()
This is deliberately a fetching and normalization example rather than a complete web application. The feed URLs in the sample are examples of public feeds; for a real aggregator, replace them with sources whose terms you have reviewed. The code does not implement durable storage, conditional HTTP requests, authentication, source-specific rate limits, or HTML sanitization for a production renderer.
Production changes to make next
- Move the source list into a database or configuration store so a source can be paused without a deployment.
- Save each successful fetch time and response validator such as ETag or Last-Modified where available; send conditional requests on the next poll.
- Use a scheduled worker or queue, not a feed fetch during a reader’s page request.
- Upsert records using a unique source plus stable item identity. Update changed titles and summaries without creating a second item.
- Sanitize feed HTML before display. Escape text by default; if excerpts permit markup, use an allowlist sanitizer.
- Keep a separate audit record for the terms version or URL and fetch time that applied when the item was ingested.
5. Choose WordPress or a custom application
WordPress is a quick way to validate a feed-based publication. Its [RSS block](https://wordpress.org/documentation/article/rss-block/) accepts a feed URL and can display titles, authors, dates, and excerpts in list or grid layouts. A plugin such as [WP RSS Aggregator](https://wordpress.org/plugins/wp-rss-aggregator/) is another prototype path; check its current features and limits before committing to it.
Choose the WordPress route when a small curated set of feeds and editorial publishing are the main needs. A custom ingestion service becomes useful when you need cross-source ranking, robust deduplication, search, alerts, authenticated feeds, multiple API shapes, or explicit source-level retry and compliance controls. You can combine them: let a separate worker ingest and normalize data, then use the [WordPress REST API](https://developer.wordpress.org/rest-api/) as the publishing or editorial interface.
| Need | WordPress block or plugin | Custom ingestion service |
|---|---|---|
| Launch a simple feed directory | Fast to configure | More setup than necessary |
| Custom ranking and deduplication | Depends on plugin behavior | Full control over rules and data |
| Many APIs or authenticated sources | May need extensions or custom code | Designed for source-specific adapters |
| Editorial publishing | Built-in authoring workflow | Can integrate with WordPress through REST |
6. Publish with attribution and crawler controls
Show the source name, publication date when available, and a prominent link to the original. Keep excerpts short and useful. Do not copy full articles, images, or media merely because a feed includes them; verify permission and API terms first. Give source owners a visible contact path for corrections and removal requests.
Respect each host’s robots.txt rules when crawling. Google’s [robots.txt specification](https://developers.google.com/crawling/docs/robots-txt/robots-txt-spec) describes retrieval by HTTP GET and how rules apply by host, scheme, and port. Robots.txt provides crawler guidance; it is not a copyright license. Review terms, rate limits, and reuse permission separately. WordPress also explains robots.txt and discovery controls in its [search engine optimization documentation](https://wordpress.org/documentation/article/search-engine-optimization/).
Prefer feeds and documented APIs over scraping page HTML. If scraping is expressly allowed and necessary, identify your crawler where appropriate, limit request rates, avoid login or access controls, and stop when a publisher objects. A feed subscription is also different from a search crawler: use the source’s published access method and terms for your use case.
7. Schedule, cache, and operate the pipeline
Do not fetch every source every time a visitor opens the homepage. Use scheduled background jobs, cache successful responses, and keep the last good copy available if a source is temporarily down. Poll sources at a cadence that matches their publishing frequency and stated limits. Use conditional requests when supported, exponential backoff for transient failures, bounded retries, and a dead-letter queue for malformed or repeatedly failing feeds.
Track these measures per source:
- Last successful fetch and current staleness
- HTTP status, parse errors, fetch duration, and retry count
- Items received, changed, rejected, and deduplicated
- Reader clicks to original sources
- Corrections, source pauses, and removal requests
Add a source-level pause switch, alert when a source becomes stale, and make ingestion idempotent so retries do not create duplicates. If an upstream server times out, retain existing items and mark the source stale instead of deleting its content or showing an empty feed.
8. Handle edge cases deliberately
- Missing dates: Keep publication time unknown; record fetch time separately. Do not label fetch time as publication time.
- Changed URLs: Preserve the prior URL and stable GUID when possible. A URL change can otherwise look like a new story.
- Repeated feed entries: Deduplicate on stable identity, then use content hashes as a review signal.
- Malformed markup: Quarantine the item or strip unsafe markup; never insert untrusted feed HTML directly into a page.
- Large feeds: Limit entries processed per poll and use pagination or API cursors where documented.
- Encoding and time zones: Parse feed dates carefully, normalize stored timestamps to UTC, and preserve the original value for debugging.
- Removed items: Decide whether a missing entry means deletion, a feed truncation, or an upstream correction. Avoid removing a previously published item based on a single incomplete fetch.
9. Troubleshooting common problems
| Symptom | Likely cause | Fix |
|---|---|---|
| No items appear | Wrong feed URL, redirect, blocked request, or source has no recent entries | Open the feed URL directly, check its HTTP status and content type, and verify the source’s official feed address. |
| Feed parses as malformed | Invalid XML, encoding mismatch, or an HTML error page served at the feed URL | Inspect the response body and parser error, then contact the publisher or use its documented API. |
| Every poll creates duplicates | The importer uses a changing URL or no stable identity | Use the feed GUID/API ID, normalize URL tracking parameters cautiously, and add a database uniqueness constraint. |
| Items are stale | Worker stopped, polling interval is too long, or errors are being swallowed | Monitor last-success time, surface per-source errors, and retry transient failures with backoff. |
| Excerpts show markup or unsafe content | Feed HTML was treated as plain text or rendered without sanitizing | Escape text; for permitted markup, sanitize with a restrictive allowlist before rendering. |
| Requests are blocked or throttled | Rate limit, authentication requirement, or crawler policy | Check API documentation and terms, reduce polling, honor Retry-After if present, and request authorized access. |
| Dates sort incorrectly | Naive timestamps, missing dates, or mixed time zones | Normalize known dates to UTC and keep missing values null rather than guessing. |
10. Performance, reliability, and cost
The largest avoidable performance mistake is doing upstream fetches in the reader’s request path. Background polling makes page latency depend on your own database and cache, not the slowest publisher. Store enough metadata to update incrementally, index source ID, publication time, and item identity, and paginate the public feed rather than loading every record.
Reliability comes from retaining last-known-good data, isolating sources from one another, bounding retries, and making writes idempotent. One malformed feed should not prevent healthy sources from updating. Expose freshness in internal monitoring and consider a user-facing “updated” timestamp when delays matter.
Costs depend on the number of sources, polling frequency, storage and search requirements, hosting, and any paid API access. Estimate fetch volume as sources multiplied by polls per source over time, then reduce it with sensible cadence, conditional requests, and caching. Check provider rate limits before scaling; the dossier supplies no universal benchmark or cost figure, so measure your own workload.
11. Add screenshots of source pages when useful
An aggregator may need a visual preview for editorial review, archived references, or a card that links to the original. A screenshot is a separate representation from the feed record: preserve the source link and attribution, and make sure your use of the captured page fits your source terms.

You can run a browser automation worker and manage navigation, consent dialogs, timeouts, and image output yourself. If you only need an image of the rendered page, ScreenshotNeo provides a website screenshot API and MCP server. It can remove known cookie banners, newsletter popups, and chat widgets before capture; each step can be turned off. Its response identifies page verdict and billing status, and bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. AI agents can use its MCP tools to take screenshots, inspect page information, and capture PDFs.
Or skip the browser setup
Make one GET request with the page URL. See the ScreenshotNeo API documentation for parameters and options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));
Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month.
12. Frequently asked questions
Can I aggregate content without storing copies?
Yes. You can store identifiers, short permitted excerpts, and source links, then render a link directory or fetch details at display time from an authorized API. Avoid making reader page loads depend on many upstream requests; cache or precompute what your terms allow.
Should an aggregator have its own editorial review?
Usually, if relevance or quality matters. Automated ingestion can admit spam, malformed records, or near-duplicates. A review queue is useful for new sources, uncertain permissions, flagged items, and corrections.
Does robots.txt tell me whether I may republish an article?
No. Robots.txt communicates crawler rules. Check the publisher’s terms, API conditions, and content permissions for reuse.
Can WordPress be both the aggregator and the public site?
It can suit a modest feed-based publication, especially when a block or plugin covers the needed workflow. A separate ingestion service is a better fit once you need more control over normalization, ranking, source-specific retries, or multiple API types.