How to Extract Metadata from Websites with Metascraper
Build a reliable Node.js metadata extractor with Metascraper, browser-rendered HTML, fallbacks, custom rules, validation, and production troubleshooting.

Metascraper extracts normalized metadata from a URL and its HTML. It can resolve a page title, description, image, author, publication date, language, logo, publisher, canonical URL, audio and video by combining Open Graph, ordinary HTML tags, JSON-LD, Microdata, RDFa, Twitter Cards and other rule bundles. The practical pattern is: retrieve accurate HTML, pass the URL and markup to Metascraper, then inspect or store the resolved object.
This guide shows a complete Node.js implementation, browser rendering for JavaScript-heavy pages, selective fields, custom rules, validation, failure handling, scaling considerations and a managed screenshot option when browser setup is unnecessary.
What Metascraper needs
Metascraper requires two inputs: the target URL and the HTML markup behind that URL. The URL is used to resolve relative links such as /images/cover.jpg, and can act as a fallback value for URL-related rules. The HTML can come from a normal HTTP request or a headless browser. Choose the lightest retrieval method that returns the same markup a user would see.
A simple server-rendered article may work with fetch. A single-page application that inserts Open Graph data after JavaScript runs needs a browser context. The official project example uses html-get with browserless, then passes the resulting HTML to Metascraper. See the official Metascraper documentation and README for the supported bundles and API.
Install the packages
mkdir metadata-extractor
cd metadata-extractor
npm init -y
npm install metascraper metascraper-author metascraper-date metascraper-description metascraper-image metascraper-lang metascraper-logo metascraper-publisher metascraper-title metascraper-url html-get browserless
The individual metascraper-* packages are rule bundles. Keeping bundles explicit makes the output predictable and avoids loading properties your application does not need.
Complete browser-backed example
The following program retrieves browser-rendered HTML and extracts common article fields. It follows the official integration pattern; it is a template to adapt to your runtime and deployment model.

const getHTML = require('html-get')
const browserless = require('browserless')()
const metascraper = require('metascraper')([
require('metascraper-author')(),
require('metascraper-date')(),
require('metascraper-description')(),
require('metascraper-image')(),
require('metascraper-lang')(),
require('metascraper-logo')(),
require('metascraper-publisher')(),
require('metascraper-title')(),
require('metascraper-url')()
])
const getContent = async url => {
const browserContext = browserless.createContext()
try {
return await getHTML(url, {
getBrowserless: () => browserContext
})
} finally {
await browserContext
.then(browser => browser.destroyContext())
.catch(() => {})
}
}
async function main () {
const url = process.argv[2]
if (!url) throw new Error('Usage: node extract.js https://example.com/article')
const html = await getContent(url)
const metadata = await metascraper({ url, html })
console.log(JSON.stringify(metadata, null, 2))
await browserless.close()
}
main().catch(async error => {
console.error(error)
await browserless.close().catch(() => {})
process.exitCode = 1
})
Save this as extract.js, then run node extract.js https://example.com/article. The result is a plain object whose properties depend on the bundles and rules that found matches. A missing field is normal: it means no configured rule produced a value.
How rule resolution and fallbacks work
Metascraper is assembled from small rule bundles. Each bundle contains selectors and transformations for one property. Rules run from specific signals to generic fallbacks; the first successful rule wins. For a title, an Open Graph value can therefore win over a generic HTML heading, while a generic HTML rule remains available when og:title is absent.
This ordering is why you should configure several sources rather than relying on one tag family. It also means the returned value is the best resolved candidate, not proof that every source agreed. Keep the source URL and, if auditability matters, retain the original HTML or the signals you used to explain a result.
Fields and rule bundles
| Property | Typical use | Bundle |
|---|---|---|
| title | Article or page heading | metascraper-title |
| description | Search and preview summary | metascraper-description |
| image | Preview or social card image | metascraper-image |
| author | Byline normalization | metascraper-author |
| date | Publication or modified date | metascraper-date |
| lang | Document language | metascraper-lang |
| logo | Publisher or site mark | metascraper-logo |
| publisher | Organization or site name | metascraper-publisher |
| url | Canonical or resolved URL | metascraper-url |
The project also lists bundles for audio, video, citation metadata, feeds, readability, manifests and vendor-specific sources such as Amazon, Instagram, Reddit, Spotify, TikTok, X and YouTube. Install only the bundles relevant to your output contract, or add more when you need broader coverage.

Select only the properties you need
Use pickPropNames to limit execution and output. It accepts a Set. This is useful for link previews that need only a title, description and image.
const metadata = await metascraper({
url: 'https://example.com/article',
html,
pickPropNames: new Set(['title', 'description', 'image'])
})
omitPropNames excludes properties. If both are supplied, pickPropNames takes precedence. Keep this distinction in mind when a shared helper adds defaults: an explicit pick list is the narrower contract.
URL validation and relative links
validateUrl defaults to true and checks WHATWG URL compliance. Leave validation enabled for user-provided URLs. If an internal pipeline already guarantees a special URL format, you can deliberately change the setting, but do so only when you understand which rules depend on a valid absolute URL.
const metadata = await metascraper({
url: 'https://example.com/article',
html,
validateUrl: true
})
Always pass the final URL after redirects when possible. Metascraper uses it to turn relative image, canonical and publisher links into usable absolute URLs. If you pass a short or pre-redirect URL, fields may still be extracted but can point at the wrong origin.
Static HTML versus browser-rendered HTML
Start with a normal HTTP client when the response already contains the metadata. It is cheaper and simpler than a browser. Use a headless browser when a site generates tags after JavaScript runs, requires interaction before content appears, or serves substantially different markup to browsers and basic clients.
- Static request: lower latency and resource use; works for server-rendered pages.
- Browser retrieval: executes JavaScript and can expose client-rendered metadata; costs more CPU and requires lifecycle cleanup.
- Hybrid strategy: fetch first, detect missing required fields, then retry with a browser only for those URLs.
Do not assume every modern site needs a browser. Compare the HTML returned by your HTTP client with the browser DOM for representative URLs, and choose based on whether the required fields are present.
Custom rules and per-call overrides
Rules are extensible. You can add a custom bundle for a publisher-specific tag, or pass additional rules when calling Metascraper. A custom rule should have a clear property name, selectors ordered from most specific to most general, and a transformation that returns either a normalized value or no value.
const metascraper = require('metascraper')([
require('metascraper-title')(),
({ url, htmlDom }) => ({
'article-id': htmlDom('meta[name="article-id"]').attr('content') || undefined
})
])
const metadata = await metascraper({ url, html })
The exact rule shape can vary with the Metascraper version and the helper packages you use, so consult the project README before publishing a custom bundle. Keep custom properties namespaced when multiple teams own the output, and write fixtures for pages where the publisher changes markup.
Handling missing or conflicting Open Graph tags
- Fetch the HTML that a real visitor receives.
- Check whether
og:title,og:description,og:imageandarticle:published_timeexist. - Allow Metascraper’s HTML, JSON-LD, Microdata, RDFa and Twitter Card rules to provide fallbacks.
- Use
pickPropNamesto enforce the small set your application actually needs. - Log the URL and missing property names so you can identify publisher-specific patterns.
When values conflict, define an application policy. For example, accept Metascraper’s title for display, but reject a record if no image is found; or store both the resolved value and a source-specific value when editorial review requires provenance.
Production checklist
- Set request and browser timeouts. A page that never finishes should become a classified failure, not a stuck worker.
- Close browser contexts in a
finallyblock. - Limit concurrency to what your CPU, memory and target sites can handle.
- Cache HTML or metadata for URLs that are requested repeatedly.
- Normalize redirects and remove tracking parameters when your application permits it.
- Persist the original URL, retrieval timestamp and parser version with the result.
- Validate required fields before writing a record.
- Protect your service from SSRF: restrict schemes, private network ranges and internal hostnames when URLs are user-controlled.
Performance, reliability and cost
Parsing HTML is usually less expensive than opening a browser. The browser retrieval step dominates latency and memory, especially when pages load many scripts, images or advertisements. A two-stage fetch-then-browser fallback often gives better throughput than rendering every URL.
Reliability depends on the retrieval layer as much as the parser. DNS failures, TLS errors, robots policies, bot challenges, login walls, infinite client-side loading and rate limits can all prevent accurate HTML. Record these outcomes separately from “no metadata found”; the remediation is different.
Metascraper itself is a library, so your infrastructure pays for HTTP traffic, browser processes, proxies and storage. At larger volumes, operating browsers, proxy rotation, antibot workarounds and paywall access can become the main operational burden. The project documentation describes the managed Microlink API as a pay-as-you-go option that starts free; check its live service terms, pricing, quotas and regional availability before choosing it.
Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
Invalid URL |
Missing scheme or malformed input | Require an absolute http or https URL and keep validateUrl enabled. |
| All fields are empty | Wrong or empty HTML | Log response status and a short HTML sample; retry with browser-rendered retrieval. |
| Image is relative | URL was omitted or not the final page URL | Pass the absolute target or post-redirect URL. |
| Browser contexts accumulate | Cleanup skipped on an exception | Destroy the context in finally and close the browserless client during shutdown. |
| Metadata changes between runs | Dynamic page, A/B test or changing source tags | Store retrieval time and HTML; define a preferred source policy. |
| Requests are blocked | Rate limit, bot check or restricted page | Back off, respect site policies and use an appropriate managed retrieval service when authorized. |
| Process runs out of memory | Too many concurrent browsers or unbounded pages | Bound concurrency, close contexts promptly and use static fetches where possible. |
Or skip the browser setup
If your goal is a clean visual record of a page alongside its metadata, ScreenshotNeo provides a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP or PDF. Before capture it accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.
See the ScreenshotNeo API documentation for all options. A minimal call is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Options include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or any viewport, retina scale, PDF paper sizes and page ranges, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, custom headers and cookies, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture for 100 URLs per call and a usage API. An MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.
There are 1,000 free screenshots each month with no card. Paid plans start at $5 for 3,000 screenshots, and every feature is available on every plan. Create a free ScreenshotNeo account.
FAQ
Does Metascraper download a URL by itself?
No. Supply both the URL and the HTML. Use an HTTP client for static pages or a browser integration for client-rendered pages.
Which source wins when tags disagree?
The first successful rule in the configured ordering wins. Configure multiple fallbacks and document the source policy your application expects.
Can I extract only one field?
Yes. Use pickPropNames with a set such as new Set(['title']).
Should every URL be rendered in a browser?
No. Browser rendering is useful when the required metadata appears only after JavaScript runs. Start with the lightest retrieval method that returns accurate HTML.
How should I handle a missing image?
Treat it as a normal incomplete result, preserve the URL and retrieval details, and decide whether your product should use another fallback image or reject the record.
Where can I find the supported bundles?
The Metascraper README lists official bundles, controls and additional integrations.


