ScreenshotNeo

BlogEngineering

Link Preview APIs: URL Unfurling and Open Graph Metadata

Build a safe link preview API that extracts Open Graph, Twitter Card, HTML metadata, and provider embeds from any URL.

By the ScreenshotNeo team1 October 20268 min read

Direct answer: A link preview API accepts a URL, fetches the page, extracts Open Graph tags first, then Twitter Card and ordinary HTML metadata, and returns normalized fields such as title, description, image, canonical URL, domain, and favicon. Production systems must treat URL fetching as a security boundary: prevent SSRF, revalidate every redirect, enforce size and time limits, handle malformed HTML and character sets, and cache results by canonical URL.

Open Graph is metadata for a static preview card. The Open Graph protocol defines properties such as og:title, og:description, og:image, and og:url. oEmbed is complementary: it returns a provider-controlled photo, video, rich, or link representation that can include embed HTML.

What URL unfurling returns

A useful response separates normalized values from the raw tags that produced them. Keeping both makes debugging possible when a publisher has conflicting or malformed metadata.

Field Preferred source Purpose
title og:title, then twitter:title, then <title> Card headline
description og:description, then twitter:description, then <meta name="description"> Summary text
image og:image, then Twitter image Preview artwork
canonical_url og:url, then <link rel="canonical">, then final response URL Stable identity and cache key
domain Parsed canonical URL Trust and attribution label
favicon <link rel="icon">, then /favicon.ico Small source icon
provider_embed oEmbed endpoint, when available Interactive provider output

Architecture: fetch, parse, normalize, cache

  1. Validate the submitted URL. Accept only http and https; reject credentials, unsupported schemes, malformed hosts, and private or link-local destinations.
  2. Resolve DNS and connect safely. Check the resolved address against an IP denylist before connecting. Repeat the check after each DNS resolution.
  3. Fetch with limits. Set connect and total timeouts, cap redirects, limit response bytes, and send a realistic user agent. Do not execute page JavaScript unless you have an isolated browser and a reason to render it.
  4. Revalidate redirects. Parse every Location, validate its scheme and destination, and apply SSRF checks again before following it.
  5. Decode and parse. Honor the HTTP charset and HTML declarations, tolerate broken markup, and parse only the head plus enough body for fallback metadata.
  6. Extract in order. Open Graph, Twitter Card, standard HTML, then provider-specific or inferred fields.
  7. Normalize and resolve URLs. Convert relative image, icon, canonical, and oEmbed URLs against the final response URL; preserve the original URL separately.
  8. Cache deliberately. Key by a canonicalized URL, store the fetch timestamp and final URL, and use an explicit freshness policy.

Minimal Node.js unfurling service

The following example uses Node.js 18+ built-in fetch. Add an HTML parser such as cheerio in a real service; the small parser below demonstrates the extraction order without pretending to be a complete HTML parser.

import express from 'express';
import * as cheerio from 'cheerio';

const app = express();
const MAX_BYTES = 2_000_000;
const TIMEOUT_MS = 10_000;

function validateUrl(input) {
  const u = new URL(input);
  if (!['http:', 'https:'].includes(u.protocol) || u.username || u.password) {
    throw new Error('Only public HTTP(S) URLs are accepted');
  }
  return u;
}

function first($, selectors) {
  for (const selector of selectors) {
    const value = $(selector).first().attr('content') || $(selector).first().text();
    if (value?.trim()) return value.trim();
  }
  return null;
}

async function readLimited(response) {
  const reader = response.body.getReader();
  const chunks = [];
  let total = 0;
  for (;;) {
    const { value, done } = await reader.read();
    if (done) break;
    total += value.byteLength;
    if (total > MAX_BYTES) throw new Error('response too large');
    chunks.push(value);
  }
  return Buffer.concat(chunks).toString('utf8');
}

app.get('/preview', async (req, res) => {
  try {
    const original = validateUrl(String(req.query.url));
    const controller = new AbortController();
    const timer = setTimeout(() => controller.abort(), TIMEOUT_MS);
    const response = await fetch(original, {
      redirect: 'follow', signal: controller.signal,
      headers: { 'user-agent': 'PreviewBot/1.0 (+https://your.example/bot)' }
    });
    clearTimeout(timer);
    if (!response.ok) throw new Error(`upstream status ${response.status}`);
    const html = await readLimited(response);
    const finalUrl = new URL(response.url);
    const $ = cheerio.load(html);
    const absolute = (value) => value ? new URL(value, finalUrl).href : null;
    const result = {
      original_url: original.href,
      final_url: finalUrl.href,
      canonical_url: absolute(first($, ['meta[property="og:url"]', 'link[rel="canonical"]'])) || finalUrl.href,
      title: first($, ['meta[property="og:title"]', 'meta[name="twitter:title"]', 'title']),
      description: first($, ['meta[property="og:description"]', 'meta[name="twitter:description"]', 'meta[name="description"]']),
      image: absolute(first($, ['meta[property="og:image"]', 'meta[name="twitter:image"]'])),
      domain: finalUrl.hostname,
      favicon: absolute($('link[rel="icon"], link[rel="shortcut icon"]').first().attr('href')) || `${finalUrl.origin}/favicon.ico`
    };
    res.json(result);
  } catch (error) {
    res.status(400).json({ error: error.message });
  }
});

app.listen(3000);

This sample omits the production SSRF resolver and redirect hook for readability. Put those checks in the HTTP client layer so every caller receives the same protection.

Equivalent API calls

cURL

curl -G 'http://localhost:3000/preview' \
  --data-urlencode 'url=https://example.com/article'

Python

import requests

r = requests.get(
    'http://localhost:3000/preview',
    params={'url': 'https://example.com/article'},
    timeout=15,
)
r.raise_for_status()
print(r.json())

Node.js

const q = new URLSearchParams({ url: 'https://example.com/article' });
const res = await fetch(`http://localhost:3000/preview?${q}`);
if (!res.ok) throw new Error(await res.text());
console.log(await res.json());

Open Graph, Twitter Cards, and oEmbed

Use Open Graph for a portable static card. A page should normally provide:

<meta property="og:title" content="Article title">
<meta property="og:description" content="Short summary">
<meta property="og:image" content="https://example.com/card.jpg">
<meta property="og:url" content="https://example.com/article">
<meta property="og:type" content="article">

Twitter Card tags can provide a platform-specific image or title, so use them as the second source. The HTML <title> and description meta tag are essential fallbacks for pages without social metadata.

Use oEmbed when you need provider-controlled interactive output, such as a playable video or an embedded post. Its response can contain HTML, dimensions, title, and thumbnail. Sanitize or sandbox returned embed HTML; never inject it into an untrusted page without an allowlist.

Security checklist for user-submitted URLs

  • Allow only HTTP and HTTPS.
  • Block loopback, link-local, private, multicast, reserved, and cloud metadata IP ranges for both IPv4 and IPv6.
  • Resolve hostnames yourself and pin the checked address for the connection to reduce DNS rebinding risk.
  • Recheck every redirect target and reject excessive redirect chains.
  • Disable automatic forwarding of user cookies and authorization headers.
  • Limit response size, decompression ratio, connection time, total time, and concurrent fetches.
  • Restrict ports if your use case does not require arbitrary ports.
  • Record URL, final URL, status, content type, byte count, duration, and error class without logging secrets.
  • Return safe text and URLs to clients; escape text in HTML and allowlist image schemes.

Rendering JavaScript-heavy pages

Many sites inject metadata after JavaScript runs. Choose between an HTTP fetch and an isolated browser based on your input mix. HTTP fetching is cheaper and faster for server-rendered pages. Browser rendering handles client-side frameworks but increases latency, memory use, attack surface, and cost. If you render, disable downloads you do not need, enforce the same SSRF checks for every request, isolate the browser process, and apply a page timeout.

Hosted services expose these controls as options. OpenGraph.io documents smart defaults such as auto_proxy, auto_render, and retry, plus plan-based concurrency limits. TryUnfurl documents SSRF protection, redirect handling, encoding, broken-HTML handling, and fallbacks. Verify current quotas, pricing, SLA, and data-processing terms before choosing a provider.

Caching, freshness, and canonical URLs

Canonicalize scheme and hostname, remove default ports, normalize the path, and preserve meaningful query parameters. Do not blindly drop tracking parameters if they change the page. Cache both successful and negative results with different TTLs: metadata can remain stable for hours, while transient failures should be retried sooner. Store the original URL, final URL, canonical URL, ETag, Last-Modified value, fetched time, parser version, and error state. Re-fetch when the canonical URL changes or the freshness policy expires.

Performance, reliability, and cost

Decision Effect
HTTP fetch before browser render Lower latency and resource use; misses client-generated tags
Connection pooling Reduces handshake time for repeated hosts
Bounded concurrency per host Protects your service and avoids overwhelming publishers
Positive and negative caching Reduces duplicate fetches; use short TTLs for transient errors
Retries Retry timeouts and 5xx responses with backoff; do not retry validation or 4xx errors
Managed API Transfers proxy, rendering, retries, and operations to a provider; compare quotas, residency, and total cost

Measure fetch latency, parse latency, cache-hit rate, status classes, render rate, bytes read, and preview completeness. Set separate budgets for DNS, connect, first byte, body download, and total request time so one slow phase is visible.

Troubleshooting

Symptom Likely cause Fix
No title Missing OG and HTML title, or metadata injected by JavaScript Use fallbacks, then enable isolated rendering for that host.
Wrong image Multiple og:image tags or relative URL Define a deterministic first-valid policy and resolve against the final URL.
Redirect loop HTTP/HTTPS or locale redirects Cap redirects, revalidate each target, and return the chain for diagnosis.
Private-network error SSRF protection blocked the destination Keep the block; require an explicit trusted-network integration for internal pages.
Garbled characters Incorrect charset detection Honor HTTP charset, then HTML declaration, then a documented fallback.
Timeout Slow origin, blocked asset, or browser render Use phase timeouts, cache prior results, and avoid waiting for every asset.
429 or 403 Publisher rate limit or bot policy Back off, identify your bot, respect robots and terms, or use a provider with permitted access.
Oversized response Unbounded download or decompression bomb Enforce byte and decompression limits before parsing.

Or skip the browser setup

If you need a reliable page image alongside metadata, ScreenshotNeo is a website screenshot API and MCP server. It removes cookie and consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Responses identify the page verdict and billing status with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for all options. A single request returns PNG, JPEG, WebP, or PDF:

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Features include full-page and element capture, dark mode, device presets, custom CSS and JavaScript, waits, request blocking, headers and cookies, geolocation, resizing, TTL caching, signed links, async webhooks, bulk capture, usage reporting, and PDF output. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Should I use Open Graph or oEmbed?

Use Open Graph for a static card that works across many sites. Use oEmbed when a provider offers an interactive representation you are allowed to embed. Supporting both gives you a static fallback.

Can I trust a page’s canonical URL?

Treat it as publisher-provided data. Validate its scheme and host, store it separately from the requested and final URLs, and decide whether your product allows cross-host canonicalization.

Do I need a headless browser?

No for server-rendered metadata. Add one only for pages whose metadata appears after JavaScript execution, and isolate it with strict network and resource limits.

How long should previews be cached?

Choose a TTL based on how quickly your product needs updates. Store freshness metadata and support explicit refreshes for users who need a current card.

What should happen when unfurling fails?

Return a typed error and the original URL, then let the client render a plain link. Do not expose internal network details or retry indefinitely.