How to Add Website Previews to a Link Directory Built with PHP in India
Build safe, cached website preview cards in PHP with Open Graph metadata, secure URL fetching, and practical fallbacks for missing data.
To add website previews to a PHP link directory, fetch a submitted page’s metadata once on your server, store the title, description, image URL, and fetch status, then render those saved fields as an escaped card. Prefer Open Graph metadata, fall back to the HTML title and hostname, and never fetch the destination again each time someone views the directory. Because URLs come from users, validate the destination and every redirect to prevent server-side request forgery (SSRF).
This pattern works the same whether your directory is hosted in India or elsewhere. Choose hosting and caching based on your audience and operational needs; the implementation itself is not country-specific.
1. Decide what the preview contains
A compact directory card usually needs a linked title, a short description, a thumbnail when one is available, and the destination’s hostname. Open Graph defines og:title, og:type, og:image, and og:url; og:description is commonly used for the description. [Open Graph protocol](https://ogp.me/?library=true&previewmode=true)
| Field | Preferred value | Fallback |
|---|---|---|
| Title | og:title |
HTML <title>, then hostname |
| Description | og:description |
HTML meta description, then empty |
| Image | og:image |
Use a local neutral placeholder |
| Destination | Normalized submitted URL | Reject invalid input |
| Site name | og:site_name |
Hostname |
A preview card does not need a rich, interactive embed. Use oEmbed only for supported services where provider-specific data or playback materially improves the entry. oEmbed supports discovery through HTML link elements and HTTP Link headers, and its response can describe link, photo, video, or rich representations. Treat provider HTML as untrusted. [oEmbed specification](https://oembed.com/)
2. Create a metadata fetcher in PHP
The example below demonstrates the shape of the fetch-and-parse code. Its host validation deliberately fails closed: production systems should implement address validation at connection time, including after DNS resolution and for each redirect, rather than trusting a one-time hostname check. Do not deploy a user-controlled fetcher without that SSRF protection.
<?php
declare(strict_types=1);
function normalizeSubmittedUrl(string $raw): string
{
$url = trim($raw);
if ($url === '' || strlen($url) > 2048) {
throw new InvalidArgumentException('URL is empty or too long.');
}
$parts = parse_url($url);
if ($parts === false || empty($parts['scheme']) || empty($parts['host'])) {
throw new InvalidArgumentException('Enter an absolute HTTP or HTTPS URL.');
}
$scheme = strtolower($parts['scheme']);
if (!in_array($scheme, ['http', 'https'], true)) {
throw new InvalidArgumentException('Only HTTP and HTTPS URLs are allowed.');
}
if (isset($parts['user']) || isset($parts['pass'])) {
throw new InvalidArgumentException('Credentials in URLs are not allowed.');
}
if (isset($parts['port']) && !in_array((int)$parts['port'], [80, 443], true)) {
throw new InvalidArgumentException('This port is not allowed.');
}
return $url;
}
function fetchHtml(string $url): string
{
// IMPORTANT: Replace this placeholder with a hardened fetcher that resolves
// and blocks private/reserved IPv4 and IPv6 addresses, pins the validated
// address used for connection, and repeats validation for every redirect.
$host = parse_url($url, PHP_URL_HOST);
if (!$host || filter_var($host, FILTER_VALIDATE_IP)) {
throw new RuntimeException('Host validation requires the hardened fetch service.');
}
$ch = curl_init($url);
curl_setopt_array($ch, [
CURLOPT_RETURNTRANSFER => true,
CURLOPT_FOLLOWLOCATION => false, // validate each redirect before following
CURLOPT_CONNECTTIMEOUT => 4,
CURLOPT_TIMEOUT => 10,
CURLOPT_MAXREDIRS => 0,
CURLOPT_PROTOCOLS => CURLPROTO_HTTP | CURLPROTO_HTTPS,
CURLOPT_HTTPHEADER => ['Accept: text/html,application/xhtml+xml;q=0.9,*/*;q=0.1'],
CURLOPT_USERAGENT => 'DirectoryPreviewBot/1.0',
CURLOPT_WRITEFUNCTION => null,
]);
$body = curl_exec($ch);
$errno = curl_errno($ch);
$status = (int)curl_getinfo($ch, CURLINFO_RESPONSE_CODE);
$type = strtolower((string)curl_getinfo($ch, CURLINFO_CONTENT_TYPE));
curl_close($ch);
if ($body === false || $errno !== 0) {
throw new RuntimeException('Could not fetch page.');
}
if ($status < 200 || $status >= 400) {
throw new RuntimeException('Remote page returned HTTP ' . $status . '.');
}
if ($type !== '' && !str_contains($type, 'text/html') && !str_contains($type, 'application/xhtml+xml')) {
throw new RuntimeException('Remote resource is not HTML.');
}
if (strlen($body) > 1_000_000) {
throw new RuntimeException('Remote HTML exceeds the size limit.');
}
return $body;
}
function metaContent(DOMXPath $xpath, string $property): ?string
{
$nodes = $xpath->query('//meta[translate(@property,"ABCDEFGHIJKLMNOPQRSTUVWXYZ","abcdefghijklmnopqrstuvwxyz")="' . strtolower($property) . '"]/@content | //meta[translate(@name,"ABCDEFGHIJKLMNOPQRSTUVWXYZ","abcdefghijklmnopqrstuvwxyz")="' . strtolower($property) . '"]/@content');
if (!$nodes || $nodes->length === 0) return null;
$value = trim(preg_replace('/\s+/u', ' ', $nodes->item(0)->nodeValue ?? ''));
return $value !== '' ? mb_substr($value, 0, 500) : null;
}
function extractPreview(string $url, string $html): array
{
$previous = libxml_use_internal_errors(true);
$dom = new DOMDocument();
$dom->loadHTML($html, LIBXML_NONET | LIBXML_NOERROR | LIBXML_NOWARNING);
libxml_clear_errors();
libxml_use_internal_errors($previous);
$xpath = new DOMXPath($dom);
$titleNode = $xpath->query('//title');
$htmlTitle = $titleNode && $titleNode->length ? trim($titleNode->item(0)->textContent) : '';
$host = strtolower((string)parse_url($url, PHP_URL_HOST));
$title = metaContent($xpath, 'og:title') ?? ($htmlTitle !== '' ? mb_substr($htmlTitle, 0, 200) : $host);
$description = metaContent($xpath, 'og:description') ?? metaContent($xpath, 'description') ?? '';
$image = metaContent($xpath, 'og:image');
$siteName = metaContent($xpath, 'og:site_name') ?? $host;
return compact('title', 'description', 'image', 'siteName', 'host');
}
$url = normalizeSubmittedUrl($_POST['url'] ?? '');
$html = fetchHtml($url); // call only through an SSRF-hardened fetch layer
$preview = extractPreview($url, $html);
// Persist URL, preview fields, fetch timestamp, and status in your database.
?>
Important limits in this illustrative code: it refuses to follow redirects but does not implement safe redirect handling or connection-time DNS pinning. The CURLOPT_WRITEFUNCTION setting is not the byte limiter; production code must stream the response and abort after a configured byte ceiling. Use a dedicated hardened fetch layer or implement those controls before enabling arbitrary user submissions. Set both a connection timeout and a total timeout, and never forward cookies, Authorization headers, or internal service credentials.
3. Harden URL fetching against SSRF
A server-side preview fetch turns a user-provided URL into an outbound network request. Attackers may submit loopback addresses, private network hosts, cloud metadata destinations, or public hostnames that resolve to those addresses. Redirects and DNS rebinding can bypass naive hostname checks.
- Accept only absolute
httpandhttpsURLs. Reject embedded credentials and ports outside your explicit policy. - Resolve the hostname and reject loopback, private, link-local, multicast, and reserved ranges for both IPv4 and IPv6.
- Ensure the validated address is the address used by the connection. Do not validate DNS once and then allow an independent resolution when connecting.
- Disable automatic redirects. For each redirect, resolve and validate the new destination before making another request; cap redirect count.
- Apply short connection and total timeouts, a strict response-byte cap, and a small set of accepted content types. Do not send user cookies or authorization.
- Run the fetcher with restricted network access where practical, and log rejected destinations without storing secrets.
The code above intentionally leaves the crucial address-pinning and safe-redirect implementation to a hardened fetch layer. A hostname check using gethostbyname() alone is not a complete defense because it may check a different answer from the one the HTTP client eventually connects to.
4. Store previews and render safely
Store the normalized URL, title, description, image URL, site name, fetch timestamp, and a status such as success, timeout, blocked, or invalid-content. Cache successes and failures for a policy-appropriate interval. Refresh via a background job or controlled schedule so third-party latency does not slow directory page views.
<?php
function h(?string $value): string {
return htmlspecialchars($value ?? '', ENT_QUOTES | ENT_SUBSTITUTE, 'UTF-8');
}
$placeholder = '/assets/link-placeholder.png';
$imageUrl = $preview['image'] ?: $placeholder;
?>
<article class="link-card">
<a href="<?= h($preview['url'] ?? '') ?>" rel="nofollow noopener noreferrer">
<img src="<?= h($imageUrl) ?>" alt="" loading="lazy">
<h2><?= h($preview['title'] ?? '') ?></h2>
<p><?= h($preview['description'] ?? '') ?></p>
<small><?= h($preview['host'] ?? '') ?></small>
</a>
</article>
Escape every text node and attribute at output time, even if the value came from a parser. Validate image URLs independently; a remote image can break, track visitors, or point to an unexpected destination. If you proxy images, apply the same SSRF and byte-limit controls. A restrictive Content Security Policy and a local placeholder help contain failures. Do not place arbitrary oEmbed rich HTML directly into your directory origin; if a rich embed is necessary, use a tightly controlled sandboxed iframe on an isolated origin. [oEmbed security considerations](https://oembed.com/)
5. Add submission handling and refresh behavior
- Validate the submitted URL before saving it. Return a clear validation error without making a network request for malformed input.
- Create the directory record with a pending preview state, then enqueue extraction. This keeps form submission responsive.
- Fetch and parse in a worker with the network limits above. Save a success or failure status and the time of the attempt.
- Show a neutral card while pending or when extraction fails. Let visitors open the submitted URL, but do not treat missing metadata as a fatal directory-entry error.
- Refresh stale records on a schedule or when an editor requests it. Add per-user and per-host rate limits to prevent abuse.
For a small low-traffic directory, synchronous extraction may be acceptable if the request is strictly time-bounded and failures do not block the listing. A queue is more reliable as volume or remote-site latency grows.
6. Handle missing metadata, redirects, and image URLs
- No Open Graph tags: use the document title, hostname, and a placeholder image; keep the description blank if none exists.
- Relative image path: resolve it against the final page URL after redirects, then validate the resulting absolute URL.
- Redirect: validate every hop, cap the number of hops, and store both the submitted URL and the final canonical destination if your product needs both.
- Wrong content type: reject PDFs, archives, and other non-HTML resources for metadata extraction rather than parsing arbitrary bodies.
- Oversized or slow page: abort at the byte or time cap and record a retryable failure state.
- Malformed HTML: use a tolerant HTML parser and accept that some fields may be unavailable.
- Image removed later: allow the browser’s image error behavior or serve a controlled placeholder. Consider a proxy only if you can operate it with the same security controls.
- Untrusted descriptions: normalize whitespace, cap lengths, and escape at render time. Do not assume HTML entities or provider values are safe.
7. When oEmbed is useful
For general directory cards, Open Graph and ordinary HTML metadata are usually enough. Add provider-specific oEmbed only when a supported service provides meaningful structured data or an interaction such as video playback. The specification describes endpoint discovery via an advertised application/json+oembed or XML link relation or a Link header; support should be allow-listed and provider-aware. [oEmbed specification](https://oembed.com/)
Rich responses may include HTML. Never inject that HTML into the directory page. A link-only response can be handled as metadata; interactive representations need a tightly sandboxed, isolated embedding design. The specification lists PHP consumer libraries such as Essence and Embera, but check current maintenance, PHP compatibility, dependency health, and provider support before adopting one. [oEmbed implementations](https://oembed.com/)
8. cURL, Python, and Node.js equivalents
The directory implementation can remain in PHP. These snippets show the equivalent HTTP fetch pattern for auxiliary workers or services; they do not implement SSRF defenses. Apply the same destination validation, redirect checks, timeouts, and byte limits in whichever language performs the fetch.
cURL
curl --max-time 10 --connect-timeout 4 --max-redirs 0 \
-H 'Accept: text/html' \
-A 'DirectoryPreviewBot/1.0' \
'https://example.com/'
Python
import requests
from urllib.parse import urlparse
url = "https://example.com/"
parts = urlparse(url)
if parts.scheme not in {"http", "https"} or not parts.hostname:
raise ValueError("Only absolute HTTP(S) URLs are accepted")
if parts.username or parts.password:
raise ValueError("URL credentials are not allowed")
with requests.get(
url,
headers={"Accept": "text/html", "User-Agent": "DirectoryPreviewBot/1.0"},
timeout=(4, 10),
allow_redirects=False,
stream=True,
) as response:
response.raise_for_status()
if "text/html" not in response.headers.get("Content-Type", "").lower():
raise ValueError("Response is not HTML")
body = bytearray()
for chunk in response.iter_content(16_384):
body.extend(chunk)
if len(body) > 1_000_000:
raise ValueError("Response too large")
html = bytes(body)
Node.js
const url = new URL('https://example.com/');
if (!['http:', 'https:'].includes(url.protocol) || url.username || url.password) {
throw new Error('Only credential-free HTTP(S) URLs are accepted');
}
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 10_000);
try {
const response = await fetch(url, {
redirect: 'manual',
signal: controller.signal,
headers: { Accept: 'text/html', 'User-Agent': 'DirectoryPreviewBot/1.0' },
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
if (!(response.headers.get('content-type') || '').toLowerCase().includes('text/html')) {
throw new Error('Response is not HTML');
}
const reader = response.body.getReader();
const chunks = [];
let size = 0;
while (true) {
const { value, done } = await reader.read();
if (done) break;
size += value.byteLength;
if (size > 1_000_000) { await reader.cancel(); throw new Error('Response too large'); }
chunks.push(value);
}
const html = Buffer.concat(chunks).toString('utf8');
} finally {
clearTimeout(timer);
}
These examples stop automatic redirects so your application can validate each Location target before following it. They still need a connection layer that pins validated DNS results; application-level parsing by itself is not sufficient protection against rebinding.
9. Performance, reliability, and operating cost
- Do not fetch on page view. Saved preview fields make directory rendering independent of third-party latency and availability.
- Use caching. Cache successful and failed extraction attempts, with different refresh policies if useful. Avoid retrying a broken host on every submission or page load.
- Queue work. Background jobs absorb spikes and allow bounded concurrency. Apply per-host limits so one destination cannot consume all workers.
- Bound resources. Timeouts, byte limits, redirect caps, and worker memory limits control compute and outbound traffic.
- Measure failure classes. Track timeout, blocked destination, HTTP error, unsupported type, parse-empty, and successful outcomes without logging sensitive request data.
- Choose image handling deliberately. Direct remote images reduce your storage and bandwidth use but rely on the remote host. Proxying gives more control but adds bandwidth, storage, security, and maintenance cost.
- Use a hosted unfurl service when appropriate. Compare supported sites, caching, privacy, latency, pricing, and operational maintenance. OpenGraph.io documents metadata extraction, caching, rendering, and proxy options; the cited documentation does not establish its current commercial terms or India performance. [OpenGraph.io API docs](https://www.opengraph.io/docs/api/site)
10. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Every URL is rejected | URL parser or scheme policy is too restrictive or input lacks a scheme | Require an absolute URL and show an example; do not silently guess a scheme. |
| Fetch times out | Slow origin, blocked egress, or overly short limit | Keep limits bounded; show a fallback card and allow a controlled retry. |
| HTTP 403 or challenge page | Destination blocks automated fetching | Store a fetch failure and use hostname/placeholder fallbacks; do not attempt to bypass access controls. |
| Title is empty | Client-rendered page, malformed HTML, or missing tags | Use fallbacks. If rendering is essential, use a purpose-built rendering service with the same URL security controls. |
| Preview image is broken | Relative URL mishandled, remote image removed, or hotlinking denied | Resolve against final page URL, validate separately, and show a local placeholder on failure. |
| Internal hosts can be reached | Only string or hostname checks are used, or redirects/DNS are unchecked | Block private and reserved IPv4/IPv6 destinations, validate every redirect, and pin the validated address at connection time. |
| Directory page becomes slow | Metadata is fetched during each visitor request | Persist results and refresh them asynchronously. |
| Database contains markup or script-like values | Untrusted metadata is rendered without contextual escaping | Escape at output using the correct context; never render provider HTML as ordinary text. |
11. “How to add website previews to a link directory built with PHP in India”: practical sequence
- Add preview columns or a related table for title, description, image URL, site name, status, fetched time, and final URL.
- Validate submitted URLs and enqueue extraction instead of making the directory visitor wait.
- Run a hardened fetcher with SSRF defenses, byte and time caps, content-type checks, and manual redirect validation.
- Parse Open Graph fields with a tolerant parser; use document title and hostname fallbacks.
- Escape every rendered value, validate image destinations, and show a placeholder when preview data is missing.
- Cache both success and failure, monitor outcomes, and refresh stale entries under controlled limits.
- Add allow-listed oEmbed only for providers where an embed is worth the added complexity.
Or skip the browser setup
For a preview image rather than metadata extraction, ScreenshotNeo offers a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF; its parameters also use names familiar from other screenshot APIs. Cookie banners, popups, and chat widgets are removed before the shot, and each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers state the page verdict and billing status. An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for configuration. Free includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card required.
FAQ
Should I store preview fields or fetch them on every page view?
Store them. This keeps directory pages fast and prevents availability of third-party sites from affecting visitors.
Do all websites provide Open Graph tags?
No. Fall back to the document title and hostname, and use a placeholder when no image is provided.
Can I put the URL fetch in the PHP form handler?
You can for a small directory if it is strictly time-bounded, but a background worker gives users a faster and more reliable submission flow.
Is this implementation specific to India?
No. The URL, metadata, and PHP patterns are the same. Select infrastructure based on audience latency and operational requirements.
When should I use oEmbed instead of Open Graph?
Use it for an allow-listed provider when structured media or interaction adds real value; ordinary cards need only metadata.


