How to Create Website Previews for a Bengali Link Directory
Build safe, useful link preview cards with Open Graph metadata, cached fetching, Bengali language support, and optional screenshots.
For most Bengali link directories, build preview cards from Open Graph and ordinary HTML metadata, not from a fresh screenshot on every page view. Accept and normalize an HTTP(S) URL, fetch it in an isolated background worker, extract and validate title, description, and image fields, cache the result, then render a card with a safe fallback. Add screenshots only when metadata is absent or does not represent the page well.
This approach keeps directory rendering independent of slow or unsafe remote sites. It does not guarantee coverage: sites differ in metadata, access controls, JavaScript behavior, and language rendering. Measure the approaches on a representative set of the sites your directory accepts.
1. Choose the preview source
| Approach | What the card shows | Trade-offs |
|---|---|---|
| Metadata-only | Publisher-provided title, description, and image | Simpler pipeline; results depend on each site’s tags and image URLs. |
| Screenshot-only | A rendered view of the page | Shows appearance when metadata is poor, but needs browser rendering, more resources, and image storage or delivery. |
| Hybrid | Metadata by default, screenshot fallback or secondary thumbnail | Lets you reserve browser work for URLs where it improves the card; adds a second capture path. |
Use metadata as the default because it gives a compact, readable card without rendering a full browser. A screenshot can be useful when the source tags are missing or misleading. Compare the options on your own URL set for success rate, response time, storage and bandwidth, stale previews, and readability at card size; available research does not establish a universal winner.
2. Build a safe preview pipeline
- Accept and normalize: require an absolute HTTP or HTTPS URL, reject credentials and unexpected schemes, normalize the hostname carefully, and define whether non-public destinations are allowed.
- Queue the fetch: do not make a directory page wait for arbitrary publisher servers. Put preview work in a worker with minimal privileges and restricted network egress.
- Validate each destination: resolve and reject prohibited address ranges at connection time, and apply the same checks after every redirect. Limit redirect count, connection and read time, response bytes, and overall resource use.
- Extract and validate: parse metadata, normalize text and image URLs, reject unsupported or unsafe values, and retain the publisher domain for context.
- Cache: key results by normalized URL and refresh on a controlled schedule. Serve the last useful result while a refresh is pending.
- Render a resilient card: show the destination, domain, title, optional description, and image or placeholder. A failed remote image should not collapse the card.
User-submitted URLs make the fetcher a security boundary. Server-side request forgery can let an attacker make your server contact internal services or consume resources; redirects can be used to evade checks. Validate schemes and destinations, revalidate redirects, restrict egress, and avoid returning fetched HTML or internal diagnostics to the requester. See [MDN’s SSRF guidance](https://developer.mozilla.org/en-US/docs/Web/Security/Attacks/SSRF).
3. Extract metadata with a priority order
Open Graph defines fields for shareable objects, including og:title, og:type, og:image, and og:url. Treat them as a strong source, not a promise that every site has complete tags. A practical priority order is:
og:title, then the document’s<title>, then the host name.og:description, then the ordinary meta description, otherwise omit the description.og:image, then an approved Twitter Card image field if your parser supports it, otherwise use a generic placeholder.- A validated canonical URL or
og:urlfor identity where appropriate; keep the submitted destination available and do not silently send users somewhere else.
Normalize whitespace, cap title and description lengths for your card design, and preserve the original Unicode text. Resolve relative image URLs against the final fetched page URL only after validating the resulting destination. Do not trust metadata as safe HTML: escape text in the rendered page, and validate image URLs according to your loading or proxy policy.
4. Minimal runnable metadata fetcher in Python
This example demonstrates the parsing shape for a trusted test URL. It is not a safe public-URL fetch service: production code must add the destination, DNS/IP, redirect, egress, byte, and time controls described above. Use a maintained HTML parser and a queue/worker in a real directory.
from html.parser import HTMLParser
from urllib.parse import urljoin, urlsplit
import requests
class Metadata(HTMLParser):
def __init__(self):
super().__init__()
self.title = None
self.in_title = False
self.meta = {}
def handle_starttag(self, tag, attrs):
attrs = dict(attrs)
if tag.lower() == "title":
self.in_title = True
if tag.lower() == "meta":
key = (attrs.get("property") or attrs.get("name") or "").lower()
value = attrs.get("content")
if key and value:
self.meta.setdefault(key, value.strip())
def handle_endtag(self, tag):
if tag.lower() == "title":
self.in_title = False
def handle_data(self, data):
if self.in_title:
text = data.strip()
if text:
self.title = (self.title or "") + text
def preview(url):
parts = urlsplit(url)
if parts.scheme not in ("http", "https") or not parts.hostname or parts.username or parts.password:
raise ValueError("Expected an absolute HTTP(S) URL without credentials")
# Demonstration only: production must validate resolved IPs and every redirect,
# and should fetch in a restricted worker with strict resource limits.
response = requests.get(url, timeout=(3, 8), allow_redirects=False,
headers={"User-Agent": "DirectoryPreviewBot/1.0"})
if 300 <= response.status_code < 400:
raise ValueError("Redirect requires destination validation before following")
response.raise_for_status()
if len(response.content) > 1_000_000:
raise ValueError("Response exceeds the demonstration size limit")
if "html" not in response.headers.get("Content-Type", "").lower():
raise ValueError("Target did not return HTML")
parser = Metadata()
parser.feed(response.text)
title = parser.meta.get("og:title") or parser.title or parts.hostname
description = parser.meta.get("og:description") or parser.meta.get("description")
image = parser.meta.get("og:image")
if image:
image = urljoin(response.url, image)
image_parts = urlsplit(image)
if image_parts.scheme not in ("http", "https") or not image_parts.hostname:
image = None
return {"title": title, "description": description,
"image": image, "domain": parts.hostname}
if __name__ == "__main__":
print(preview("https://example.com/"))
The sample intentionally refuses redirects instead of following them without checks. A production worker can follow a bounded number only if it validates each new destination and enforces its network and resource policy.
5. Render cards accessibly and support Bengali
Set the page language to the language of the directory interface, and mark passages in another language when needed. For a Bengali-language interface, declare lang="bn" on the root HTML element; for an English interface, use lang="en" and mark Bengali excerpts with lang="bn". Language declarations help assistive technologies choose pronunciation and processing rules. See [W3C WAI technique H57](https://www.w3.org/WAI/WCAG21/Techniques/html/H57).
Preserve Bengali source text, including conjunct characters. Test long Bengali titles, truncation, fallback fonts, and mixed-language content on the devices your readers use. The language declaration is defined by the standard, but the right font and rendering behavior should be checked on actual target devices.
Give the card an accessible name that makes sense without its image. If the image is decorative because the title already names the destination, use empty alternative text; if it conveys distinct information, provide appropriate alternative text. Keep the destination visibly identifiable. Handle image load failure by replacing it with a placeholder without changing card dimensions.
6. Load images without layout jumps
For cards below the fold, use lazy loading and reserve a stable image box with declared dimensions or an aspect ratio:
<img src="/preview-image/42" alt="" loading="lazy" width="320" height="180">
A fixed ratio can be used when image dimensions vary. Avoid lazy loading the first visible card image if that would delay above-the-fold content. See [MDN’s image performance guidance](https://developer.mozilla.org/en-US/docs/Web/HTML/Element/img#loading).
7. Choose image delivery and CSP deliberately
If browsers load publisher images directly, your Content Security Policy must allow the origins you intend to use in img-src. Arbitrary per-site origins can make that policy broad. A proxy or controlled image store can give the page a narrower image origin, but it adds another server-side fetch path, plus storage, bandwidth, privacy, and security responsibilities. Apply URL and network protections to that path too. The img-src directive controls valid image sources; see [MDN’s CSP reference](https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Headers/Content-Security-Policy/img-src).
8. Add screenshots only where they help
Use a screenshot as a fallback when metadata is missing or unhelpful, or as an alternate card mode if the visual appearance is the feature users need. A screenshot renderer must load remote pages in an isolated environment and apply the same URL and resource safety policy. Decide whether to capture on submission, on a background refresh, or only after metadata fails. Cache the result so visits do not trigger a browser render each time.
Evaluate metadata-only, screenshot-only, and hybrid behavior on the same representative URL set. Track extraction success, useful-card rate, time to ready, stale results, bandwidth, storage, and any vendor or compute cost. These are decision measures, not outcomes established by the research.
Or skip the browser setup
For the screenshot fallback, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. Its capture can accept cookie and consent banners like a visitor and remove 60+ known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
See the ScreenshotNeo API documentation. Example request:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await require('node:fs/promises').writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo includes full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets and arbitrary viewports, retina scale, PDF options, HTML/CSS capture, custom CSS and JavaScript, click-before-capture, selector hiding, wait controls, request and resource blocking, custom headers, cookies, user agent and Authorization, timezone and geolocation, transparent background, image resizing, configurable cache TTL, signed links, asynchronous jobs with signed webhooks, bulk capture for 100 URLs per call, a usage API, and an OpenAPI spec. Parameter names used by other screenshot APIs also work. Every feature is on every plan. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots, with higher tiers available and yearly billing giving two months free.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
9. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| No title or a generic title | Missing tags, malformed HTML, or page title added by JavaScript. | Try the document title and hostname fallbacks. If a rendered title is essential, consider a screenshot/browser mode and measure coverage. |
| Image missing or broken | Relative or invalid image URL, publisher blocks hotlinking, image removed, or CSP disallows its origin. | Resolve against the validated final page URL, handle image errors, check img-src, and consider a protected proxy or placeholder. |
| Fetch hangs or consumes worker capacity | Slow host, oversized response, too many redirects, or expensive resource loads. | Use connection/read/overall timeouts, byte and redirect caps, and queue limits; restrict worker egress. |
| Redirect reaches a private or internal address | Only the submitted URL was checked. | Validate each redirect destination and resolved address at connection time; reject disallowed ranges and restrict network egress. |
| Preview is stale | Cache refresh interval is too long or a failed refresh replaced good data. | Use a controlled refresh schedule, retain the last known good preview on failures, and expose a manual refresh policy if needed. |
| Bengali text displays as boxes or clips | Fallback font lacks glyphs or truncation/layout assumes Latin text widths. | Test actual devices and fonts, allow wrapping, and check long titles and conjunct characters. |
| Screenshot is blank or blocked | Bot check, CAPTCHA, slow rendering, or site-specific browser behavior. | Keep a metadata or domain fallback, set suitable waits and bounded retries, and do not treat a screenshot as guaranteed coverage. |
10. Performance, reliability, and cost
- Keep remote fetches off page render: a queue prevents a slow publisher from holding up the directory response.
- Reuse results: cache by normalized URL, refresh on a schedule, and retain good cached data after transient errors.
- Bound work: cap concurrency, redirects, time, response size, and screenshot frequency. These controls protect both reliability and cost.
- Separate metadata and image caching: metadata refreshes and image lifetimes may have different needs. Track failures without exposing internal diagnostics.
- Measure the workload you have: compare useful-card rate and time to ready across representative target sites. No benchmark in this research establishes Bengali-site coverage or a universally cheaper method.
FAQ
Should every directory entry get a screenshot?
Usually not as the first step. Start with metadata and reserve screenshot rendering for entries where a visual capture improves the card.
Can Open Graph tags be trusted to exist?
No. Use them as the first metadata source, with document-title and placeholder fallbacks.
Should I store the submitted URL as the card destination?
Keep the intended destination visible and apply your product’s URL policy. Metadata such as a canonical URL can inform identity, but should not silently replace the user-submitted destination.
Does Bengali require a different preview schema?
No special preview schema is established here. Preserve Unicode text, declare document and passage languages, and test fonts and layouts on target devices.


