ScreenshotNeo

BlogEngineering

Embedding Generated Document Previews

Build a searchable index of PDF and page previews with multimodal embeddings, OCR, metadata, citations, and reliable retrieval.

By the ScreenshotNeo team1 October 202610 min read

Embedding a generated document preview means converting each rendered page (or preview image) into a vector that represents both what the page says and what it looks like. Store that vector with the document ID, page number, revision, preview-render version, and access metadata. At query time, embed the user’s question with the matching retrieval task, search the vector index, and return the preview together with a citation to the source page.

For PDFs, multimodal embedding is useful when charts, tables, diagrams, handwriting, or layout carry meaning that plain text extraction loses. Google states that PDF embedding uses both visual and text features, while scanned PDFs go through OCR. Cohere describes the same approach as one embedding derived from textual and visual elements.

Architecture at a glance

  1. Render: create a stable image or one-page PDF preview for every page or preview state.
  2. Preserve provenance: keep the original document and attach document ID, page number, revision, source URI, access policy, and renderer version.
  3. Extract and embed: send the page PDF or image to a multimodal embedding model. For scanned pages, include OCR or use a service that performs it automatically.
  4. Index: store the vector and metadata in a vector database or managed retrieval service.
  5. Retrieve: embed a text or image query with the matching retrieval convention, run nearest-neighbor search, and return the page preview plus a source citation.
  6. Rebuild when needed: re-embed after content, layout, OCR output, or embedding-model version changes.

How to embed a PDF preview

1. Generate deterministic page previews

Render pages at a fixed scale and record the rendering inputs. A preview should not silently change because a browser, font, or rendering package was upgraded.

# Example command with Poppler installed
pdftoppm -png -r 150 input.pdf preview/page

For large documents, process one page at a time. Google’s Gemini PDF embedding workflow accepts at most six pages per file and recommends one page per PDF for best quality. A rendered PDF page consumes 258 visual tokens; the shared input limit is 8,192 tokens, so oversized inputs can be truncated. See the Gemini embedding documentation.

2. Keep metadata beside every vector

{
  "id": "report-2026-r3-p017",
  "document_id": "report-2026",
  "page_number": 17,
  "revision": 3,
  "preview_uri": "s3://previews/report-2026/r3/page-017.png",
  "source_uri": "s3://documents/report-2026/r3/original.pdf",
  "preview_renderer": "poppler-24.02-150dpi",
  "ocr_quality": 0.96,
  "embedding_model": "MODEL_VERSION",
  "access_policy": "tenant-42"
}

Store the source citation fields even if your vector database supports only a small metadata schema. You need them to explain why a result matched and to enforce authorization before displaying it.

3. Embed documents and queries with consistent task formatting

For asymmetric retrieval, use a document task when indexing and a query task when searching. Google’s example formats a query as task: search result | query: ... and a document as title: ... | text: .... Keep this convention identical across all pages and queries. Inconsistent task instructions can reduce retrieval quality.

The following program is complete apart from the provider-specific embedding call. Set EMBEDDING_ENDPOINT to your multimodal embedding endpoint and provide the request shape required by that endpoint. The rest of the pipeline writes a JSONL index, preserves metadata, and performs cosine search without a separate database.

import base64
import json
import math
import os
from pathlib import Path
from typing import Any

import requests

ENDPOINT = os.environ["EMBEDDING_ENDPOINT"]
API_KEY = os.environ["EMBEDDING_API_KEY"]
MODEL = os.environ.get("EMBEDDING_MODEL", "MODEL_VERSION")
INDEX_PATH = Path("preview-index.jsonl")


def embed(content: dict[str, Any], task: str) -> list[float]:
    """Call your multimodal provider and return its float vector.

    Adapt only the JSON field names to your provider. The content contains
    both the rendered page image and extracted/OCR text when available.
    """
    payload = {"model": MODEL, "task": task, "content": content}
    response = requests.post(
        ENDPOINT,
        headers={"Authorization": f"Bearer {API_KEY}"},
        json=payload,
        timeout=90,
    )
    response.raise_for_status()
    body = response.json()
    return body["embedding"]


def cosine(a: list[float], b: list[float]) -> float:
    dot = sum(x * y for x, y in zip(a, b))
    na = math.sqrt(sum(x * x for x in a))
    nb = math.sqrt(sum(y * y for y in b))
    return dot / (na * nb) if na and nb else 0.0


def page_content(image_path: Path, text: str) -> dict[str, Any]:
    encoded = base64.b64encode(image_path.read_bytes()).decode("ascii")
    return {
        "image_base64": encoded,
        "mime_type": "image/png",
        "text": text,
    }


def index_page(image_path: Path, text: str, metadata: dict[str, Any]) -> None:
    vector = embed(
        page_content(image_path, text),
        task="retrieval_document",
    )
    record = {"vector": vector, "metadata": metadata}
    with INDEX_PATH.open("a", encoding="utf-8") as handle:
        handle.write(json.dumps(record) + "\n")


def search(question: str, limit: int = 5) -> list[dict[str, Any]]:
    query_vector = embed({"text": question}, task="retrieval_query")
    hits = []
    with INDEX_PATH.open(encoding="utf-8") as handle:
        for line in handle:
            record = json.loads(line)
            hits.append((cosine(query_vector, record["vector"]), record["metadata"]))
    hits.sort(key=lambda item: item[0], reverse=True)
    return [{"score": score, **metadata} for score, metadata in hits[:limit]]


if __name__ == "__main__":
    index_page(
        Path("preview/page-017.png"),
        "Revenue by region; the table shows EMEA growing faster than North America.",
        {
            "id": "report-2026-r3-p017",
            "document_id": "report-2026",
            "page_number": 17,
            "revision": 3,
            "source_uri": "s3://documents/report-2026/r3/original.pdf",
            "preview_uri": "s3://previews/report-2026/r3/page-017.png",
            "embedding_model": MODEL,
            "access_policy": "tenant-42",
        },
    )
    print(json.dumps(search("Which region grew fastest?"), indent=2))

Install the only dependency with python -m pip install requests. If your provider accepts a PDF directly, send a one-page PDF instead of image_base64. If it accepts images and text separately, keep both fields so the model can use layout and extracted text.

cURL: send a preview to an embedding endpoint

curl -X POST "$EMBEDDING_ENDPOINT" \
  -H "Authorization: Bearer $EMBEDDING_API_KEY" \
  -H "Content-Type: application/json" \
  --data-binary @- <<'JSON'
{
  "model": "MODEL_VERSION",
  "task": "retrieval_document",
  "content": {
    "image_base64": "BASE64_PAGE_IMAGE",
    "mime_type": "image/png",
    "text": "OCR or extracted text for this page"
  }
}
JSON

Replace the endpoint, model, and response-field names with those documented by your provider. Keep the task value and query/document convention stable for every request.

Node.js: index a preview

import { readFile } from "node:fs/promises";

const endpoint = process.env.EMBEDDING_ENDPOINT;
const apiKey = process.env.EMBEDDING_API_KEY;
const image = (await readFile("preview/page-017.png")).toString("base64");

const response = await fetch(endpoint, {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${apiKey}`,
    "Content-Type": "application/json"
  },
  body: JSON.stringify({
    model: process.env.EMBEDDING_MODEL || "MODEL_VERSION",
    task: "retrieval_document",
    content: {
      image_base64: image,
      mime_type: "image/png",
      text: "OCR or extracted text for this page"
    }
  })
});

if (!response.ok) {
  throw new Error(`${response.status}: ${await response.text()}`);
}

const body = await response.json();
console.log(JSON.stringify({
  vector: body.embedding,
  metadata: {
    document_id: "report-2026",
    page_number: 17,
    revision: 3,
    preview_renderer: "poppler-24.02-150dpi"
  }
}, null, 2));

Should you embed each page or the whole PDF?

Unit Use it when Trade-off
One page Users need precise citations, tables, diagrams, or page previews More vectors and indexing requests, but clearer results
Small page group A concept routinely spans adjacent pages Fewer vectors, less precise citations
Whole PDF The provider supports it and documents are short May exceed page or token limits and can blur page-level relevance

For Gemini’s documented PDF workflow, use one page per PDF when quality matters. A six-page maximum applies per file, and 8,192 shared input tokens constrain large pages or long extracted text. Keep page vectors even when you also create document-level summaries so you can return an auditable citation.

Scanned PDFs, OCR, and layout quality

Scanned pages contain pixels rather than selectable text. Google’s Gemini Developer API automatically enables OCR for PDF inputs. If you need explicit control, Document AI Enterprise OCR can return blocks, paragraphs, lines, words, symbols, and page numbers, with rotation correction and image-quality scores. Use OCR confidence or quality metadata to reprocess, flag, or exclude poor pages. See Document AI Enterprise OCR.

  • Keep the original page image even after OCR so visual evidence remains available.
  • Store OCR language, confidence, rotation, and preprocessing version.
  • Do not discard tables or diagrams just because OCR extracted little text; their visual structure may still be the strongest signal.
  • Re-embed after correcting rotation, replacing a low-resolution scan, or changing OCR output.

Query formatting and result citations

Use the same embedding space and retrieval task family for documents and queries. A typical convention is title: ... | text: ... for an indexed page and task: search result | query: ... for a user question. Return the top page IDs, scores, preview URI, source URI, and page number to the application. Enforce access policy before displaying a preview or source link.

Google’s managed File Search can handle storage, chunking, embeddings, vector search, and citation injection for supported workflows. It returns citations identifying the document passages used. See the Google File Search announcement.

Choosing a model or service

Option Relevant capability Questions to verify
Gemini Embedding 2 / Gemini API Direct PDF input, visual and text processing, automatic OCR for scanned PDFs, task instructions, adjustable dimensions, and vector-store integrations Page, token, quota, retention, and regional limits
Cohere Embed v4 One multimodal embedding from textual and visual PDF elements Accepted formats, dimensions, pricing, and storage policy
Gemini File Search Managed storage, chunking, embeddings, vector search, broad file-format support, and citations Supported formats, lifecycle controls, quotas, and data residency
Document AI Enterprise OCR Explicit OCR preprocessing, layout structure, rotation correction, and image-quality signals OCR cost, latency, region, and retention

Compare vendors on visual-plus-text support, OCR and layout fidelity, page/file/token limits, output dimensions, index cost, task instructions, metadata and citation support, data residency, retention, quotas, and operational pricing. Gemini Embedding 2 has a documented default of 3,072 dimensions with adjustable output dimensions; reducing dimensions can lower index storage, but measure retrieval quality on your corpus before changing them. See Google Cloud’s multimodal embeddings documentation.

Or skip the browser setup

If your source is a live document viewer or web page, ScreenshotNeo can generate the preview before you embed it. One GET request returns PNG, JPEG, WebP, or PDF; it can capture a full page or a CSS-selected element, load lazy images, apply a device preset or custom viewport, use a retina scale, and run custom CSS or JavaScript. Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.

See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. Every feature is on every plan: 1,000 shots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Performance, reliability, and cost

  • Batch work: render and embed pages asynchronously, with bounded concurrency and retries for transient provider errors.
  • Cache by content: hash the source bytes, renderer settings, OCR output, and model version. Reuse a vector when the hash is unchanged.
  • Control dimensions: vector dimensions affect index storage and nearest-neighbor cost. Record the chosen dimension with every vector.
  • Respect limits: split PDFs before the provider’s page and token ceilings. Reject or downscale oversized pages deliberately instead of allowing silent truncation.
  • Make writes idempotent: key records by document ID, revision, page number, renderer version, and model version. Upsert on retries.
  • Monitor quality: track OCR confidence, empty-page rates, embedding failures, query latency, and citation accuracy on a small labeled test set.
  • Protect data: apply tenant filters before vector search results reach the user, and verify provider retention and residency requirements.

Troubleshooting

Symptom Likely cause Fix
Results ignore charts or tables Only extracted text was embedded Send the rendered page image or PDF together with text.
Scanned pages return weak matches OCR failed, rotated, or produced low-confidence text Run OCR with rotation correction, store quality signals, and reprocess low-quality pages.
Requests fail on long PDFs Page or shared token limit exceeded Render one page per input or split into smaller files; keep page metadata.
Relevant pages rank below unrelated pages Query and document task formatting differs Use one documented convention consistently and re-embed affected vectors.
Duplicate or stale results appear Old renderer, revision, or model vectors remain active Include version fields in the key and delete or filter superseded records.
Unauthorized previews are returned Access policy was omitted from the filter Apply tenant or document authorization before returning hits.
Index storage grows unexpectedly Every rerender created another vector Content-address previews and upsert by the full versioned identity.

Checklist

  • Each page has a stable preview and a reproducible renderer version.
  • Vectors include document, page, revision, model, and access metadata.
  • Images and extracted/OCR text are embedded together when layout matters.
  • Page and token limits are enforced before requests are sent.
  • Query and document task formatting is consistent.
  • Search results return a source citation and pass authorization checks.
  • OCR confidence and embedding failures are observable.
  • Re-embedding is triggered by content, layout, OCR, or model changes.

FAQ

Can embeddings understand charts and tables?

Yes, multimodal document embeddings are designed to use visual features as well as text. Preserve the rendered page so chart structure and table layout reach the model.

How do I search scanned PDFs semantically?

Run OCR as part of preprocessing or use a PDF embedding service with automatic OCR, then embed the page image and extracted text together. Track OCR quality so weak pages can be reprocessed.

Which vector database should store preview embeddings?

Use a store that supports your vector dimension, metadata filters, tenant isolation, and citation fields. Google lists Vector Search 2.0, BigQuery, AlloyDB, Cloud SQL, and third-party vector databases as options.

When should I re-embed a page?

Re-embed when the source content, page layout, OCR output, preview renderer, or embedding model changes. Keep the previous version if you need reproducible historical results.