ScreenshotNeo

BlogHow-to

How to Detect Text and Layout Changes in Website Screenshots with OCR

Combine repeatable browser captures, pixel diffs, OCR text, and word-box geometry to catch meaningful website changes while filtering capture noise.

By the ScreenshotNeo team4 October 202611 min read

To detect text and layout changes in website screenshots, capture the same page in a controlled browser environment, compare each image with a reviewed baseline, and combine three signals: a pixel diff for broad visual changes, OCR text comparison for changed or missing copy, and OCR bounding boxes for text that moved or changed size. OCR alone cannot tell you whether the overall page changed, while a pixel diff cannot explain which words changed. Using both is a practical workflow inferred from the capabilities of browser screenshot comparison and OCR tools.

The workflow below uses Playwright for repeatable captures and visual comparison, then Tesseract for OCR text and geometry. It keeps the images and OCR output available for review so capture noise or recognition errors do not silently become accepted regressions.

1. Make captures reproducible

Visual comparison is meaningful only when the baseline and current capture use comparable conditions. Browser rendering can vary with the operating system, browser version, settings, hardware, power source, and headless mode. Use the same CI image or host environment for creating and checking baselines where practical.

  • Pin the browser version and run baseline updates in the same environment as regular checks.
  • Fix the viewport dimensions and device scale factor. Keep screenshot dimensions identical before comparing OCR coordinates.
  • Load the same fonts and wait for them to finish loading.
  • Put the page into a known state: use stable test data, disable or freeze rotating content, and avoid time-dependent content where possible.
  • Wait for the specific page state your test needs. “Network idle” alone may be unsuitable for pages with persistent requests.
  • Mask, hide, or restyle irrelevant volatile regions, such as timestamps, rotating ads, or animations.

Playwright documents that screenshots may vary across environments and supports stylesheets for changing or hiding elements during capture. Keep such exclusions narrow: hiding a large region can conceal a real regression.

2. Capture a reviewed baseline with Playwright

Playwright Test can create a reference screenshot on the first run of a screenshot assertion and compare later captures against it. Commit the reviewed reference files. Treat a baseline update as a code review decision: the new image is not correct just because it was generated by the test.

Install Playwright Test and its browser binaries using the official installation guide. Add a test such as the following as tests/page-visual.spec.ts:

import { test, expect } from '@playwright/test';

test('homepage matches the approved visual baseline', async ({ page }) => {
  await page.setViewportSize({ width: 1440, height: 1000 });
  await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
  await page.evaluate(() => document.fonts.ready);

  // Hide only content known to vary and irrelevant to this check.
  await page.addStyleTag({ content: `
    .rotating-ad, .last-updated { visibility: hidden !important; }
    *, *::before, *::after {
      animation-duration: 0s !important;
      transition-duration: 0s !important;
      caret-color: transparent !important;
    }
  ` });

  await expect(page).toHaveScreenshot('homepage.png', {
    fullPage: true,
    animations: 'disabled',
    maxDiffPixelRatio: 0.001,
  });
});

Replace the URL and volatile selectors with values from your application. The threshold above is an example configuration, not a universal tolerance. Playwright also offers a per-pixel perceptual threshold and a maximum differing-pixel count; choose one or both based on representative changes and noise in your environment. A permissive threshold can hide small but meaningful changes. Inspect the generated diff when an assertion fails.

Useful capture and assertion choices include:

  • fullPage: capture the full scrollable page when below-the-fold content matters. Viewport-only captures are faster and localize changes to the visible area.
  • animations: disable finite animations for stable output. Prefer controlling application state for continuously changing content.
  • mask and maskColor: cover specific volatile elements while retaining their position and area.
  • stylePath: apply a stylesheet to hide or restyle elements during screenshot capture.
  • maxDiffPixels and maxDiffPixelRatio: bound tolerated changed pixels. Tune them against your baseline dimensions and actual expected variation.

See the Playwright screenshot assertions documentation for supported assertion options and baseline behavior.

3. Install Tesseract and extract text with coordinates

Tesseract can run locally and emits plain text as well as structured formats. TSV output includes word-level bounding boxes, confidence, and text; hOCR can also represent geometry and confidence. Install the Tesseract executable and its language data for the languages present in your screenshots, then produce TSV for each image:

tesseract baseline.png baseline -l eng --psm 3 tsv
tesseract current.png current -l eng --psm 3 tsv

These commands write baseline.tsv and current.tsv. The --psm 3 setting treats the input as a page with automatic segmentation. For a tightly cropped button, banner, or other small region, use a segmentation mode suited to that crop; Tesseract’s quality guidance notes that its default page assumption may not fit small regions. Skew can also reduce line segmentation quality. See the Tesseract image quality guidance and command-line usage.

A compact Python script can compare normalized word sequences and approximate word-box movement from the two TSV files. Save as compare_ocr.py:

import csv
import re
import sys
from pathlib import Path


def read_words(path):
    rows = []
    with Path(path).open(newline='', encoding='utf-8') as f:
        for row in csv.DictReader(f, delimiter='\t'):
            text = row.get('text', '').strip()
            if not text:
                continue
            try:
                confidence = float(row['conf'])
                box = tuple(int(row[k]) for k in ('left', 'top', 'width', 'height'))
            except (KeyError, ValueError):
                continue
            rows.append({'text': text, 'conf': confidence, 'box': box})
    return rows


def normalize(text):
    # Keep case and punctuation by default; relax only if those changes
    # do not matter to your application.
    return re.sub(r'\s+', ' ', text).strip()


if len(sys.argv) != 3:
    raise SystemExit('usage: python compare_ocr.py baseline.tsv current.tsv')

before = read_words(sys.argv[1])
after = read_words(sys.argv[2])
before_text = [normalize(w['text']) for w in before]
after_text = [normalize(w['text']) for w in after]

if before_text != after_text:
    print('OCR word sequence changed')
    print('BASELINE:', ' '.join(before_text))
    print('CURRENT: ', ' '.join(after_text))
else:
    print('OCR word sequence unchanged')

# This positional comparison assumes the OCR word order is similar in both
# captures. For large insertions/deletions, align words or lines first.
for i, (old, new) in enumerate(zip(before, after)):
    if normalize(old['text']) != normalize(new['text']):
        continue
    ob, nb = old['box'], new['box']
    delta = tuple(n - o for o, n in zip(ob, nb))
    if any(abs(v) > 3 for v in delta):
        print(f'BOX MOVED word={old["text"]!r} index={i} delta={delta}')

if len(before) != len(after):
    print(f'WORD COUNT changed: {len(before)} -> {len(after)}')

Run it with:

python compare_ocr.py baseline.tsv current.tsv

This is a minimal comparison example, not a complete word-alignment algorithm. It reports sequence changes and flags boxes at the same word index that shifted more than three pixels. For robust layout checks, align corresponding lines or regions before comparing boxes, and normalize coordinates by image width and height if dimensions can differ. Keep confidence scores: a low-confidence recognition difference should lead to visual inspection of the affected crop, not automatic acceptance or rejection.

4. Compare text and geometry in a useful way

Text and layout are related but distinct checks. Decide which changes matter for the page and compare at the appropriate level:

Signal Detects well Can miss or misclassify
Pixel diff Broad visual changes, color shifts, missing regions, and rendering differences May flag harmless antialiasing or environment noise; does not explain changed words
OCR text Added, deleted, or changed visible copy Recognition mistakes, reading-order changes, and layout movement without text edits
OCR boxes Text movement and size changes, when corresponding regions are matched Unstable coordinates if dimensions or scale differ; ambiguous matching after insertions
  1. Compare text by stable regions or lines when possible, not only as one page-wide string. This makes an addition near the top less likely to shift every subsequent word pairing.
  2. Report additions, deletions, and replacements. Normalize whitespace only when whitespace itself is not under test.
  3. Retain OCR confidence and link each text or geometry alert to the corresponding screenshot crop.
  4. Compare box coordinates only after ensuring equal screenshot dimensions and scale, or convert positions to fractions of image width and height.
  5. Review the original image, the pixel diff, and OCR output together before classifying a change.

The positional Python sample is suitable for a stable page with similar word order. For pages where copy may be inserted or deleted, use line or region matching first, then compare boxes for matched content. OCR coordinates are measurements from an image; they do not identify the DOM element that produced the text.

5. Choose an OCR path based on your constraints

There is no universal winner for website screenshot OCR established by the official documentation cited here. Evaluate candidate tools against representative screenshots from your own pages, fonts, languages, and contrast conditions.

Option Deployment and output Considerations
Tesseract Local OCR; text, TSV word coordinates and confidence, hOCR geometry Useful when images should remain local; recognition depends on image quality, language data, and segmentation settings
Google Cloud Vision Hosted OCR; text detection and document text detection with page, block, paragraph, word, and break structure Consider data handling, latency, quotas, and service cost; assess on screenshots, not just documents
Amazon Textract Hosted text detection and document analysis with layout blocks Its documented focus is document analysis; test website screenshot behavior and visually confirm consequential low-confidence detections

Official references: Google Cloud Vision OCR, Amazon Textract overview, and Textract best practices. Compare local versus hosted processing, required languages, hierarchy and geometry outputs, confidence reporting, privacy requirements, latency, quotas, and cost. Do not infer comparative accuracy from feature descriptions.

6. Run it in CI and review changes

  1. Run the browser capture in a pinned environment with stable test data.
  2. Generate the screenshot assertion result and OCR TSV for both the approved baseline and current capture.
  3. Fail or flag the job according to your policy, but retain the baseline, current image, pixel diff, OCR output, and affected crops as artifacts.
  4. Have a reviewer decide whether a difference is a regression, intended change, capture noise, or OCR error.
  5. Update the baseline only after the new appearance is approved. Keep the baseline update in the same reviewable change as the UI change where possible.

Capture less than the full page when only a specific component matters; smaller images usually reduce capture, storage, and OCR work. Full-page captures provide broader coverage but can be slower and more sensitive to content loaded lower on the page. Hosted OCR adds network and service dependencies; local OCR avoids sending screenshots to a service but requires managing the executable and language data. The dossier does not establish universal latency, accuracy, or pricing comparisons, so measure these against your own image volume and operational requirements.

7. Troubleshooting

Symptom Likely cause Fix
Nearly every run has pixel differences Different browser or host environment, fonts, scale, animation, or dynamic page state Pin the environment, wait for fonts and stable state, disable animations, and mask only known volatile regions
Small text edits do not fail the screenshot assertion Pixel tolerance is too permissive or the changed area is too small relative to the threshold Review threshold settings and add the OCR text comparison; validate against known small edits
OCR reports missing or garbled words Low contrast, small text, unsuitable segmentation, skew, or missing language data Inspect the crop, improve image quality, select appropriate language data and page segmentation, and retain confidence values
Every box after an inserted sentence appears moved Word-by-word positional matching lost alignment after an insertion or deletion Match lines or regions first, then compare boxes within matched units
Boxes shift even though the page looks unchanged Image dimensions or device scale differ, or OCR segmentation changed Keep dimensions and scale identical; normalize coordinates and inspect the source image
Baseline updates keep hiding regressions Reference images are refreshed without visual review Require a reviewer to inspect the new screenshot and diff before approving an update
Tests hang waiting for network idle The site maintains long-running requests or background polling Wait for a specific selector or application-ready state, then capture

8. Or skip the browser setup

If you need repeatable captures without maintaining browser setup, ScreenshotNeo is a website screenshot API and MCP server. Its one-call API returns an image or PDF for a URL; the response headers report the page verdict and whether the capture was billed. The screenshot still needs to be passed through your chosen OCR tool for text and geometry comparison. See the ScreenshotNeo docs for API options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res); // Bun example; use your runtime's file API in Node.js

For Node.js, save the returned bytes with the built-in filesystem API:

import { writeFile } from 'node:fs/promises';
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for 1,000 free screenshots a month, with no card required.

FAQ

Can OCR detect a layout change if the words stay the same?

Text strings alone cannot. Compare OCR bounding boxes for matched text regions, and use the pixel diff to catch non-text layout changes.

Should OCR output be the only CI pass/fail signal?

Usually not. OCR can misread text, and its geometry needs reliable matching. Preserve visual artifacts and review consequential or low-confidence differences.

Does this approach work for languages other than English?

Yes, provided the OCR engine has suitable language data or language support and the chosen configuration matches the text. Validate recognition and reading order on your actual screenshots.

How should I handle a legitimate UI redesign?

Review the current capture and diff, approve the intended appearance, then update the baseline and any expected OCR text or geometry reference as part of the change.

References