ScreenshotNeo

BlogHow-to

How to Extract Ranking Positions from Google SERP Screenshots with OCR

Extract text and coordinates from Google SERP screenshots, group text into results, and assign positions using a clear counting rule.

By the ScreenshotNeo team4 October 202610 min read

To extract ranking positions from a Google SERP screenshot, run OCR that returns text with bounding boxes, group nearby title, attribution, and snippet text into candidate result blocks, decide which blocks count under a written rule, then sort eligible blocks by vertical position. OCR reads and locates text; it does not determine rankings by itself.

Keep the screenshot and its capture context with the result. A position extracted from one screenshot describes that captured page, not a universal rank: Google results can vary by device, location, query, and language. Google’s visual elements guide describes text results and their optional elements, such as attribution, title links, snippets, and sitelinks.

1. Define what position means

Before writing code, specify the measurement. A practical default for organic text rankings is:

  • Count eligible organic text-result blocks in the main results column from top to bottom.
  • Exclude paid ads, local/map results, image and video modules, featured answers, and other non-text modules.
  • Count a result with sitelinks as one result, not one result per sitelink.
  • Record excluded or ambiguous blocks separately rather than silently treating them as organic results.

There is no single counting convention imposed by OCR or by the visual layout. If your reporting convention includes a featured result or another module, document that and use it consistently. Google documents multiple visual result types and notes that text-result elements can vary.

2. Preserve screenshot context

Store enough information to reproduce or audit each extraction:

  • Screenshot file and a stable identifier or checksum.
  • Search query, capture timestamp, device or viewport dimensions, and screenshot scale.
  • Country or location, interface/query language, and known session or personalization context.
  • OCR engine and version, language model, preprocessing steps, and counting-rule version.

Do not compare ranks as if they came from the same conditions when the device, location, or query context differs. Google says its results can be displayed differently based on these factors.

3. Choose OCR output with geometry

A plain text dump loses layout. Request word boxes or structured layout so your code can reconstruct lines and blocks. Google Cloud Vision supports TEXT_DETECTION, which returns detected text and bounding boxes, and DOCUMENT_TEXT_DETECTION, which returns page, block, paragraph, word, and break structure. See the Cloud Vision OCR documentation. The documentation does not establish which feature is more accurate on SERP screenshots; compare both on a labeled sample if choosing between them.

Tesseract is an open-source local OCR engine with command-line and API options. It can generate TSV output with word boxes. Its documentation does not establish SERP-specific comparative accuracy. Choose based on your needs for local processing, supported languages, output geometry, integration effort, and results on your own labeled screenshots.

4. Runnable local example with Tesseract and Python

This example runs Tesseract’s TSV output, groups words into lines, and prints lines ordered from top to bottom. It does not claim to classify every SERP module automatically: use its output to review candidate blocks and apply your documented inclusion rule. Install Tesseract and its language data, then install Pillow:

python -m pip install Pillow
# Install the Tesseract executable using your operating system's package manager.
# Confirm it is available with: tesseract --version

Save as serp_ocr.py and run python serp_ocr.py screenshot.png:

import csv
import io
import subprocess
import sys
from collections import defaultdict
from PIL import Image

if len(sys.argv) != 2:
    raise SystemExit("Usage: python serp_ocr.py screenshot.png")

image_path = sys.argv[1]
# Tesseract TSV includes one box per recognized word. Change eng if needed.
proc = subprocess.run(
    ["tesseract", image_path, "stdout", "-l", "eng", "tsv"],
    check=True, capture_output=True, text=True,
)

# Group words by Tesseract's page/block/paragraph/line identifiers.
lines = defaultdict(list)
reader = csv.DictReader(io.StringIO(proc.stdout), delimiter="\t")
for row in reader:
    if row["level"] != "5" or not row["text"].strip():
        continue
    try:
        confidence = float(row["conf"])
        left, top = int(row["left"]), int(row["top"])
        width, height = int(row["width"]), int(row["height"])
    except (ValueError, TypeError):
        continue
    key = (row["page_num"], row["block_num"], row["par_num"], row["line_num"])
    lines[key].append({
        "text": row["text"], "left": left, "top": top,
        "right": left + width, "bottom": top + height,
        "confidence": confidence,
    })

with Image.open(image_path) as im:
    image_width, image_height = im.size

output = []
for words in lines.values():
    words.sort(key=lambda w: w["left"])
    text = " ".join(w["text"] for w in words)
    output.append({
        "top": min(w["top"] for w in words),
        "left": min(w["left"] for w in words),
        "bottom": max(w["bottom"] for w in words),
        "text": text,
        "mean_confidence": round(sum(w["confidence"] for w in words) / len(words), 1),
    })

for line in sorted(output, key=lambda item: (item["top"], item["left"])):
    y = line["top"]
    normalized_y = y / image_height
    print(f'y={y:4d} ({normalized_y:.3f}) '
          f'conf={line["mean_confidence"]:5.1f}  {line["text"]}')

The printed coordinate is the top of each recognized line. Use the boxes and vertical spacing to group lines into candidate results. OCR confidence is a tool-specific signal, not a calibrated probability that a result or rank is correct.

5. Turn OCR lines into result blocks

  1. Normalize coordinates. Keep pixel coordinates alongside normalized x/y coordinates (coordinate divided by image width/height) if screenshots have different dimensions. This makes thresholds easier to transfer between sizes.
  2. Rebuild lines and columns. Group words with similar vertical centers and nearby horizontal positions. Identify the main result column; do not mix it with side panels or separate columns.
  3. Find candidate result starts. Look for title-like lines, attribution such as a site name or visible domain, and a snippet below. Use proximity, alignment, and whitespace as cues. These cues are implementation heuristics, not a Google-published OCR classifier.
  4. Attach optional child content. Sitelinks, rich attributes, dates, or a result image may belong to the preceding result. They do not automatically represent additional ranking positions.
  5. Mark page modules. Ads, maps, featured answers, image packs, and other modules can interrupt the vertical flow. Label them explicitly and apply the chosen rule.
  6. Assign rank only after filtering. Sort the eligible result blocks by their top coordinate in the main column and number them consecutively. Keep the original y-coordinate, block decision, and rule version with each assigned position.

For a reproducible pipeline, separate OCR output, block grouping, module labels, and rank assignment into distinct data fields. That way a changed counting rule does not require rerunning OCR, and reviewers can see why a block was included.

6. Cloud Vision option

For a hosted OCR workflow, send the image to Cloud Vision using either text or document text detection and consume the returned geometry. The following is the documented request shape; replace the image URI with an image accessible to your project and configure Google Cloud authentication for the request.

POST https://vision.googleapis.com/v1/images:annotate
Content-Type: application/json

{
  "requests": [
    {
      "image": { "source": { "imageUri": "gs://YOUR_BUCKET/serp.png" } },
      "features": [ { "type": "DOCUMENT_TEXT_DETECTION" } ]
    }
  ]
}

Use TEXT_DETECTION when you want general image text detection and its word boxes; use DOCUMENT_TEXT_DETECTION when the structural hierarchy of dense text is useful. Treat this as a choice to evaluate, not a guarantee that one mode will yield better SERP rankings. Follow the current API setup and authentication instructions for a working project and credentials.

7. Validate the extracted positions

Create a representative set of screenshots and have a person label the candidate result blocks, module types, and positions under your rule. Include different viewport sizes, languages, and page layouts that occur in your actual workflow. Compare the automated output with those labels and report the sample size and error definition alongside any accuracy claim. The reviewed sources provide no published benchmark for OCR accuracy specifically on Google SERP screenshots.

Track separate error types: missed or misread text, incorrect line grouping, incorrect result-block grouping, module misclassification, and position-counting errors. This tells you whether to adjust image preprocessing, OCR choice, grouping logic, or the counting definition. Send low-confidence text and ambiguous modules to manual review instead of silently emitting a certain-looking rank.

8. Edge cases to handle

  • Mobile layouts: narrower columns cause more line wrapping and change vertical placement. Calibrate grouping for the actual viewport.
  • Low resolution or scaling: small text can disappear or merge. Capture at a readable size and preserve the original image; do not upscale and assume that lost detail returns.
  • Multi-column or side content: vertical sorting across the entire image can interleave unrelated content. Segment columns first.
  • Sitelinks and rich details: child elements may resemble separate results. Associate them with the parent block where visual structure supports it.
  • Repeated domains: a domain appearing in a title, attribution, or snippet is not enough to identify a new result. Require a candidate result start and block boundary.
  • Clipped or partial results: if a screenshot begins or ends mid-result, label the block partial and define whether it is excluded.
  • Localized text: choose matching OCR language data and validate title and domain recognition on those screenshots.
  • Overlapping boxes: retain ambiguity and request review; do not resolve overlaps solely by sorting top coordinates.

9. Troubleshooting

Symptom Likely cause Fix
No OCR output or command not found Tesseract executable is missing or not on PATH. Install Tesseract, reopen the shell, and verify tesseract --version.
Wrong characters or missing words Image is small, compressed, blurry, or uses an unsupported language model. Use the original capture at higher readable resolution and install/select the correct language data.
Text order looks scrambled OCR reading order is being mistaken for page layout; multiple columns or modules are mixed. Use bounding boxes, partition the screenshot into columns/modules, and sort within the main result column.
One result counted as several Sitelinks, snippets, or rich attributes were treated as separate blocks. Group by proximity and alignment; count parent result blocks according to the documented rule.
Position shifts between runs Capture context or counting rules changed, or borderline blocks were grouped differently. Persist viewport, location/language context, rule version, OCR output, and ambiguous-block decisions.
Cloud request returns an error Credentials, project setup, image access, or request format is invalid. Check the current Cloud Vision authentication setup, confirm the image is accessible to the project, and inspect the API error response.

10. Performance, reliability, and cost

OCR work scales with the number and size of images and the selected engine/service. Process screenshots in batches where your chosen tool supports it, avoid rerunning OCR when only the counting rule changes, and cache intermediate OCR output using an image checksum plus engine/configuration version. Hosted processing adds network and service dependencies; local processing keeps the workflow on your machine but requires installing and maintaining the engine and language data. Evaluate total processing time, failure rate, manual review load, and data-handling requirements on your own sample.

Do not use OCR output alone as a reliable rank record. Preserve source images and intermediate boxes, make extraction idempotent, and expose uncertain cases for review. Costs depend on the OCR option and its current terms; check the provider’s current pricing rather than extrapolating from an unmeasured workload.

11. Screenshot capture with ScreenshotNeo

If you need consistent source screenshots, ScreenshotNeo is a website screenshot API and MCP server. It can return PNG, JPEG, WebP, or PDF from one request. Capture context such as viewport, locale, and query still needs to be recorded by your workflow, and OCR plus position classification remains your responsibility.

Or skip the browser setup

Capture a screenshot, then pass the image to your OCR pipeline. The one-call API example below uses the documented endpoint; see the ScreenshotNeo API documentation for options and parameters.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server provides screenshot tools for AI agents, and 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000 screenshots.

Start with 1,000 free screenshots a month—no card required.

FAQ

Can OCR identify Google positions without custom logic?

No. OCR supplies text and geometry. Your pipeline must decide which text belongs to each result and which result types count.

Should ads count as ranking positions?

That depends on the metric. For organic ranking, exclude ads and state that convention; for a page-order inventory, label ads separately rather than mixing them into organic positions.

Which OCR engine is most accurate for SERP screenshots?

The cited documentation does not establish a winner for this task. Compare engines on manually labeled screenshots matching your devices, languages, and page layouts.

Does a rank from a screenshot represent a universal Google position?

No. It is an observation from a particular query and capture context, and results may vary by location, language, and device.

Sources