ScreenshotNeo

BlogHow-to

How to Use an Image Search API for OCR Text Extraction

Image search APIs find similar pictures. For text extraction, use OCR with Google Vision or Azure Read, then parse text, layout, and coordinates.

By the ScreenshotNeo team1 October 20268 min read

Direct answer: an image-search API that finds visually similar images does not normally read characters. For OCR text extraction, send the image to a vision/OCR endpoint such as Google Cloud Vision OCR or Azure AI Vision Read. Choose Google TEXT_DETECTION for ordinary images with sparse text and DOCUMENT_TEXT_DETECTION for dense, scanned documents where page, block, paragraph, word, and line structure matters.

1. Image search and OCR solve different problems

Image search services return similar or matching images, metadata, or ranking results. OCR services inspect pixels and convert visible characters into a text response. If your application needs an invoice number, a sign, a screenshot label, or a scanned paragraph, call an OCR operation.

Need Operation Typical output
Find visually similar images Image search or reverse-image search Image URLs, similarity results, metadata
Read text in a photo Google TEXT_DETECTION or Azure Read Full text, words, bounding polygons
Read a dense scan or multipage document Google DOCUMENT_TEXT_DETECTION or Azure Read for PDF Pages, blocks, paragraphs, words, breaks, coordinates

Google describes Cloud Vision as providing optical character recognition capabilities for text detection from images. The endpoint is https://vision.googleapis.com/v1/images:annotate.

2. Choose the OCR mode and input source

TEXT_DETECTION for ordinary images

Use this mode for signs, screenshots, product labels, or photos with a modest amount of text. The response includes a complete detected string plus individual text annotations and bounding polygons.

DOCUMENT_TEXT_DETECTION for documents

Use document mode for dense pages, forms, receipts, and scans. Its response exposes a hierarchy you can traverse: pages, blocks, paragraphs, words, and symbol or break information. This makes it easier to rebuild reading order or map values to regions.

Cloud Storage versus a web URL

Google accepts a Cloud Storage URI such as gs://BUCKET/path/image.jpg or a web URL. A third-party URL can fail when the host denies requests or throttles traffic, so controlled Cloud Storage is safer for production. Download remote images into storage first when you need predictable access, retention, and regional control.

3. Google Cloud Vision OCR: complete request

Create or select a Google Cloud project, enable the Vision API, configure billing, and obtain an OAuth access token. The following request uses a Cloud Storage image. Replace the placeholders and choose one feature type.

curl -X POST \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -H "x-goog-user-project: YOUR_PROJECT_ID" \
  -H "Content-Type: application/json; charset=utf-8" \
  https://vision.googleapis.com/v1/images:annotate \
  -d '{
    "requests": [{
      "image": {"source": {"imageUri": "gs://YOUR_BUCKET/path/image.jpg"}},
      "features": [{"type": "TEXT_DETECTION"}]
    }]
  }'

For a dense document, change the feature to DOCUMENT_TEXT_DETECTION:

"features": [{"type": "DOCUMENT_TEXT_DETECTION"}]

The same request shape works with a remote URL:

"image": {"source": {"imageUri": "https://example.com/receipt.jpg"}}

Use a URL only when the host is dependable and permits Google’s fetchers. Otherwise copy the object to Cloud Storage and submit the gs:// URI.

4. Runnable Python example

import json
import os
import requests

project_id = os.environ["GOOGLE_CLOUD_PROJECT"]
access_token = os.environ["GOOGLE_ACCESS_TOKEN"]
image_uri = "gs://YOUR_BUCKET/path/image.jpg"

payload = {
    "requests": [{
        "image": {"source": {"imageUri": image_uri}},
        "features": [{"type": "DOCUMENT_TEXT_DETECTION"}]
    }]
}

response = requests.post(
    "https://vision.googleapis.com/v1/images:annotate",
    headers={
        "Authorization": f"Bearer {access_token}",
        "x-goog-user-project": project_id,
        "Content-Type": "application/json",
    },
    json=payload,
    timeout=60,
)
response.raise_for_status()
data = response.json()

annotation = data["responses"][0]
full_text = annotation.get("fullTextAnnotation", {}).get("text", "")
print(full_text)

# Word-level coordinates in document mode.
for page in annotation.get("fullTextAnnotation", {}).get("pages", []):
    for block in page.get("blocks", []):
        for paragraph in block.get("paragraphs", []):
            for word in paragraph.get("words", []):
                text = "".join(
                    symbol.get("text", "")
                    for symbol in word.get("symbols", [])
                )
                vertices = word.get("boundingBox", {}).get("vertices", [])
                print(text, vertices)

Install the only dependency with python -m pip install requests. In a service, keep credentials in a secret manager and return a normalized result rather than the provider’s entire response.

5. Runnable Node.js example

const projectId = process.env.GOOGLE_CLOUD_PROJECT;
const accessToken = process.env.GOOGLE_ACCESS_TOKEN;

const payload = {
  requests: [{
    image: { source: { imageUri: 'gs://YOUR_BUCKET/path/image.jpg' } },
    features: [{ type: 'TEXT_DETECTION' }]
  }]
};

const response = await fetch('https://vision.googleapis.com/v1/images:annotate', {
  method: 'POST',
  headers: {
    Authorization: `Bearer ${accessToken}`,
    'x-goog-user-project': projectId,
    'Content-Type': 'application/json'
  },
  body: JSON.stringify(payload)
});

if (!response.ok) {
  throw new Error(`${response.status}: ${await response.text()}`);
}

const result = await response.json();
const item = result.responses?.[0] ?? {};
console.log(item.textAnnotations?.[0]?.description ?? '');

for (const annotation of item.textAnnotations?.slice(1) ?? []) {
  console.log(annotation.description, annotation.boundingPoly?.vertices ?? []);
}

This uses the built-in fetch available in current Node.js releases. For document hierarchy, switch the feature type and traverse fullTextAnnotation.pages as in the Python example.

6. Parsing text, words, and bounding boxes

For simple extraction, read the first text annotation’s description (or fullTextAnnotation.text in document mode). For coordinates, iterate word annotations and read their polygon vertices. Coordinates are pixel positions in the submitted image; preserve the original dimensions if you need to draw highlights later.

  • Full text: use when storing searchable content or sending text to a parser.
  • Words and polygons: use for redaction, click targets, table extraction, or highlighting.
  • Document hierarchy: use pages and paragraphs to retain reading order and page boundaries.
  • Confidence: treat provider confidence values as signals, then validate critical fields such as totals, dates, and IDs.

Normalize line endings, trim repeated whitespace, and keep the raw provider response for debugging. Do not assume polygon vertices are always returned in the same orientation; calculate the minimum and maximum x/y values when creating a rectangle.

7. Azure AI Vision Read as an alternative

Azure AI Vision Read accepts an image or PDF and performs OCR asynchronously. Post an image URL or binary input with an Ocp-Apim-Subscription-Key, retain the operation URL returned by the service, and query that URL until the operation completes. Azure supports selecting pages or page ranges, which is useful for large PDFs. Choose it when your application already uses Azure identity, networking, monitoring, or storage.

Compare providers on authentication, URL versus storage input, synchronous versus asynchronous behavior, layout detail, page selection, regional processing, SDK support, quotas, and current price. Neither provider’s documentation in this research supplies a directly comparable accuracy percentage, so evaluate representative images from your own workload.

8. Batch and asynchronous processing

For small interactive requests, synchronous annotation is simplest. For offline archives, Google documents asynchronous batch annotation for up to 2,000 image files, with response JSON written to Cloud Storage. Batch jobs reduce request overhead but require job tracking, output cleanup, and retry handling. Azure Read is asynchronous by design for its comparable workflow.

  1. Store source images in a controlled bucket.
  2. Submit a batch or asynchronous operation and record its operation ID.
  3. Poll with exponential backoff, or consume the provider’s completion mechanism where available.
  4. Validate each result and associate it with the original object name.
  5. Move successful output to durable storage and quarantine failures for review.

9. Reliability, performance, and cost notes

  • Image preparation: crop irrelevant borders, keep text large enough to read, and avoid unnecessary recompression. Upscaling a tiny source cannot recover missing detail.
  • Network reliability: Cloud Storage avoids third-party URL denials and throttling. Set request timeouts and retry transient 5xx responses with exponential backoff.
  • Idempotency: derive a content hash and cache OCR results so a retry does not create duplicate downstream records.
  • Throughput: parallelize within the provider’s documented quotas, then back off on rate-limit responses.
  • Regional needs: Google documents global, US, and EU regional OCR endpoints. Select the endpoint and storage location required by your data policy.
  • Cost: check the current provider pricing and quota pages before launch. Measure pages or images processed, retries, storage, and egress rather than assuming one request equals one business record.

10. Troubleshooting common OCR failures

Symptom Likely cause Fix
401 or 403 Expired token, disabled API, or missing project billing Refresh credentials, enable Vision, and verify the project header and billing account.
Image cannot be fetched Remote host blocks Google or throttles requests Copy the file to controlled Cloud Storage and submit its gs:// URI.
Empty annotations Text is too small, blurred, rotated, or low contrast Use a higher-resolution source, crop the text, improve contrast, and retry with the appropriate mode.
Paragraph order is wrong Reading order is ambiguous in a complex layout Use document mode, inspect block coordinates, and sort or group regions using your document rules.
Only part of a PDF is read Page range or asynchronous output was limited Check the requested pages and enumerate every output JSON file.
429 or quota errors Concurrency exceeds the project or provider quota Reduce parallelism, add exponential backoff, and request a quota increase if appropriate.
Characters are misread Stylized fonts, glare, handwriting, or compression artifacts Preprocess the image, use field-level validation, and route uncertain records for human review.

11. Or skip the browser setup

If the image is a webpage, first capturing a clean, stable screenshot can make the OCR input more consistent. ScreenshotNeo is a website screenshot API; it is not an OCR engine, so send its returned image to your chosen vision API.

One GET request returns PNG, JPEG, WebP, or PDF. The API can wait for a selector, delay, or network idle; load lazy images; set a viewport or device preset; apply custom CSS or JavaScript; and block ads, trackers, requests, or resource types. It accepts cookie and header settings when the page requires authentication.

See the ScreenshotNeo API documentation for all options. A minimal capture looks like this:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

12. Implementation checklist

  • Confirm you need OCR rather than image similarity search.
  • Select TEXT_DETECTION for sparse text or DOCUMENT_TEXT_DETECTION for dense documents.
  • Prefer controlled Cloud Storage over unreliable third-party URLs.
  • Parse both full text and coordinates when downstream fields need verification.
  • Use asynchronous batches for archives and track every operation.
  • Add retries, quotas, caching, and regional controls before production.
  • Validate critical values against your business rules.

FAQ

Can an image-search API extract text?

Usually no. Similarity search and OCR are separate operations; use a vision/OCR endpoint for characters.

Should I upload a file or send an image URL?

Both are supported by Google, but a Cloud Storage URI is more predictable when a third-party host may block or throttle requests.

Which Google feature reads scanned documents?

DOCUMENT_TEXT_DETECTION; it exposes page and paragraph structure in addition to text.

Can OCR return where each word appears?

Yes. Read the word-level bounding polygons and map their vertices to the source image dimensions.

When should I use Azure Read?

Use it when Azure identity, storage, monitoring, regional controls, or page-range processing fit your existing system better.

Does ScreenshotNeo perform OCR?

No. It captures a clean webpage image or PDF; pass that output to Google Vision, Azure Read, or another OCR service.