ScreenshotNeo

BlogHow-to

How to Extract Text From Images and Video

Use OCR to turn text in images into searchable text, and video text detection to capture visible words with timestamps. Here are local and cloud workflows.

By the ScreenshotNeo team4 October 202610 min read

Use optical character recognition (OCR) to extract visible text from a still image. For video, use video text detection to recognize words displayed in frames and retain their timestamps. OCR reads text in the picture; it does not transcribe spoken dialogue. For speech, use a speech recognition system.

For a quick local image workflow, install Tesseract and run tesseract input.png output -l eng. It writes recognized text to output.txt. For dense pages, PDFs, or video, choose a document or video OCR workflow that fits the input and preserves the structure or timing you need.

1. Choose an OCR workflow

Input and goal Good starting point What you get
One photograph, sign, or screenshot Local Tesseract or a general image OCR API Recognized text; some services also return word boxes
Dense scanned page or document image Document OCR such as Google Cloud Vision DOCUMENT_TEXT_DETECTION or Microsoft Document Intelligence Read Text organized into pages, blocks, lines, or words, often with locations
PDF or TIFF document A document OCR service, or convert pages to images before local Tesseract Text, potentially linked to page and layout
Video titles, signs, or burned-in subtitles Video OCR, such as Google Video Intelligence TEXT_DETECTION Recognized visible text associated with frame location and timestamp
Spoken dialogue Speech-to-text transcription Words from audio, usually with timing; this is not OCR

Google distinguishes TEXT_DETECTION, intended for text within general images, from DOCUMENT_TEXT_DETECTION, whose response is structured for dense documents. The latter includes page, block, paragraph, word, and break information. See Google Cloud Vision OCR documentation. Microsoft similarly describes Azure image OCR for general images and Document Intelligence Read for text-heavy scanned or digital documents in its Read model documentation.

2. Extract text from an image locally with Tesseract

Install Tesseract and language data

Install the Tesseract executable using the package manager or installer for your operating system, then install the trained language data you need. Confirm it is available on your command path with tesseract --version and check installed languages with tesseract --list-langs. Installation steps differ by operating system; use the package source you trust and ensure the executable and language files are installed together.

Use a clear, supported image such as PNG or JPEG. Tesseract’s documentation also lists TIFF, GIF, WebP, BMP, and PNM among supported formats, subject to the relevant image libraries. It does not read PDF input directly; convert PDF pages to images or use a PDF OCR tool. See the Tesseract input formats guide.

Run OCR and save plain text

tesseract input.png output -l eng
cat output.txt

The final argument is an output base name, so the default text output is output.txt. Replace eng with the language code for the installed model. To recognize multiple installed languages, combine codes with a plus sign, for example -l eng+deu. Check language availability with tesseract --list-langs. Incorrect or missing language data can produce poor output or an error.

Choose an output format

  • Plain text: use the basic command above when you only need the recognized words.
  • TSV: request structured output with tesseract input.png output -l eng tsv. This is useful when you need word-level boxes and confidence values for review or downstream processing.
  • hOCR: request hocr output to retain HTML-based layout and position information.
  • Searchable PDF: request the pdf config to create a PDF with a searchable text layer: tesseract input.png output -l eng pdf.

Confirm the available options for your installed version with tesseract --help. Tesseract’s command-line guide documents the output formats and invocation patterns.

Python: run Tesseract through pytesseract

This example uses the Tesseract executable and the Python wrapper. Install both Tesseract with the required language data and the Python packages first:

python -m pip install pytesseract pillow
from PIL import Image
import pytesseract

image = Image.open("input.png")
text = pytesseract.image_to_string(image, lang="eng")
with open("output.txt", "w", encoding="utf-8") as output:
    output.write(text)
print(text)

If the executable is not on PATH, configure pytesseract.pytesseract.tesseract_cmd with its installed path. The lang value must match language data installed for Tesseract.

Node.js: run the Tesseract command

With Tesseract installed and available on PATH, Node.js can invoke the command line and read the generated text file:

import { execFile } from "node:child_process";
import { readFile } from "node:fs/promises";
import { promisify } from "node:util";

const execFileAsync = promisify(execFile);
await execFileAsync("tesseract", ["input.png", "output", "-l", "eng"]);
const text = await readFile("output.txt", "utf8");
console.log(text);

Pass command arguments as an array rather than assembling a shell command from user-controlled filenames. This avoids shell interpretation of those values.

3. Use cloud OCR for images and documents

Cloud OCR is useful when an application needs a managed API, structured results, or batch processing. Choose the operation based on the input: general image text for photos and signs, document OCR for dense pages, and verify current service requirements, accepted inputs, regions, limits, and prices before production use.

Google Cloud Vision

For a general image, choose TEXT_DETECTION. For a dense document image, choose DOCUMENT_TEXT_DETECTION. Both return extracted text and geometry; document detection returns a hierarchy suited to page content. Google documents offline asynchronous batch annotation for up to 2,000 image files per request; confirm live limits and pricing when planning a production job. See the OCR guide and language support page.

A REST request for a publicly accessible image URI uses an authenticated Google Cloud project and the Vision API endpoint:

curl -X POST \
  -H "Authorization: Bearer $(gcloud auth print-access-token)" \
  -H "Content-Type: application/json; charset=utf-8" \
  https://vision.googleapis.com/v1/images:annotate \
  -d '{"requests":[{"image":{"source":{"imageUri":"https://example.com/image.jpg"}},"features":[{"type":"TEXT_DETECTION"}]}]}'

For a dense document, change the feature type to DOCUMENT_TEXT_DETECTION. Use an image URI accessible to the service or send image content as base64 according to the API documentation. The example assumes the Google Cloud CLI is authenticated and the API and billing configuration are in place.

Microsoft Azure OCR

Microsoft recommends its Azure image OCR capability for general non-document images and Document Intelligence Read for scanned or digital text-heavy documents. Document Intelligence returns recognized lines and words with locations and confidence values. These are documented service capabilities, not a guarantee that any specific image will be read correctly. Review the current Read model guidance and OCR edition selection and API guidance for the API version and input requirements you intend to use.

4. Extract visible text from video

Video OCR recognizes text rendered in frames: examples include burned-in subtitles, title cards, credits, and signs. It does not infer the spoken words from the soundtrack. If you need both, run video OCR and a separate speech transcription workflow, then align their outputs as needed.

Google Video Intelligence accepts a video stored in Cloud Storage and documents a TEXT_DETECTION annotation request. Its output associates recognized text with frame-level locations and timestamps. See Google’s video text detection guide.

Upload the video to a Cloud Storage bucket and save this request as request.json, replacing the URI and optional language hint:

{
  "inputUri": "gs://YOUR_BUCKET/video.mp4",
  "features": ["TEXT_DETECTION"],
  "videoContext": {
    "textDetectionConfig": {
      "languageHints": ["en-US"]
    }
  }
}
curl -X POST \
  -H "Authorization: Bearer $(gcloud auth print-access-token)" \
  -H "Content-Type: application/json; charset=utf-8" \
  "https://videointelligence.googleapis.com/v1/videos:annotate" \
  -d @request.json

This starts an asynchronous operation. Keep the returned operation name, poll the operation endpoint until it finishes, then inspect the annotation results. The Google guide includes client library examples and shows how to read text segments, timestamps, and rotated bounding boxes. A local video can also be sent as base64 in inputContent, but a Cloud Storage URI avoids embedding a large file in the request body.

5. Improve recognition and preserve useful output

  • Start with the original: use the clearest available source. Repeated resizing or compression can erase character detail.
  • Make text legible: crop to the relevant area when practical, and avoid including large unrelated regions if they make layout interpretation harder.
  • Check orientation and language: rotate sideways images and install or specify the correct language model. For Google Vision, language hints are optional; incorrect hints can hinder recognition.
  • Select layout-aware output for documents: a plain string may lose reading order, columns, line breaks, or page boundaries. Keep boxes or document structure when downstream tasks rely on layout.
  • Keep video timestamps: retain the returned segment and frame timing when exporting text. Repeated detections across adjacent frames may refer to the same title or subtitle; deduplicate carefully without discarding the interval during which it appears.
  • Review consequential text: names, serial numbers, punctuation, small print, handwriting, columns, and mixed languages deserve visual checking before they drive a decision or enter a database.

OCR is a recognition result, not a guarantee of correctness. No engine or accuracy rate is universally best for every image, font, language, and capture condition.

6. Or skip the browser setup

If the image you need to read is on a webpage, first capture the page as an image, then pass that image to your OCR workflow. ScreenshotNeo is a website screenshot API and MCP server; it captures a webpage but does not perform OCR. A single GET request returns an image or PDF. See the ScreenshotNeo website and API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write("shot.webp", res);

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed, and response headers report the page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan. After capturing, feed shot.webp to Tesseract or your selected OCR service. Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

7. Troubleshooting

Symptom Likely cause Fix
tesseract: command not found Tesseract is not installed or is missing from PATH. Install the executable, reopen the terminal, and verify with tesseract --version. In Python, set the wrapper’s executable path if needed.
Language data error or wrong-language output The requested trained data is absent or the language code is wrong. Check tesseract --list-langs, install the required language data, and pass the correct code with -l.
PDF input is rejected by Tesseract Tesseract reads image formats, not PDF input. Convert each PDF page to a supported image format or use a PDF OCR workflow such as OCRmyPDF.
Output is empty or badly garbled Text may be too small, blurred, low contrast, rotated, obscured, or in an unsupported language. Return to the best source, correct orientation, crop or improve legibility, and verify the language model. Inspect the output against the image.
Paragraph order or columns are mixed Plain text output does not preserve enough page layout for the use case. Use document-focused OCR and retain blocks, lines, or bounding boxes rather than only the flattened string.
Video request returns an operation instead of text Video annotation runs asynchronously. Poll the returned operation to completion, then read its annotation results. Check that the input URI is in Cloud Storage and accessible to the project.
Video OCR misses spoken words OCR only reads text visible in frames. Run a separate speech-to-text transcription service for dialogue.
Cloud API returns permission or authentication errors Credentials, API enablement, project permissions, or billing setup are incomplete. Confirm the active cloud project and credentials, enable the relevant API, and follow that provider’s current authentication setup.

8. Performance, reliability, and cost

Local Tesseract avoids sending images to a cloud OCR service, which can matter for privacy and recurring workloads, but you manage installation, language files, CPU capacity, and processing. Cloud services reduce local setup and can return structured results or process batches, but add network latency, provider quotas, data handling considerations, and service charges. Review current pricing and limits directly with the provider; this guide makes no cost or speed comparison.

For repeatable pipelines, separate image acquisition, OCR, validation, and storage. Use bounded concurrency rather than launching an unbounded job for every file. Retain source identifiers and page or frame timestamps alongside recognized text so results can be traced back. Retry transient network or service errors with limits and backoff; do not retry permanent input, permission, or authentication failures unchanged. Cache results only when the source content and OCR settings are unchanged, and protect uploaded images and credentials according to your data policy.

For large image sets, Google documents asynchronous batch annotation for up to 2,000 image files per request. Check the live quota and price before relying on that capacity. Video processing is asynchronous, so design for operation polling and delayed completion rather than assuming an immediate response.

9. Frequently asked questions

Can I copy text from a picture without installing software?

Yes. Use an image OCR feature in a cloud service or another OCR-capable application. For private or offline work, a local engine such as Tesseract is an option.

Can OCR read handwriting?

Some OCR services document handwriting support, but readability varies with writing, image quality, language, and layout. Verify the recognized result before relying on it.

How do I get subtitles from a video?

If subtitles are burned into the picture, use video OCR and keep its frame timestamps. If subtitles are spoken dialogue, use speech recognition. A video may need both workflows.

Can Tesseract create a searchable PDF?

Yes. Tesseract can write PDF output from supported image input, embedding a searchable text layer. It does not use a PDF as its input format.

Which OCR engine is most accurate?

There is no universal answer. Compare candidates using representative samples from your own images, languages, and layouts, and review errors that matter to your application.