ScreenshotNeo

BlogHow-to

How to Convert a JPG Image to Text

Convert a JPG to editable text with Google Drive, OneNote, Acrobat or local Tesseract OCR, plus accuracy fixes and automation examples.

By the ScreenshotNeo team1 October 20267 min read

Use optical character recognition (OCR) to convert the pixels in a JPG into selectable, editable text. For a small, clear image, the quickest method is Google Drive: upload the JPG, right-click it, choose Open with > Google Docs, then review the extracted text.

Choose another method when your requirements differ:

  • Google Drive: easiest general-purpose option for one or a few images.
  • OneNote: convenient if you already work in Microsoft 365 and only need copied text.
  • Adobe Acrobat: best when the result must be a searchable PDF or you have several documents.
  • Tesseract: free, scriptable OCR that runs locally.
  • SharePoint OCR: suited to organizational indexing and compliance workflows.

What you need before converting a JPG

OCR accuracy depends more on the source image than on the button you click. Prepare the JPG as follows:

  • Keep the file around 2 MB or smaller when using Google Drive.
  • Make text at least 10 pixels high.
  • Rotate the image upright.
  • Use sharp focus, even lighting, strong contrast and common fonts.
  • Crop away borders, backgrounds and unrelated objects.
  • For multilingual text, identify every language before recognition.

Handwriting, skew, shadows, unusual fonts and complex layouts require extra proofreading. No OCR method should be treated as an unquestioned transcription.

Method 1: Convert a JPG to text with Google Drive

  1. Open Google Drive and upload the JPG.
  2. Right-click the uploaded file.
  3. Select Open with > Google Docs.
  4. Google Docs creates a document containing the image and the recognized text.
  5. Read the output against the original image and correct errors.
  6. Copy the text or download the document in the format you need.

Google documents this workflow for JPEG photo files and recommends a file of 2 MB or smaller, upright text, at least 10 pixels high, sharp focus, even lighting and clear contrast. Formatting may not transfer perfectly. Google’s image-to-text instructions explain the supported workflow and image guidance.

When Google Drive is the right choice

  • You need a quick result without installing software.
  • The JPG contains a modest amount of printed text.
  • You can upload the image to a cloud service.

Method 2: Copy text from a JPG with OneNote

  1. Insert the JPG into a OneNote page.
  2. Right-click the picture.
  3. Choose Copy Text from Picture.
  4. Paste the result into OneNote, Word or another editor.
  5. Compare the pasted text with the picture and repair recognition errors.

For a multi-page printout, OneNote can copy text from the current page or all pages. Microsoft recommends checking the recognized text because image quality affects OCR effectiveness. See Microsoft’s OneNote OCR guidance.

Method 3: Create searchable text with Adobe Acrobat

  1. Open Acrobat and choose All tools > Scan & OCR > In this file.
  2. Select the page range and recognition language.
  3. Choose Recognize Text.
  4. Save the file. Acrobat adds a searchable text layer.
  5. Search through the PDF and correct important fields or passages.

For Acrobat’s web workflow, use Convert > Recognize text with OCR, select the file and choose Recognize text. Acrobat is useful when the deliverable is a searchable or editable PDF rather than a plain text snippet. Sources: Acrobat desktop OCR and Acrobat web OCR.

Method 4: Run Tesseract locally

Tesseract is an open-source OCR engine that accepts JPEG and other raster formats. It is a good fit when the image cannot be uploaded to a cloud service or when you need a repeatable command-line workflow.

Basic command

tesseract myscan.jpg output

This writes recognized text to output.txt.

Specify a language

tesseract myscan.jpg output -l eng

Combine installed languages when needed:

tesseract receipt.jpg receipt -l eng+deu

Produce a searchable PDF

tesseract myscan.jpg searchable -l eng pdf

Python wrapper around the Tesseract command

This example keeps OCR local and calls the same executable used above:

from pathlib import Path
import subprocess

image = Path("myscan.jpg")
output_base = Path("output")

subprocess.run(
    ["tesseract", str(image), str(output_base), "-l", "eng"],
    check=True,
)

text = output_base.with_suffix(".txt").read_text(encoding="utf-8")
print(text)

Install Tesseract through your operating system’s package manager, then confirm it is available with tesseract --version. Language data must also be installed for every code you pass to -l.

Node.js wrapper around the Tesseract command

import { execFile } from "node:child_process";
import { readFile } from "node:fs/promises";
import { promisify } from "node:util";

const execFileAsync = promisify(execFile);
await execFileAsync("tesseract", ["myscan.jpg", "output", "-l", "eng"]);
const text = await readFile("output.txt", "utf8");
console.log(text);

What about cURL?

Tesseract is a local executable, not an HTTP service, so there is no general cURL endpoint in this workflow. Use cURL only when the OCR provider you choose documents an HTTP API with an upload endpoint and authentication scheme. Do not send sensitive images to an unverified endpoint.

Method 5: SharePoint OCR for organizational libraries

Microsoft SharePoint OCR can extract and index text from JPG/JPEG and many other image formats. Microsoft documents support for more than 150 languages and images below 50 MB, with dimensions from 50 × 50 to 16,000 × 16,000 pixels. This is an organizational search and compliance workflow rather than the simplest personal conversion method. Review Microsoft’s current SharePoint OCR documentation before configuring a tenant.

Choosing the right method

Method Output Privacy and control Best fit
Google Drive Editable Google Doc Cloud upload Fast one-off conversion
OneNote Copied text Cloud or Microsoft 365 workspace Existing OneNote users
Acrobat Searchable PDF text layer Desktop or web workflow PDF production and document batches
Tesseract TXT, searchable PDF and other formats Local processing Automation, privacy and repeatability
SharePoint OCR Indexed organizational text Enterprise tenant controls Libraries, search and compliance

Improve OCR accuracy

  1. Deskew: rotate lines so they are horizontal.
  2. Crop: remove margins, backgrounds and unrelated objects.
  3. Increase contrast: separate characters from the page.
  4. Reduce noise: remove speckles and compression artifacts.
  5. Use the correct language: install and select the appropriate OCR language data.
  6. Segment complex layouts: process columns, tables or captions separately when a single pass mixes their reading order.
  7. Review high-risk fields: names, dates, totals, account numbers, URLs and code symbols.

Common errors and fixes

Symptom Likely cause Fix
Almost no text is returned Low resolution, blur or poor contrast Use a sharper source, crop it, improve lighting and enlarge the text.
Characters are consistently wrong Wrong language model or unusual font Select the correct language and compare representative characters with the source.
Text order is scrambled Columns, tables or mixed regions Crop each region and OCR it separately, then rebuild the reading order.
Output contains many spaces Background texture, shadows or skew Clean the image, deskew it and increase contrast before recognition.
OneNote’s option is missing The picture was not inserted as an image or the client differs Insert the JPG directly, update OneNote and try the picture’s context menu again.
tesseract: command not found Tesseract is not installed or not on PATH Install Tesseract, reopen the terminal and run tesseract --version.
Tesseract reports a missing language Language data is not installed Install the required trained data and verify the language code.
Searchable PDF has wrong words OCR errors were embedded in the text layer Proofread before distributing or relying on the PDF.

Performance, reliability and cost

  • Small, clear images are fastest: preprocessing can reduce retries and improve recognition.
  • Batch work: Acrobat and SharePoint are better suited to document collections; Tesseract can be scripted for repeatable local jobs.
  • Reliability: retain the original JPG, OCR output and a review status so errors can be corrected later.
  • Privacy: cloud tools upload the image; local Tesseract keeps processing on your machine.
  • Cost: Google Drive, OneNote, Acrobat and SharePoint availability depends on the account or subscription you use. Tesseract itself is open source, but installation and maintenance remain your responsibility.

Or skip the browser setup

If the JPG comes from a webpage, ScreenshotNeo can capture the page before you run OCR. It is a website screenshot API, so it does not replace OCR; it gives you a clean image source through one request.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for capture options. Cookie banners, newsletter popups and chat widgets are removed before the shot. Bot checks, blank pages and failed loads are never billed. An MCP server lets AI agents take screenshots. You get 1,000 screenshots a month free with no card, and paid plans start at $5 for 3,000. Sign up free for ScreenshotNeo.

FAQ

Can I convert a JPG to text without installing software?

Yes. Google Drive, OneNote and Acrobat provide GUI workflows. Uploading an image to a cloud service may not suit confidential material.

Does OCR preserve the original formatting?

Usually not perfectly. Plain paragraphs are easier than columns, tables, forms and mixed layouts. Preserve the image and proofread the reconstructed document.

Is Tesseract accurate enough for invoices or IDs?

It can help, but verify every critical field manually. Blur, glare, small type and unusual fonts can change numbers or characters.

Which method keeps the image on my computer?

Tesseract can run locally. Google Drive, OneNote, Acrobat web and SharePoint process images through cloud services or organizational infrastructure.

Can OCR read handwriting?

Results vary widely. The documented workflows are most dependable for clear printed text, so handwritten output needs especially careful review.