ScreenshotNeo

BlogHow-to

How to Convert an Image into an HTML Table

Learn how to turn a table screenshot or photo into accurate, accessible HTML with OCR, cell detection, merged-cell handling, validation, and runnable code.

By the ScreenshotNeo team1 October 202610 min read

To convert an image into an HTML table, solve two separate problems: recognize the words and reconstruct the table layout. OCR alone can read text while losing rows, columns, headers, and merged cells.

A reliable pipeline is:

  1. Prepare and normalize the image.
  2. Detect the table and its cell boundaries.
  3. Run OCR that returns text and coordinates.
  4. Assign words to cells and rebuild spans.
  5. Generate semantic, escaped HTML.
  6. Validate the result against the source image.

For a one-off table, a managed table-recognition service is fastest. For privacy or cost control, use local OCR such as Tesseract and implement the structure step yourself.

1. Prepare the source image

Keep the original file unchanged so every generated cell can be checked later. Work on a copy.

  • Crop away surrounding page content.
  • Deskew a photographed or scanned page.
  • Upscale small text before OCR.
  • Increase contrast and convert to grayscale.
  • Remove shadows, glare, and compression artifacts.
  • Preserve faint lines long enough for table detection, then create a separate line-free image for OCR if needed.

Do not over-process the image. Aggressive thresholding can erase decimal points, minus signs, and thin characters.

2. Choose an extraction approach

Approach Structure fidelity Privacy Implementation Best use
Amazon Textract Returns cells, merged-cell relationships, headers, titles, footers, and table type information. Managed service; check your AWS region and data requirements. Low to medium Production tables where structure matters.
Google Vision or Document AI Vision returns document hierarchy, words, and bounding boxes. Google recommends Document AI for scanned-document parsing and structured extraction. Managed service; check project and region requirements. Medium Google Cloud pipelines and document workflows.
Tesseract Returns text positions through hOCR or TSV; you implement table detection and cell grouping. Local processing Medium to high Privacy, offline processing, and predictable infrastructure.
Table Transformer Detects tables and recognizes structure; its HTML export does not retain cell bounding boxes. Local model or hosted deployment Medium Separating table detection from OCR.

Amazon documents Textract table entities and relationships in its table analysis guide. AWS Samples shows HTML generation with the Textractor package in its Textractor repository. Google documents OCR responses and recommends Document AI for scanned-document parsing in its OCR documentation. Tesseract documents positional hOCR and TSV output in its command-line guide. The Table Transformer project documents HTML and CSV export in its repository.

3. Local conversion with Tesseract and Python

This example uses OpenCV to find horizontal and vertical rules, Tesseract TSV output for word coordinates, and HTML escaping before insertion. It is intentionally explicit so you can replace the line detector with a table model when your images are irregular.

Install dependencies

python -m pip install opencv-python pytesseract pandas
# Install the Tesseract executable with your operating system package manager.
# Debian/Ubuntu example:
sudo apt-get install tesseract-ocr

Complete script

from __future__ import annotations

import html
import sys
from pathlib import Path

import cv2
import pandas as pd
import pytesseract
from pytesseract import Output


def clean_image(path: str) -> tuple[object, object]:
    original = cv2.imread(path)
    if original is None:
        raise FileNotFoundError(path)
    gray = cv2.cvtColor(original, cv2.COLOR_BGR2GRAY)
    gray = cv2.resize(gray, None, fx=2, fy=2, interpolation=cv2.INTER_CUBIC)
    normalized = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(8, 8)).apply(gray)
    return original, normalized


def line_boxes(gray: object) -> tuple[list[tuple[int, int, int, int]], list[int], list[int]]:
    binary = cv2.adaptiveThreshold(
        gray, 255, cv2.ADAPTIVE_THRESH_MEAN_C, cv2.THRESH_BINARY_INV, 21, 10
    )
    width = binary.shape[1]
    height = binary.shape[0]
    horizontal_kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (max(20, width // 30), 1))
    vertical_kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (1, max(20, height // 30)))
    horizontal = cv2.morphologyEx(binary, cv2.MORPH_OPEN, horizontal_kernel)
    vertical = cv2.morphologyEx(binary, cv2.MORPH_OPEN, vertical_kernel)
    h_contours, _ = cv2.findContours(horizontal, cv2.RETR_LIST, cv2.CHAIN_APPROX_SIMPLE)
    v_contours, _ = cv2.findContours(vertical, cv2.RETR_LIST, cv2.CHAIN_APPROX_SIMPLE)
    h_lines = sorted({cv2.boundingRect(c)[1] for c in h_contours if cv2.boundingRect(c)[2] > width * 0.15})
    v_lines = sorted({cv2.boundingRect(c)[0] for c in v_contours if cv2.boundingRect(c)[3] > height * 0.15})
    if len(h_lines) < 2 or len(v_lines) < 2:
        raise ValueError("Could not find enough table rules; use a table detector or tune the kernels")
    boxes = []
    for r in range(len(h_lines) - 1):
        for c in range(len(v_lines) - 1):
            x1, x2 = v_lines[c], v_lines[c + 1]
            y1, y2 = h_lines[r], h_lines[r + 1]
            boxes.append((x1, y1, x2, y2))
    return boxes, h_lines, v_lines


def ocr_words(gray: object) -> pd.DataFrame:
    data = pytesseract.image_to_data(gray, output_type=Output.DATAFRAME, config="--psm 6")
    data = data.dropna(subset=["text"])
    data["text"] = data["text"].astype(str).str.strip()
    return data[data["text"] != ""]


def assign_cells(words: pd.DataFrame, h_lines: list[int], v_lines: list[int]) -> list[list[str]]:
    rows = len(h_lines) - 1
    cols = len(v_lines) - 1
    cells = [[[] for _ in range(cols)] for _ in range(rows)]
    for _, word in words.iterrows():
        cx = float(word.left) + float(word.width) / 2
        cy = float(word.top) + float(word.height) / 2
        col = next((i for i in range(cols) if v_lines[i] <= cx < v_lines[i + 1]), None)
        row = next((i for i in range(rows) if h_lines[i] <= cy < h_lines[i + 1]), None)
        if row is not None and col is not None:
            cells[row][col].append((float(word.top), float(word.left), str(word.text)))
    return [[" ".join(t[2] for t in sorted(cell, key=lambda x: (x[0], x[1]))) for cell in row] for row in cells]


def to_html(cells: list[list[str]], header_rows: int = 1) -> str:
    out = ["<table>"]
    if header_rows:
        out.append("<thead>")
        for row in cells[:header_rows]:
            out.append("<tr>" + "".join(f"<th scope=\"col\">{html.escape(value)}</th>" for value in row) + "</tr>")
        out.append("</thead>")
    out.append("<tbody>")
    for row in cells[header_rows:]:
        out.append("<tr>" + "".join(f"<td>{html.escape(value)}</td>" for value in row) + "</tr>")
    out.extend(["</tbody>", "</table>"])
    return "\n".join(out)


if __name__ == "__main__":
    image_path = sys.argv[1] if len(sys.argv) > 1 else "table.png"
    _, gray = clean_image(image_path)
    _, h_lines, v_lines = line_boxes(gray)
    words = ocr_words(gray)
    cells = assign_cells(words, h_lines, v_lines)
    Path("table.html").write_text(to_html(cells, header_rows=1), encoding="utf-8")
    print("Wrote table.html")

This baseline assumes visible, mostly straight rules and one header row. It will not infer merged cells correctly. For borderless tables, rotated pages, nested headers, or row and column spans, use a table-structure model or a managed API and retain the returned geometry.

4. Managed extraction with Amazon Textract

Textract returns table cells and relationships that are difficult to reproduce with line morphology. A typical workflow is:

  1. Upload or provide the image bytes according to the Textract API you use.
  2. Call table analysis.
  3. Read cell blocks, relationships, confidence values, and entity types.
  4. Map row and column indexes to HTML rows and cells.
  5. Convert merged-cell relationships to rowspan and colspan.

AWS Samples’ Textractor package can render a detected table with to_html(). Treat that output as a starting point: inspect header semantics, escape text, and validate spans before publishing.

5. Reconstruct rows, columns, and merged cells

Assign OCR words to cells

With word bounding boxes, calculate each word’s center point and assign it to the cell rectangle containing that point. For words crossing a boundary, assign by the largest intersection area or the nearest cell center, then flag the decision for review.

Handle line wraps

Sort words by their top coordinate and then left coordinate. Group words on nearby baselines into lines, preserving spaces between words. Never join numeric tokens without checking decimal separators and thousands separators.

Represent spans

Use colspan when one cell covers multiple columns and rowspan when it covers multiple rows. Keep a grid of occupied coordinates so later cells are placed in the next unoccupied slot.

Choose header semantics

Use <thead> for column headers and <th scope="row"> for row labels. Multi-row headers may need headers and id attributes when scope alone is ambiguous.

6. Generate safe, accessible HTML

Escape every OCR value before inserting it into markup. OCR output is untrusted input and can contain characters that change the document.

import html
safe_value = html.escape(ocr_value, quote=True)
cell = f'<td>{safe_value}</td>'

Add a caption when the source image has a meaningful title, preserve empty cells as empty <td> elements, and avoid replacing missing values with guesses. Add CSS separately rather than embedding presentation into extracted text.

7. Validation checklist

  • Row count and column count match the image.
  • Every header is in the correct row or column.
  • Numbers, signs, decimal separators, and dates were checked manually.
  • Empty cells remain empty.
  • Wrapped text was not merged with an adjacent cell.
  • Rowspan and colspan positions do not overlap.
  • Low-confidence or ambiguous cells have been reviewed.
  • HTML passes a parser and an accessibility checker.
  • The original image, OCR coordinates, and transformation version are retained when the table supports financial, medical, legal, or operational decisions.

8. Difficult images and edge cases

Borderless tables

Infer columns from aligned text centers or use a structure-recognition model. Do not assume equal-width columns.

Photographs and perspective

Detect the page quadrilateral, apply a perspective transform, then deskew. Re-run detection after transformation.

Rotated or multilingual content

Detect orientation before OCR and install the required Tesseract language data or configure the managed service’s language options.

Handwriting and faint text

Expect lower confidence and human review. OCR confidence is a prioritization signal, not proof of semantic correctness.

Nested tables

Detect the outer table first, then process nested regions independently. Decide whether nested markup belongs inside a cell or should become a separate table.

9. Troubleshooting

Symptom Likely cause Fix
No cells detected Lines are faint, broken, or the image is borderless. Improve contrast, tune morphology kernels, or use a table-structure model.
Words appear in adjacent cells Deskew or coordinate scaling is wrong. Use the same resized image for line detection and OCR; verify bounding-box units.
Rows are merged OCR line grouping used text order without geometry. Group by baseline and assign through cell rectangles.
Numbers are wrong Low resolution, thresholding, or language mismatch. Upscale, preserve grayscale, configure language, and manually review numeric cells.
HTML is malformed Unescaped OCR characters or overlapping spans. HTML-escape values and validate the occupied-cell grid.
Header is inaccessible Headers were emitted as ordinary <td> cells. Use <th>, scope, or explicit headers/id relationships.
Cloud request fails Credentials, region, permissions, or payload limits. Check the provider’s request logs, IAM permissions, selected region, and supported file format.

10. Performance, reliability, and cost

Preprocessing and OCR are usually CPU-bound locally. Upscaling improves recognition but increases memory and processing time. Cache OCR and geometry results by a hash of the original image so repeated HTML renders do not repeat extraction.

For batch work, process images concurrently within provider quotas, retry transient failures with exponential backoff, and store confidence and coordinates alongside the HTML. Keep deterministic preprocessing settings so regenerated tables can be compared.

Managed APIs reduce engineering effort but charge per page, image, or operation according to their current pricing. Local Tesseract avoids per-request service charges but shifts the cost to compute, model maintenance, and review time. Compare structure fidelity, merged-cell support, language coverage, data residency, throughput, confidence reporting, HTML export, and auditability before choosing.

11. Or skip the browser setup

If the source is a web page rather than a local photo, ScreenshotNeo can create a clean image for the OCR step with one request. Its consent handling removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing result.

See the ScreenshotNeo API documentation for options such as full-page capture, element selection, custom CSS and JavaScript, waiting conditions, device presets, and image formats.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes an MCP server so Claude, Cursor, and other MCP clients can take screenshots, inspect pages, and capture PDFs. The free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

12. FAQ

Can OCR alone convert an image to a table?

No. OCR recognizes text; table conversion also requires geometry, row and column grouping, and span reconstruction.

Which format preserves the original layout?

Keep the source image and coordinate data. HTML represents semantics but can discard exact geometry.

How do I handle a table with no visible borders?

Use aligned text positions or a table-structure model, then validate inferred columns against the image.

Should I trust OCR confidence scores?

Use them to prioritize review. They do not prove that a value has the correct meaning or belongs to the correct cell.

Can generated HTML be used directly in a public page?

Only after escaping all extracted text, validating markup, checking accessibility, and reviewing sensitive values.