ScreenshotNeo

BlogHow-to

HTML Table to JSON: A Complete Guide for Browser, Python, and Node.js

Convert HTML tables to reliable JSON with browser JavaScript, Python, and Node.js examples, plus handling spans, headers, types, and errors.

By the ScreenshotNeo team1 October 202610 min read

To convert a regular HTML table to JSON, select the table, read its header cells as keys, map each body row to an object, decide how to handle duplicate or empty headings, and serialize the result with JSON.stringify(). The simple mapping works when every row has the same columns. Tables with rowspan, colspan, multi-level headers, nested markup, or missing cells need an explicit schema or a converter that understands those structures.

An HTML table is not always a flat grid. The HTML specification defines captions, column groups, header, body, and footer sections, while the browser exposes the structure through HTMLTableElement and related row and cell interfaces. See the WHATWG table specification before assuming that visual position alone identifies a field.

1. Decide the JSON shape first

The most common output is an array of row objects:

[{"name":"Ada","team":"Platform","active":true},{"name":"Lin","team":"Data","active":false}]

This shape is a design choice, not an automatic property of HTML. Before writing a parser, decide:

  • Which table is the intended input when a page contains several tables.
  • Whether headings come from the first row, an explicit thead, or a supplied schema.
  • How blank headings are named.
  • What to do with duplicate headings: suffix them, combine values, or reject the table.
  • Whether cell values remain strings or become numbers, booleans, dates, or null.
  • Whether hidden columns, footer rows, totals, and pagination controls belong in the result.

2. Browser JavaScript for a regular table

Given one header row and one or more body rows, this function converts the selected table into objects. It preserves values as strings by default, trims surrounding whitespace, and gives blank or repeated headings deterministic names.

function tableToJson(table, { parseValues = false } = {}) {
  const headerCells = table.querySelectorAll('thead tr:last-child th');
  const firstRow = table.querySelector('tr');
  const cells = headerCells.length
    ? [...headerCells]
    : firstRow
      ? [...firstRow.querySelectorAll('th, td')]
      : [];

  const counts = new Map();
  const keys = cells.map((cell, index) => {
    const base = cell.textContent.trim() || `column_${index + 1}`;
    const count = (counts.get(base) || 0) + 1;
    counts.set(base, count);
    return count === 1 ? base : `${base}_${count}`;
  });

  const rows = table.tBodies.length
    ? [...table.tBodies].flatMap(body => [...body.rows])
    : [...table.rows].slice(headerCells.length ? 0 : 1);

  return rows.map(row => {
    const values = [...row.cells].map(cell => cell.textContent.trim());
    return Object.fromEntries(keys.map((key, index) => [
      key,
      parseValues ? parseCell(values[index] ?? '') : (values[index] ?? '')
    ]));
  });
}

function parseCell(value) {
  if (value === '') return null;
  if (/^(true|false)$/i.test(value)) return value.toLowerCase() === 'true';
  if (/^-?(?:0|[1-9]\d*)(?:\.\d+)?$/.test(value)) return Number(value);
  return value;
}

const table = document.querySelector('#orders');
if (!table) throw new Error('Could not find #orders');
const data = tableToJson(table, { parseValues: true });
console.log(JSON.stringify(data, null, 2));

Use textContent when you want readable cell text. If the cell contains links, images, or markup whose meaning matters, extract those fields separately instead of assuming that all HTML should become text.

Selecting the right rows

A page may contain a header row in thead, data rows in several tbody elements, and a totals row in tfoot. The example reads body rows and avoids the footer. If the page has no semantic sections, select rows explicitly and skip known header or summary rows.

const table = document.querySelector('table.results');
const rows = [...table.querySelectorAll('tbody tr:not(.total)')];

3. A complete HTML example

<table id="people">
  <thead>
    <tr><th>Name</th><th>Age</th><th>Active</th></tr>
  </thead>
  <tbody>
    <tr><td>Ada</td><td>36</td><td>true</td></tr>
    <tr><td>Lin</td><td>29</td><td>false</td></tr>
  </tbody>
</table>
<script>
  const result = tableToJson(document.querySelector('#people'), { parseValues: true });
  console.log(JSON.stringify(result, null, 2));
</script>

Output:

[
  {"Name":"Ada","Age":36,"Active":true},
  {"Name":"Lin","Age":29,"Active":false}
]

4. Handling difficult table structures

Duplicate or empty headings

Two <th> elements with the same text cannot safely become the same object key. A suffix policy such as price, price_2 preserves both columns. For blank headings, use stable names such as column_1, or reject the input when a schema is required.

colspan and rowspan

A spanning cell occupies multiple visual columns or rows, so row.cells[index] is not necessarily the same logical column in every row. Build a rectangular occupancy grid before assigning values, or provide a schema that says which logical field each cell belongs to. Do not silently shift values left when a cell is missing.

Multi-level headers

For headers such as “Sales” spanning “Units” and “Revenue”, derive paths like Sales.Units and Sales.Revenue, or pass an explicit list of keys. A robust converter must associate header cells by their scope and spans; taking only the last visible header row can lose meaning.

Missing cells and extra cells

Choose a policy: fill missing fields with null, fill with an empty string, or report a row error. If a row has more cells than the schema, preserve the extras under an error field or reject the row. Silent truncation makes downstream data difficult to audit.

Nested markup

Use textContent for plain text, but inspect links, data-* attributes, images, and buttons when they carry the real value. For example, a displayed “View” button may have an order ID only in its URL or data attribute.

Multiple tables and rendered grids

Use a stable selector, table caption, or nearby heading. Some JavaScript grids are styled div elements rather than real table elements; in that case, inspect the grid’s data model or accessibility roles instead of assuming table APIs will work.

5. Python: convert saved HTML with Beautiful Soup

Install the parser once:

python -m pip install beautifulsoup4

This script reads a local file, selects one table, applies a duplicate-heading policy, and keeps values as strings. The explicit conversion step makes type handling visible.

from bs4 import BeautifulSoup
import json
from pathlib import Path


def unique_keys(cells):
    counts = {}
    keys = []
    for index, cell in enumerate(cells, start=1):
        base = cell.get_text(" ", strip=True) or f"column_{index}"
        counts[base] = counts.get(base, 0) + 1
        keys.append(base if counts[base] == 1 else f"{base}_{counts[base]}")
    return keys


def html_table_to_json(path, selector="table", parse_values=False):
    soup = BeautifulSoup(Path(path).read_text(encoding="utf-8"), "html.parser")
    table = soup.select_one(selector)
    if table is None:
        raise ValueError(f"No table matched {selector!r}")

    header = table.select_one("thead tr")
    if header is None:
        rows = table.select("tr")
        if not rows:
            return []
        header, rows = rows[0], rows[1:]
    else:
        rows = table.select("tbody tr") or table.select("tr")[1:]

    keys = unique_keys(header.select("th, td"))
    output = []
    for row_number, row in enumerate(rows, start=1):
        values = [cell.get_text(" ", strip=True) for cell in row.select("th, td")]
        if len(values) != len(keys):
            raise ValueError(
                f"Row {row_number} has {len(values)} cells; expected {len(keys)}"
            )
        record = {}
        for key, value in zip(keys, values):
            if parse_values and value == "":
                record[key] = None
            elif parse_values and value.lower() in {"true", "false"}:
                record[key] = value.lower() == "true"
            else:
                record[key] = value
        output.append(record)
    return output

records = html_table_to_json("page.html", "table#orders", parse_values=True)
print(json.dumps(records, indent=2, ensure_ascii=False))

For remote HTML, fetch it with an HTTP client, check the status code and content type, then pass the response text to Beautiful Soup. Respect the site’s access rules and do not assume that a server response contains rows rendered later by JavaScript.

6. Node.js: parse HTML with Cheerio

Install Cheerio in a project:

npm install cheerio
const fs = require('node:fs');
const cheerio = require('cheerio');

function uniqueKeys($, cells) {
  const counts = new Map();
  return cells.map((_, cell) => {
    const base = $(cell).text().replace(/\s+/g, ' ').trim() || `column_${cells.index(cell) + 1}`;
    const count = (counts.get(base) || 0) + 1;
    counts.set(base, count);
    return count === 1 ? base : `${base}_${count}`;
  });
}

function tableToJson(html, selector = 'table') {
  const $ = cheerio.load(html);
  const table = $(selector).first();
  if (!table.length) throw new Error(`No table matched ${selector}`);

  const header = table.find('thead tr').last();
  const headerCells = header.length ? header.find('th,td').toArray() : table.find('tr').first().find('th,td').toArray();
  const keys = uniqueKeys($, headerCells);
  const rows = table.find('tbody tr').length
    ? table.find('tbody tr').toArray()
    : table.find('tr').slice(1).toArray();

  return rows.map((row, rowIndex) => {
    const values = $(row).find('th,td').toArray().map(cell => $(cell).text().replace(/\s+/g, ' ').trim());
    if (values.length !== keys.length) {
      throw new Error(`Row ${rowIndex + 1} has ${values.length} cells; expected ${keys.length}`);
    }
    return Object.fromEntries(keys.map((key, i) => [key, values[i]]));
  });
}

const html = fs.readFileSync('page.html', 'utf8');
console.log(JSON.stringify(tableToJson(html, 'table#orders'), null, 2));

Cheerio parses the HTML you provide. It does not execute page JavaScript, wait for network requests, or simulate a browser. Use a browser automation tool when rows appear only after client-side rendering.

7. Standards-based conversion versus DOM mapping

The W3C documents “Generating JSON from Tabular Data on the Web” and the related tabular data model describe annotated tables, cell parsing, metadata, and standard and minimal conversion modes. The conversion document states: “A conformant JSON conversion application MUST produce output conforming to this algorithm according to the chosen mode of conversion: standard or minimal.” That requirement applies to the annotated tabular-data model; it does not define every ad hoc DOM-to-object mapping.

Use a simple mapper when you control the markup and need a small array of records. Evaluate a standards-oriented workflow when metadata, annotations, typed values, language information, or interoperable table descriptions matter. In either case, document how blanks, invalid values, dates, numbers, and parse errors are represented.

8. Library option: tabletojson

The tabletojson package documents conversion from HTML markup or a URL and options covering duplicate headings, row and column spans, complex headers, HTML in cells, ignored columns, and row limits. Treat those options as evaluation checkpoints, not as a guarantee that every site will convert correctly. Pin and verify the version you deploy, then validate output against your own schema.

9. Troubleshooting

Symptom Likely cause Fix
Empty array The selector matches no table, or rows are inserted after load. Check the selector in developer tools; wait for rendering or use the page’s data source.
Headers become column_1 The table has no usable text in its header cells. Use an explicit schema or extract accessible labels and attributes.
Values appear under the wrong key rowspan/colspan changed the logical grid. Expand a span-aware occupancy grid or use a converter that supports spans.
Duplicate fields disappear Object keys overwrite earlier values. Suffix duplicate headings, store arrays, or reject ambiguous input.
Numbers are wrong Locale separators, currency symbols, or footnotes were parsed as plain numbers. Keep strings first; implement locale-aware parsing with validation.
HTML appears in output The parser returned markup instead of text. Use text extraction, or deliberately map links, attributes, and nested fields.
Only the first page is exported Pagination or virtual scrolling hides other rows. Iterate pages through the application interface or obtain the underlying dataset.
Server-side parser sees no rows The table is generated by browser JavaScript. Render it in a browser, call the underlying API, or export after rendering.

10. Validation checklist

  • Confirm the selected table by caption, ID, or surrounding heading.
  • Check that every data row has the expected logical column count.
  • Test duplicate, blank, and multi-row headings.
  • Test at least one rowspan and colspan example if the source uses them.
  • Define string, number, boolean, date, blank, and invalid-value rules.
  • Compare a sample of exported records with the visible table and source markup.
  • Record row-level errors instead of silently dropping data.
  • Escape or safely encode JSON before inserting it into another document.

11. Performance, reliability, and cost

DOM traversal is normally linear in the number of cells. The expensive parts are usually downloading the page, rendering client-side JavaScript, waiting for network requests, and handling large tables in memory. For large exports, stream rows to a file or process pages incrementally instead of building one enormous array.

Cache source HTML only when the data may be reused and its freshness permits it. Record the source URL, retrieval time, parser version, selector, and row errors so a later schema change is diagnosable. Never treat a successful HTTP response as proof that the table was complete.

12. Or skip the browser setup

If your goal is a visual capture of a page containing a table, ScreenshotNeo provides a single GET request for a PNG, JPEG, WebP, or PDF. It is a screenshot service rather than an HTML-to-JSON parser, so use the code below when you need the rendered page image for review, archiving, or an AI workflow.

See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify the page verdict and billing status.
  • An MCP server lets Claude, Cursor, and other MCP clients use take_screenshot, get_page_info, and capture_pdf.
  • The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account to get started.

13. FAQ

Can JSON contain duplicate keys?

Object member names should be unique for predictable use. Apply a suffix, array, or schema rule before serialization.

Should table values always be strings?

Preserving strings is safest until you define locale, currency, date, and error rules. Convert only fields whose format you control.

Can I convert a table loaded from a URL?

Yes, if you can retrieve the HTML and have permission to access it. A server-side parser will not see rows created later by browser JavaScript.

What is the safest way to handle a complex header?

Build header paths from the complete header hierarchy or supply an explicit schema. Do not infer keys from the last visual row alone.