ScreenshotNeo

BlogHow-to

HTML Table Capture with Ruby: Extract Rows, Handle Spans, and Export CSV

Capture HTML tables in Ruby with Nokogiri: select the right table, normalize spans, preserve UTF-8, export CSV, and automate screenshots.

By the ScreenshotNeo team29 September 20264 min read

HTML Table Capture with Ruby: Extract Rows, Handle Spans, and Export CSV

Direct answer: use Nokogiri to parse the HTML, scope your selector to the intended table, then iterate through each tr and extract its th and td cells. This captures the cells present in the DOM. If the table uses rowspan or colspan, add a normalization pass when you need a rectangular grid that matches the visual columns.

Nokogiri is the usual Ruby starting point because it parses HTML and supports both CSS and XPath queries. Its documentation describes the parser and search APIs in detail: Nokogiri HTML/XML parsing. The examples below work on local files, downloaded responses, or HTML returned by another service.

1. Install Nokogiri and choose your input

gem install nokogiri

In a Bundler project, add this to your Gemfile:

gem "nokogiri"

Then run:

bundle install

You can parse a file, a string, or an IO object. Keep the source encoding in mind: Nokogiri returns text as UTF-8, so verify non-ASCII values when the page contains accented characters, symbols, or non-Latin scripts.

require "nokogiri"

html = File.read("page.html", encoding: "UTF-8")
doc = Nokogiri::HTML(html)

# For a response body:
# doc = Nokogiri::HTML(response.body)

puts doc.at_css("title")&.text

2. Capture a table with CSS selectors

Start by selecting the exact table. A page may contain navigation tables, pricing tables, layout tables, and the data table you actually want. An ID is ideal; a distinctive class or a nearby heading is the next best option.

Nokogiri turns table markup into rows that Ruby can validate and export.
Nokogiri turns table markup into rows that Ruby can validate and export.
require "nokogiri"

html = File.read("page.html", encoding: "UTF-8")
doc = Nokogiri::HTML(html)

table = doc.at_css("table#results")
raise "table not found" unless table

rows = table.css("tr").map do |row|
  row.css("th, td").map { |cell| cell.text.strip }
end

p rows

Given a table with a header row, the result is an array of arrays:

[
  ["Name", "Status", "Total"],
  ["Ada", "Paid", "$42.00"],
  ["Lin", "Pending", "$18.50"]
]

cell.text includes descendant text, so links, spans, and formatted numbers are included. strip removes indentation and surrounding whitespace. For more precise cleanup, use cell.xpath(".//text()").map(&:text).join(" ") or normalize whitespace with gsub(/\s+/, " ").

3. Use XPath when the table needs a structural rule

CSS is readable for most selectors. XPath is useful when you need conditions such as “the table following this heading” or “rows containing a cell with a particular label.” Nokogiri supports both styles.

heading = doc.at_xpath("//h2[normalize-space()='Orders']")
raise "heading not found" unless heading

table = heading.at_xpath("following::table[1]")
raise "table not found" unless table

rows = table.xpath(".//tr").map do |row|
  row.xpath("./th | ./td").map { |cell| cell.text.gsub(/\s+/, " ").strip }
end

Notice the relative XPath expressions beginning with .. Without that scope, a query can accidentally select rows or cells from every table in the document.

4. Separate headers from data rows

Do not assume the first row is always a header. Prefer explicit thead and tbody sections, then fall back to the first row only when the source is inconsistent.

header_cells = table.css("thead tr").first&.css("th, td")
headers = header_cells&.map { |cell| cell.text.gsub(/\s+/, " ").strip }

body_rows = table.css("tbody tr")
body_rows = table.css("tr")[1..] if body_rows.empty? && table.css("tr").any?

records = body_rows.map do |row|
  values = row.css("th, td").map { |cell| cell.text.gsub(/\s+/, " ").strip }
  headers ? headers.zip(values).to_h : values
end

p records

Headers can be duplicated or blank. Before converting rows to hashes, make them unique and supply a fallback name:

def unique_headers(values)
  counts = Hash.new(0)
  values.map.with_index do |value, index|
    base = value.empty? ? "column_#{index + 1}" : value
    counts[base] += 1
    counts[base] == 1 ? base : "#{base}_#{counts[base]}"
  end
end

headers = unique_headers(headers || [])

5. Normalize rowspan and colspan into a rectangular grid

The simple loop returns the cells that exist in each DOM row. It does not expand a cell with rowspan="2" into both rows or place a colspan="3" value into three columns. If downstream code expects fixed columns, use a grid builder.

def table_grid(table)
  grid = []

  table.xpath(".//tr").each_with_index do |row, row_index|
    grid[row_index] ||= []
    column = 0

    row.xpath("./th | ./td").each do |cell|
      column += 1 while grid[row_index][column]

      text = cell.text.gsub(/\s+/, " ").strip
      rowspan = [cell["rowspan"].to_i, 1].max
      colspan = [cell["colspan"].to_i, 1].max

      rowspan.times do |row_offset|
        target_row = row_index + row_offset
        grid[target_row] ||= []
        colspan.times do |column_offset|
          grid[target_row][column + column_offset] = text
        end
      end

      column += colspan
    end
  end

  width = grid.map(&:length).max || 0
  grid.map { |row| row.fill(nil, row.length...width) }
end

grid = table_grid(table)

This produces a rectangular representation, but it is still a policy choice: repeating a spanning value may be correct for reporting and wrong for some semantic data models. Inspect representative tables before relying on the result.

6. Export captured rows as CSV

Use Ruby’s standard CSV library for quoting commas, quotes, and line breaks. Joining values with join(",") corrupts data as soon as a cell contains punctuation.

require "csv"

rows = table_grid(table)

CSV.open("table.csv", "wb", write_headers: false) do |csv|
  rows.each { |row| csv << row }
end

With headers and hashes:

require "csv"

CSV.open("orders.csv", "wb", write_headers: true, headers: headers) do |csv|
  records.each do |record|
    csv << headers.map { |header| record[header] }
  end
end

Ruby’s CSV::Table documentation covers header, row, and column operations. Extraction and CSV serialization are separate steps; keeping them separate makes validation easier.

7. Parse HTML5 safely and account for runtime differences

Nokogiri::HTML5 is documented as available since Nokogiri 1.12.0. The HTML5 API is not available on JRuby, and Nokogiri documents differences among parser implementations and runtimes. Record your Ruby version, Nokogiri version, parser choice, and representative fixture when reproducibility matters.

require "nokogiri"

html = File.read("page.html", encoding: "UTF-8")

doc = if defined?(Nokogiri::HTML5)
  Nokogiri::HTML5.parse(html)
else
  Nokogiri::HTML.parse(html)
end

Check your environment before depending on HTML5-only behavior:

puts RUBY_VERSION
puts Nokogiri::VERSION
puts RUBY_PLATFORM

Nokogiri treats input as untrusted by default and does not load external DTDs or access external resources during normal parsing. Keep those protections enabled for scraped or user-supplied HTML. Parser safety does not give permission to fetch a site or bypass its access controls.

8. When the table is rendered by JavaScript

Nokogiri parses the HTML you give it; it does not execute JavaScript. If the initial response contains an empty table and the browser fills it later, inspect the network response that carries the data, use an approved browser automation workflow, or capture the rendered page with a screenshot service.

Before changing tools, view the raw response and search for expected headers:

body = File.read("page.html", encoding: "UTF-8")
puts body.include?("Order ID")

If the data is in JSON, parse that JSON directly when possible. It is usually more reliable than scraping the browser’s final markup.

9. Or skip the browser setup

ScreenshotNeo can return a clean screenshot or PDF from one GET request, which is useful when you need a visual record of a rendered table instead of only its text values. Its capture options include full-page shots with lazy images loaded, element capture by CSS selector, custom JavaScript and CSS, waits, headers, cookies, user agents, blocking rules, caching, PDFs, and bulk capture. See the ScreenshotNeo API documentation for the parameter reference.

A rendered capture service can remove common overlays before saving the page.
A rendered capture service can remove common overlays before saving the page.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
await Bun.write("shot.webp", data);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

10. Troubleshooting checklist

“table not found”

Check the selector, confirm the table is present in the input string, and print a small fragment of the document. IDs may be generated dynamically, and the table may be inside an iframe or absent from the server response.

Rows are empty or missing

Inspect whether the source uses tbody, custom elements, or JavaScript rendering. Use table.xpath(".//tr") to test the structure, then scope back to direct children when you need strict semantics.

Columns shift between rows

Look for rowspan, colspan, hidden cells, and repeated header rows. Use the grid normalizer above and validate the expected column count.

Accented characters are corrupted

Read the source with the correct encoding, inspect the HTTP Content-Type charset, and verify the resulting Ruby strings with value.encoding. Do not blindly transcode unknown bytes.

JRuby raises an HTML5 API error

Use the supported parser for your runtime, or run the HTML5-specific code on CRuby. Pin and record Nokogiri versions so deployments use the same behavior.

CSV opens as one column

Write with Ruby’s CSV library, check the delimiter expected by the receiving spreadsheet, and avoid manually concatenating fields.

11. Performance, reliability, and cost notes

  • Parse once and reuse the document when extracting several tables.
  • Scope selectors early; searching one table is cheaper and less error-prone than scanning the whole document repeatedly.
  • For large files, avoid retaining unnecessary intermediate strings and write CSV rows incrementally.
  • Keep fixtures for tables with nested markup, missing cells, spans, duplicate headers, and non-ASCII text.
  • Validate row widths before loading data into a database.
  • For remote pages, network time dominates parsing time. Cache permitted responses and respect site policies.
  • With ScreenshotNeo, choose a cache TTL for repeat captures, use bulk capture for up to 100 URLs per call, and inspect X-Page-Verdict and X-Billed to reconcile usage.

12. FAQ

Can Nokogiri turn a table directly into hashes?

No single universal conversion exists because headers may be missing, duplicated, or span multiple columns. Extract rows first, normalize headers, then zip each row deliberately.

Should I use CSS or XPath?

Use CSS for stable IDs and classes. Use XPath when selection depends on relationships, text, or document order.

Does Nokogiri execute JavaScript?

No. Parse the server HTML or obtain the underlying data endpoint. Use a rendered capture workflow when the visual result is the requirement.

Extract attributes separately, such as cell.at_css("a")&.[]("href"). Text extraction alone discards markup semantics.

What is the safest default for untrusted HTML?

Keep Nokogiri’s default protections enabled and do not enable external DTDs, entity expansion, or network access merely to make malformed input parse.