How to Import an Existing HTML File in Rust
Read an HTML file, parse documents or fragments, extract with CSS selectors, handle encoding and errors, and choose the right Rust crate.
Importing an existing HTML file in Rust has two separate steps: read the file, then parse the resulting text (or bytes). For selector-based extraction, scraper is the simplest high-level choice. Use Html::parse_document for a complete page and Html::parse_fragment for a snippet. Choose Kuchiki when you need a mutable DOM-like tree; use html5ever directly only when you need its lower-level WHATWG parser callbacks.
This guide shows each approach, complete Cargo projects, encoding choices, common failures, and a ScreenshotNeo option when the HTML you need is a live website rather than a local file.
1. Read and parse a complete HTML file with scraper
Add the crate:
[dependencies]
scraper = "0.20"
Create page.html:
<!doctype html>
<html>
<head><title>Rust docs</title></head>
<body>
<main id="content">
<h1>Importing HTML</h1>
<a class="docs" href="/guide">Guide</a>
</main>
</body>
</html>
Then compile and run:
use scraper::{Html, Selector};
use std::{error::Error, fs};
fn main() -> Result<(), Box<dyn Error>> {
let html = fs::read_to_string("page.html")?;
let document = Html::parse_document(&html);
let title_selector = Selector::parse("title")?;
if let Some(title) = document.select(&title_selector).next() {
let text = title.text().collect::<String>();
println!("title: {text}");
}
let link_selector = Selector::parse("main#content a.docs")?;
for link in document.select(&link_selector) {
let label = link.text().collect::<String>();
let href = link.value().attr("href").unwrap_or("");
println!("{label} -> {href}");
}
Ok(())
}
read_to_string reads the entire file and requires valid UTF-8. It returns an I/O error for a missing path, permissions failure, or invalid UTF-8; the ? operator propagates that error instead of silently producing partial data. See the Rust read_to_string documentation.
2. Parse an HTML fragment
A fragment is markup such as a table row or list item without <html> and <body>. Parse it explicitly:
use scraper::{Html, Selector};
fn main() -> Result<(), Box<dyn std::error::Error>> {
let fragment = Html::parse_fragment("<li>One</li><li>Two</li>");
let item = Selector::parse("li")?;
for node in fragment.select(&item) {
println!("{}", node.text().collect::<String>());
}
Ok(())
}
parse_document may insert implied document structure while fragment parsing keeps the input in a fragment context. The scraper API documents both entry points and selector traversal at docs.rs/scraper.
3. Extract attributes, text, and serialized HTML
use scraper::{Html, Selector};
fn main() -> Result<(), Box<dyn std::error::Error>> {
let document = Html::parse_document(
r#"<article data-id="42"><h2>Hello</h2><p>Body</p></article>"#,
);
let article = Selector::parse("article")?;
for node in document.select(&article) {
let id = node.value().attr("data-id").unwrap_or_default();
let text = node.text().collect::<Vec<_>>().join(" ");
println!("id={id}, text={text}");
println!("inner/outer serialization: {node}");
}
Ok(())
}
Use CSS selectors for descendants, classes, IDs, attributes, and structural matches. Parse a selector once and reuse it in loops. If Selector::parse fails, fix the selector string; do not unwrap user-provided selectors in a long-running service.
4. When to use Kuchiki
Kuchiki builds an html5ever-backed, DOM-like tree that is convenient when you must inspect and mutate nodes. Its parse_html handles full documents and parse_fragment handles fragments. A minimal example:
[dependencies]
kuchiki = "0.8"
use kuchiki::traits::*;
use std::{error::Error, fs};
fn main() -> Result<(), Box<dyn Error>> {
let html = fs::read_to_string("page.html")?;
let document = kuchiki::parse_html().one(html);
for css_match in document.select("main h1")? {
let mut text = css_match.text_contents();
text.push_str(" (imported)");
css_match.as_node().children().detach();
css_match.as_node().append(kuchiki::NodeRef::new_text(text));
}
print!("{document}");
Ok(())
}
Use scraper when you mainly query. Use Kuchiki when edits, node ownership, or DOM-style traversal are central. Both rely on html5ever parsing behavior; Kuchiki provides the tree around it.
5. Lower-level html5ever
html5ever parses and serializes HTML according to WHATWG HTML5 specifications, but it exposes callbacks rather than a ready-made DOM tree. That makes it suitable for custom streaming or tree-building code, with more implementation work than scraper or Kuchiki.
6. Handle non-UTF-8 files deliberately
When the file may use an unknown encoding, read bytes first. std::fs::read returns a complete Vec<u8> without a UTF-8 check; decode according to your input contract.
use std::{error::Error, fs};
fn main() -> Result<(), Box<dyn Error>> {
let bytes = fs::read("page.html")?;
let html = String::from_utf8(bytes)?; // strict UTF-8
println!("{} bytes of valid HTML", html.len());
Ok(())
}
For lossy fallback (only when replacement characters are acceptable), use String::from_utf8_lossy. For Windows-1252 or another declared encoding, decode with an appropriate encoding crate before passing text to the parser. Never guess silently when extracted text is used for signatures, IDs, or compliance records. See the Rust read documentation.
7. A reusable importer function
use scraper::{Html, Selector};
use std::{error::Error, path::Path};
pub fn page_title(path: impl AsRef<Path>) -> Result<Option<String>, Box<dyn Error>> {
let html = std::fs::read_to_string(path)?;
let document = Html::parse_document(&html);
let selector = Selector::parse("title")?;
Ok(document
.select(&selector)
.next()
.map(|node| node.text().collect::<String>().trim().to_owned()))
}
fn main() -> Result<(), Box<dyn Error>> {
println!("{:?}", page_title("page.html")?);
Ok(())
}
For many files, walk paths outside this function, bound file size before reading, and process one file at a time. Parsing creates an in-memory tree, so a very large document can exceed available memory.
8. Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
No such file or directory |
Relative path is resolved from the process working directory. | Print std::env::current_dir(), use an absolute/configured path, and verify deployment layout. |
stream did not contain valid UTF-8 |
read_to_string rejects non-UTF-8 bytes. |
Use fs::read and decode with the file's declared encoding. |
| Selector returns no nodes | Wrong selector, fragment context, or malformed source. | Log a small input sample, validate the selector, and choose parse_fragment for snippets. |
| Expected elements are missing | The file contains a template, script-generated content, or an iframe; Rust parsers do not execute JavaScript. | Obtain rendered HTML from the producing system, or capture the live page with a browser-capable service. |
| Process runs out of memory | Whole-file loading plus a DOM tree multiplies memory use. | Reject oversized files, stream your own preprocessing, or parse only needed sections. |
| Mutation is awkward in scraper | Scraper is optimized for querying, not general DOM editing. | Switch to Kuchiki for tree mutation, or generate a transformed output explicitly. |
9. Performance, reliability, and cost
- Performance: read once, parse once, and reuse compiled selectors. Avoid collecting every node's text when an attribute is enough.
- Reliability: propagate I/O and selector errors, cap input size, and record the source path and encoding decision. Add fixtures for malformed HTML and missing elements.
- Cost: local parsing has no API charge; its practical limits are CPU, memory, and storage. Browser rendering of a live URL adds network and browser startup costs.
10. Or skip the browser setup
If your actual input is a live URL and you need a clean image or PDF rather than the source DOM, ScreenshotNeo provides a single GET request. It accepts consent banners before capture and removes 60+ known consent platforms, newsletter popups, and chat widgets. Only clean shots are billed; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Responses identify the page verdict and billing with X-Page-Verdict and X-Billed headers. See the ScreenshotNeo API docs.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also has an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. It supports full-page and element captures, custom CSS/JavaScript, waits, blocking rules, devices, retina scale, PDF controls, caching, signed links, async jobs, bulk capture, and usage reporting. Plans include 1,000 shots/month free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
11. FAQ
Does Rust execute JavaScript in the HTML file?
No. scraper, Kuchiki, and html5ever parse markup; they do not run scripts or fetch iframe content.
Should I use scraper or Kuchiki?
Choose scraper for read-only CSS selection and text extraction. Choose Kuchiki when you need to mutate a DOM-like tree.
Can I parse a partial snippet?
Yes. Pass it to Html::parse_fragment and select within the returned fragment.
What if the file is not UTF-8?
Read bytes with std::fs::read, then decode using the known source encoding before parsing.


