Web Scraping with Html Agility Pack
Learn how to fetch HTML and extract reliable data with Html Agility Pack in C#, including XPath, malformed markup, errors, and JavaScript limits.
Html Agility Pack (HAP) parses HTML; it does not fetch pages or run a browser. A reliable scraper therefore has two separate stages: obtain the response with an HTTP client, then load that HTML into HAP and query its DOM with XPath. HAP builds a read/write DOM, supports XPath and XSLT, and is designed to tolerate malformed real-world markup. Check the current NuGet package metadata before publishing because versions and framework compatibility can change.
What Html Agility Pack can and cannot do
- Can: parse an HTML string, file, or stream into a DOM; query nodes with XPath; read and modify elements and attributes; and handle many forms of imperfect markup.
- Cannot by itself: make HTTP requests, execute JavaScript, render a browser viewport, solve CAPTCHAs, or bypass authentication and access controls.
If the values are present in the server response, HAP is often enough. If a page inserts them only after JavaScript runs, obtain the site’s documented API or use a rendering/browser service first, then parse the resulting HTML if needed.
1. Create a .NET project and install HAP
dotnet new console -n HapScraper
cd HapScraper
dotnet add package HtmlAgilityPack
Pin a version when your build requires reproducibility. At research time the NuGet listing showed 1.13.0 and included .NET 8.0 and .NET Standard 2.0 among its target frameworks; verify the registry for the version you intend to use.
2. Parse HTML and extract values with XPath
The following complete console program uses an inline document so it can run without a network dependency. It demonstrates safe node and attribute access, whitespace normalization, and validation.
using System;
using System.Collections.Generic;
using System.Linq;
using HtmlAgilityPack;
var html = """
<!doctype html>
<html>
<body>
<article class='product' data-id='sku-42'>
<h1> Example & Widget </h1>
<span class='price'>$19.99</span>
<a class='buy' href='/buy/sku-42'>Buy</a>
</article>
<!-- Deliberately imperfect markup is tolerated by HAP. -->
</body>
</html>
""";
var document = new HtmlDocument();
document.LoadHtml(html);
var product = document.DocumentNode.SelectSingleNode("//article[contains(concat(' ', normalize-space(@class), ' '), ' product ')]");
if (product is null)
throw new InvalidOperationException("Required product element was not found.");
var title = RequiredText(product.SelectSingleNode(".//h1"), "product title");
var priceText = RequiredText(product.SelectSingleNode(".//span[contains(@class, 'price')]"), "product price");
var id = RequiredAttribute(product, "data-id");
var buyUrl = product.SelectSingleNode(".//a[contains(@class, 'buy')]")?.GetAttributeValue("href", "");
Console.WriteLine($"Id: {id}");
Console.WriteLine($"Title: {title}");
Console.WriteLine($"Price: {priceText}");
Console.WriteLine($"Buy URL: {buyUrl}");
static string RequiredText(HtmlNode? node, string field)
{
var value = Normalize(node?.InnerText);
if (value.Length == 0)
throw new InvalidOperationException($"Missing or empty {field}.");
return value;
}
static string RequiredAttribute(HtmlNode node, string name)
{
var value = node.GetAttributeValue(name, "").Trim();
if (value.Length == 0)
throw new InvalidOperationException($"Missing attribute {name}.");
return value;
}
static string Normalize(string? value) =>
string.Join(" ", (value ?? "").Split((char[]?)null, StringSplitOptions.RemoveEmptyEntries));
HtmlNode.InnerText gives decoded text for the selected node. Normalize it before storing values because indentation and line breaks are part of the source markup. Use SelectNodes for collections and handle the possibility that it returns null.
3. Fetch the response before parsing it
Use an HttpClient that you reuse for multiple requests. Check the status code, set a reasonable timeout, and retain the response URL because redirects can change the document you received.
using System;
using System.Net.Http;
using System.Threading;
using System.Threading.Tasks;
using HtmlAgilityPack;
using var client = new HttpClient { Timeout = TimeSpan.FromSeconds(30) };
client.DefaultRequestHeaders.UserAgent.ParseAdd("HapScraper/1.0 (+https://example.invalid/contact)");
using var response = await client.GetAsync("https://example.com/", HttpCompletionOption.ResponseHeadersRead);
response.EnsureSuccessStatusCode();
var finalUrl = response.RequestMessage?.RequestUri?.ToString();
var html = await response.Content.ReadAsStringAsync();
var document = new HtmlDocument();
document.LoadHtml(html);
var heading = document.DocumentNode.SelectSingleNode("//h1")?.InnerText.Trim();
Console.WriteLine($"Fetched: {finalUrl}");
Console.WriteLine($"Heading: {heading ?? "(missing)"}");
Replace the example user agent and URL with values appropriate for your application. Follow the target site’s terms, robots policy, authentication requirements, and applicable law.
4. XPath patterns you will use often
| Need | XPath | Notes |
|---|---|---|
| First heading | //h1 |
Use SelectSingleNode. |
| All product cards | //article[contains(concat(' ', normalize-space(@class), ' '), ' product ')] |
The class test avoids matching product-card-old. |
| Descendant price | .//span[contains(@class, 'price')] |
The leading dot keeps the search inside the current node. |
| Exact data attribute | //div[@data-id='sku-42'] |
Escape or construct values carefully when generating XPath. |
| Links with href | //a[@href] |
Read with GetAttributeValue. |
| Text containing a phrase | //p[contains(normalize-space(.), 'shipping')] |
Whitespace normalization improves matching. |
var cards = document.DocumentNode.SelectNodes("//article[contains(concat(' ', normalize-space(@class), ' '), ' product ')]") ?? new HtmlNodeCollection(null);
foreach (var card in cards)
{
var name = Normalize(card.SelectSingleNode(".//h2")?.InnerText);
var href = card.SelectSingleNode(".//a[@href]")?.GetAttributeValue("href", "").Trim();
if (name.Length == 0 || href.Length == 0)
continue;
Console.WriteLine($"{name}: {href}");
}
For a simpler collection pattern, avoid constructing an empty collection and use a nullable result:
var nodes = document.DocumentNode.SelectNodes("//article") ?? Enumerable.Empty<HtmlNode>();
foreach (var node in nodes)
Console.WriteLine(node.InnerText.Trim());
5. Load files and streams
var document = new HtmlDocument();
document.Load("page.html");
using var stream = File.OpenRead("page.html");
var streamed = new HtmlDocument();
streamed.Load(stream);
For downloaded bytes with uncertain encoding, prefer the HTTP response’s declared charset when available. If you already decoded the response to a string, pass that string to LoadHtml.
6. Make extraction resilient
Missing nodes and attributes
SelectSingleNode can return null; an absent attribute returns the default supplied to GetAttributeValue. Decide field by field whether absence is an error, a nullable value, or a reason to skip the record.
var image = card.SelectSingleNode(".//img");
var imageUrl = image?.GetAttributeValue("data-src", null)
?? image?.GetAttributeValue("src", null);
if (string.IsNullOrWhiteSpace(imageUrl))
Console.WriteLine("Card has no image URL");
Relative URLs
if (Uri.TryCreate(new Uri("https://example.com/catalog/"), href, out var absolute))
Console.WriteLine(absolute);
Duplicate or changed markup
Prefer stable attributes such as data-* values. Keep selectors in one place, record the source URL and retrieval time, and validate required fields so a layout change fails visibly instead of silently producing incorrect data.
Tables and entities
Iterate rows, then cells, and normalize each cell independently. HAP handles common entities while parsing; validate the final value for the data type your application expects.
7. Validate against the real response
An XPath that works on a copied example can fail when the server returns a consent page, an error document, a different locale, or a logged-out variant. Save a small redacted response fixture and check:
- HTTP status and final URL;
- content type and character encoding;
- presence of a page-specific marker;
- minimum record count;
- required attributes and parseable numeric or date values.
if (!html.Contains("<article", StringComparison.OrdinalIgnoreCase))
throw new InvalidOperationException("Unexpected response: product markup is absent.");
var products = document.DocumentNode.SelectNodes("//article") ?? new HtmlNodeCollection(null);
if (products.Count == 0)
throw new InvalidOperationException("No products found; inspect the saved response.");
8. JavaScript-rendered pages
HAP sees only the HTML supplied to LoadHtml. It does not execute scripts or wait for client-side requests. Inspect the raw response first. If the data is loaded from a JSON endpoint, use that endpoint when permitted. Otherwise, use a browser or rendering service to produce HTML or a screenshot, then process the result. Treat a successful HTTP status as insufficient evidence that the desired content exists.
9. CSS selectors, HTML5 parsing, and alternatives
HAP’s documented query model is XPath. A separate Universal.HtmlAgilityPack package advertises CSS selector support by converting selectors to XPath. AngleSharp is another option when HTML5 specification-oriented parsing and CSS selectors are central requirements. Compare libraries against your actual input, target frameworks, selector style, and maintenance needs; the available research does not establish a universal performance or accuracy winner.
10. cURL, Python, and Node.js fetch examples
These snippets show the separate acquisition step. They save the response for inspection; HAP remains the C# parsing component.
curl -L --fail --compressed \
-A 'HapScraper/1.0' \
'https://example.com/' \
-o page.html
import requests
response = requests.get(
"https://example.com/",
headers={"User-Agent": "HapScraper/1.0"},
timeout=30,
)
response.raise_for_status()
with open("page.html", "w", encoding=response.encoding or "utf-8") as file:
file.write(response.text)
const response = await fetch('https://example.com/', {
headers: { 'User-Agent': 'HapScraper/1.0' },
signal: AbortSignal.timeout(30_000),
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
await Bun.write('page.html', await response.text());
11. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
SelectSingleNode returns null |
Selector does not match this response, namespace or class assumptions are wrong, or content is absent. | Save and inspect the exact response; simplify the XPath; check the final URL and page variant. |
Empty SelectNodes result |
Collection is absent and the result is null. | Use ?? Enumerable.Empty<HtmlNode>() and decide whether zero items is valid. |
| Expected text is blank | Whitespace-only text, wrong node, or value inserted by JavaScript. | Normalize InnerText, inspect the raw HTML, and check the site’s data endpoint. |
| HTTP 403 or 429 | Server policy, authentication, rate limit, or bot protection. | Follow the site’s access rules, authenticate as documented, slow requests, and do not attempt to bypass controls. |
| HTTP 200 but no records | Consent, login, error, or challenge page. | Validate a page marker and inspect the saved response before parsing. |
| Garble or broken characters | Incorrect response decoding or malformed source encoding. | Use the declared charset, preserve response bytes when diagnosing, and test with a fixture containing non-ASCII text. |
| Works locally, fails in production | Different cookies, headers, locale, redirects, network policy, or HTML variant. | Log status, final URL, content type, and a safe response fingerprint; make those inputs explicit. |
12. Performance, reliability, and cost
- Reuse clients: keep one
HttpClientrather than creating one per URL. - Bound work: set request timeouts, cancellation tokens, response-size limits, and concurrency limits.
- Parse once: load each response into one
HtmlDocumentand run the required XPath queries against it. - Cache carefully: cache responses only when freshness and authorization rules allow it.
- Retry selectively: retry transient network failures and selected 5xx responses with backoff; do not blindly retry 4xx responses.
- Measure your workload: HAP’s source material does not provide a benchmark, so profile your document sizes, selector count, and network time instead of assuming a speed ranking.
- Control cost: HTTP bandwidth, proxy or browser services, storage, and downstream processing can dominate the cost. A parser library does not make those resources free.
Or skip the browser setup
If your goal is a clean image or PDF of a page rather than DOM fields, ScreenshotNeo provides a single GET request. It accepts the cookie or consent banner like a visitor, removes more than 60 known consent platforms plus newsletter popups and chat widgets before capture, and lets each cleanup step be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.
See the ScreenshotNeo API documentation for the available options, including full-page capture, CSS element selection, device presets, dark mode, custom CSS and JavaScript, waits, blocking rules, cookies and headers, PDFs, signed links, asynchronous jobs, bulk capture, and usage information.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also has an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Does HAP download a website?
No. Fetch the response with HttpClient or another HTTP client, then give the HTML to HAP.
Does HAP execute JavaScript?
No. Use an API endpoint or a browser/rendering step when the required content is created after page load.
Is malformed HTML a problem?
HAP is designed to tolerate malformed real-world HTML, but your XPath still needs to match the actual DOM it builds.
Should I use XPath or CSS selectors?
Use XPath when it fits your team and document structure. Consider a CSS-to-XPath add-on or AngleSharp when CSS selectors and HTML5-oriented behavior are requirements.
Can I assume a selector will keep working?
No. Validate against saved response fixtures and monitor required fields whenever the target site changes.


