Web Scraping in C#: From Basics to Production-Ready Code in 2026
Learn how to scrape static and JavaScript pages in C# with HttpClient, HtmlAgilityPack, AngleSharp and Playwright, then harden the scraper for production.
Direct answer: use HttpClient plus an HTML parser when the data is present in the server response. Use Playwright for .NET when JavaScript must run, a browser interaction is required, or the content appears only after rendering. A production scraper also needs connection reuse, cancellation, bounded concurrency, response validation, resilient selectors, observability and an explicit policy for robots.txt, terms and authorization.
This guide answers “How do I scrape a website with C#?” with runnable .NET examples, then covers the operational decisions that matter when a one-off script becomes a service.
1. How do I scrape a website with C#?
- Identify whether the required content is in the initial HTML response.
- Fetch the page with a reused
HttpClient. - Check cancellation, status code, content type and response size.
- Parse with Html Agility Pack (XPath) or AngleSharp (DOM and CSS selectors).
- Extract optional fields defensively and validate required fields.
- Persist normalized records and log parsing failures.
Install a parser in a new console project:
dotnet new console -n CSharpScraper
cd CSharpScraper
dotnet add package HtmlAgilityPack
The following complete example fetches a page, selects links, normalizes text and writes JSON. Replace the URL and selectors for the target site.
using System.Net;
using System.Net.Http.Headers;
using System.Text.Json;
using HtmlAgilityPack;
var target = new Uri("https://example.com/");
using var handler = new SocketsHttpHandler
{
PooledConnectionLifetime = TimeSpan.FromMinutes(5),
AutomaticDecompression = DecompressionMethods.All
};
using var http = new HttpClient(handler)
{
Timeout = TimeSpan.FromSeconds(60)
};
http.DefaultRequestHeaders.UserAgent.ParseAdd("ExampleResearchBot/1.0 (+https://example.com/contact)");
using var response = await http.GetAsync(target, HttpCompletionOption.ResponseHeadersRead);
response.EnsureSuccessStatusCode();
if (response.Content.Headers.ContentLength is long length && length > 10_000_000)
throw new InvalidDataException("Response is larger than the configured limit.");
var html = await response.Content.ReadAsStringAsync();
var doc = new HtmlDocument();
doc.LoadHtml(html);
var rows = doc.DocumentNode
.SelectNodes("//a[@href]")?
.Select(node => new
{
Text = HtmlEntity.DeEntitize(node.InnerText).Trim(),
Href = node.GetAttributeValue("href", "")
})
.Where(x => x.Text.Length > 0 && Uri.TryCreate(target, x.Href, out _))
.ToArray() ?? Array.Empty<object>();
await File.WriteAllTextAsync("links.json", JsonSerializer.Serialize(rows, new JsonSerializerOptions { WriteIndented = true }));
Console.WriteLine($"Extracted {rows.Length} links.");
HttpClient owns a connection pool. Microsoft recommends either a long-lived client with PooledConnectionLifetime or clients created by IHttpClientFactory; creating and disposing a client for every request is not the recommended pattern. See Microsoft’s HttpClient guidelines.
2. Should I use HttpClient, HtmlAgilityPack, AngleSharp, or Playwright?
| Tool | Use it when | Selection model | Operational cost |
|---|---|---|---|
| HttpClient | You need the HTTP response | None; pair with a parser | Low; connection pooling |
| Html Agility Pack | You want a forgiving HTML parser and XPath | XPath and node traversal | Low |
| AngleSharp | You prefer a standards-oriented DOM and CSS selectors | CSS selectors and DOM APIs | Low |
| Playwright for .NET | JavaScript, browser state or interaction is required | Locators, selectors and page APIs | Higher; browser binaries and runtime |
There is no supported benchmark here that makes one parser universally faster or better. Choose based on the target markup, selector style, document quirks and team familiarity. Microsoft’s ASP.NET Core integration-test example names both AngleSharp and Html Agility Pack as parser options, but it is not a current comparative benchmark.
3. Fetching reliably with HttpClient
Long-lived client or IHttpClientFactory
A long-lived client can retain a DNS-resolved endpoint. PooledConnectionLifetime causes connections to be replaced so DNS can be resolved again. Microsoft’s 15-minute value is illustrative, not a universal production setting. Select a lifetime that matches the DNS behavior of the service you call.
builder.Services.AddHttpClient<ScrapeClient>(client =>
{
client.Timeout = TimeSpan.FromSeconds(60);
client.DefaultRequestHeaders.UserAgent.ParseAdd("CatalogBot/1.0");
});
Timeouts, cancellation and status codes
public sealed class ScrapeClient(HttpClient http)
{
public async Task<string> GetHtmlAsync(Uri uri, CancellationToken cancellationToken)
{
using var request = new HttpRequestMessage(HttpMethod.Get, uri);
using var response = await http.SendAsync(
request, HttpCompletionOption.ResponseHeadersRead, cancellationToken);
if ((int)response.StatusCode == 429 || (int)response.StatusCode >= 500)
throw new HttpRequestException($"Transient HTTP status {(int)response.StatusCode}");
response.EnsureSuccessStatusCode();
return await response.Content.ReadAsStringAsync(cancellationToken);
}
}
Bound concurrency and pace requests according to the target site’s capacity and instructions. Do not apply a universal retry count or delay. Retry only failures for which repeating the operation is appropriate, and avoid retrying authentication failures, malformed requests or deterministic parser errors.
Redirects, headers and content limits
Decide whether redirects are acceptable, preserve an explicit user agent, and reject unexpectedly large responses before loading them into memory. Treat content type as a signal rather than proof: servers sometimes mislabel HTML. Record final URLs after redirects when provenance matters.
4. Parsing HTML safely
Html Agility Pack with XPath
var doc = new HtmlDocument();
doc.LoadHtml(html);
var title = doc.DocumentNode.SelectSingleNode("//title")?.InnerText.Trim();
var priceNode = doc.DocumentNode.SelectSingleNode("//span[contains(@class,'price')]");
var price = priceNode is null ? null : HtmlEntity.DeEntitize(priceNode.InnerText).Trim();
AngleSharp with CSS selectors
dotnet add package AngleSharp
using AngleSharp;
var context = BrowsingContext.New(Configuration.Default);
var page = await context.OpenAsync(req => req.Content(html).Address(target.ToString()));
var heading = page.QuerySelector("h1")?.TextContent.Trim();
var links = page.QuerySelectorAll("a[href]")
.Select(a => new { Text = a.TextContent.Trim(), Href = a.GetAttribute("href") })
.Where(x => !string.IsNullOrWhiteSpace(x.Href));
Make extraction tolerant of redesigns
- Use stable attributes or semantic structure instead of generated class names.
- Handle missing nodes as normal input, not as an unhandled null reference.
- Normalize whitespace, currency and dates before persistence.
- Validate required fields and record the URL, selector and parser error when validation fails.
- Keep fixture HTML for representative page variants so selector changes are reviewable.
5. Can C# scrape JavaScript-rendered pages?
Yes, with browser automation. If the required value is absent from the ordinary HTTP response and inserted by JavaScript, an HTML parser cannot execute that JavaScript. Playwright’s official .NET port automates Chromium, Firefox and WebKit. Its APIs expose request, response, completion and failure events; an HTTP 404 or 503 can still be a completed browser request, so inspect both HTTP status and page state.
dotnet add package Microsoft.Playwright
playwright install
using Microsoft.Playwright;
using var playwright = await Playwright.CreateAsync();
await using var browser = await playwright.Chromium.LaunchAsync(new BrowserTypeLaunchOptions
{
Headless = true
});
var page = await browser.NewPageAsync();
page.Response += (_, response) =>
Console.WriteLine($"{response.Status} {response.Url}");
await page.GotoAsync("https://example.com/app", new PageGotoOptions
{
WaitUntil = WaitUntilState.NetworkIdle,
Timeout = 60_000
});
await page.Locator("[data-product-name]").WaitForAsync();
var name = await page.Locator("[data-product-name]").InnerTextAsync();
Console.WriteLine(name);
Use a browser only when rendering or interaction is necessary. Browser processes and binaries add deployment, memory and startup work. If a page calls a public JSON endpoint, inspect requests during development and consider calling that documented endpoint directly when the site’s terms and authorization permit it.
6. Browser interactions, sessions and dynamic data
Interactions belong in Playwright: accept a consent control when appropriate, fill a form, click a pagination button, wait for a locator, then extract the resulting DOM. Keep session state isolated per job when cookies or authentication are involved. Never put credentials in source code or logs.
await page.GetByRole(AriaRole.Button, new() { Name = "Load more" }).ClickAsync();
await page.Locator("article").Last.WaitForAsync();
var cards = await page.Locator("article").AllInnerTextsAsync();
7. Is robots.txt permission to scrape?
No. RFC 9309 defines the Robots Exclusion Protocol as crawler instructions. The IETF standard states: “These rules are not a form of access authorization.” Review the target’s terms, authorization requirements, access controls, copyright and privacy obligations for your jurisdiction and use case.
- Retrieve and parse
/robots.txtaccording to RFC 9309 groups and matching rules. - A successful retrieval means a crawler that honors REP should follow the applicable rules.
- The standard says cached files generally should not be used for more than 24 hours unless the file is unreachable.
- A server or network error that makes robots.txt unreachable requires assuming complete disallow under the protocol; a 4xx unavailable response has different treatment.
- RFC 9309 specifies a 500 KiB minimum parsing limit and discusses a 30-day example for unavailable or cached files. These are protocol details, not request-rate recommendations.
Read RFC 9309 and do not treat a missing or malformed file as blanket permission.
8. Production architecture checklist
- Queue: accept jobs and apply bounded concurrency per host.
- Fetch: reuse clients, set cancellation and timeout policies, and capture status and final URL.
- Parse: isolate selectors in versioned code and validate required fields.
- Persist: use idempotent keys, source timestamps and a schema that records missing values.
- Observe: track request duration, status classes, bytes, retries, parse failures and output counts without logging secrets.
- Protect: limit response sizes, validate URLs to prevent internal-network access, and restrict outbound protocols where appropriate.
- Operate: support cancellation, graceful shutdown and replay of failed jobs.
9. Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Socket exhaustion | A new HttpClient is created per request | Use a long-lived client with a suitable connection lifetime or IHttpClientFactory. |
| Old IP after DNS change | Existing pooled connection retained the old endpoint | Configure PooledConnectionLifetime or factory handler rotation. |
| Empty selector result | Content is JavaScript-rendered or markup changed | Inspect the raw response; update selectors or use Playwright. |
| 403 or 429 | Access policy, authentication or rate limiting | Check authorization and terms, identify yourself, reduce concurrency and honor retry guidance. |
| Timeout | Slow origin, blocked resource or browser wait condition | Capture timing and failed requests, use cancellation, and wait for a specific locator instead of an arbitrary delay. |
| Malformed output | Missing nodes or changed text formats | Normalize defensively, validate required fields and retain failing fixtures. |
| Playwright launch failure | Browser binaries are absent or unavailable in the deployment image | Install the required browser during image build and verify sandbox/runtime dependencies. |
10. Performance, reliability and cost
HTTP plus a parser is usually the simpler operational path because it avoids browser startup and rendering. Browser automation costs more CPU, memory and deployment effort, so reserve it for pages that need it. Measure your own target: no universal concurrency, timeout, retry count or parser benchmark is established by the sources used here.
Reuse connections, stream large responses where practical, avoid downloading unnecessary resources, and cache data only when freshness requirements allow. Retries improve recovery from transient failures but can amplify load; use bounded, observable policies and make writes idempotent. A scraper’s main cost is often engineering and browser runtime rather than the parser package itself.
11. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. It can capture a URL as PNG, JPEG, WebP or PDF with one GET request, including pages that need browser rendering.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for options. Cookie banners, newsletter popups and chat widgets are removed before the shot. Bot checks, blank pages and failed loads are never billed, and response headers identify the page verdict and billing status. An MCP server lets Claude, Cursor and other MCP clients call take_screenshot, get_page_info and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.
Create a free ScreenshotNeo account.
12. FAQ
Which C# library should I use for web scraping?
Use HttpClient for transport, then choose Html Agility Pack for XPath-oriented extraction or AngleSharp for CSS and DOM APIs. Use Playwright when a real browser is required.
Can I scrape a site without JavaScript?
Yes. If the required markup is in the HTTP response, a parser is enough and is usually simpler to operate.
How do I know whether a page is JavaScript-rendered?
Compare the raw response with the browser’s rendered DOM. If the value exists only after scripts run, use browser automation or an authorized underlying data endpoint.
Should I retry every failed request?
No. Retry only transient failures and operations that are safe to repeat, with bounded concurrency and logging.
Does robots.txt make scraping legal?
No. It is a crawler protocol, not authorization. Evaluate the site’s terms, authorization, access controls and applicable law.


