ScreenshotNeo

BlogEngineering

How to Build Optimized Web Scraping Actors in Go

Build a concurrent, resource-bounded Go scraping worker with reusable HTTP connections, measured tuning, profiling, and practical troubleshooting.

By the ScreenshotNeo team29 September 202611 min read

How to Build Optimized Web Scraping Actors in Go

A useful Go scraping actor is a bounded worker pipeline: it accepts jobs, fetches pages through a shared HTTP client, extracts structured data, and reports results. Optimize it by measuring a representative crawl, then changing one bottleneck at a time. Reusing a client and transport is the right default; increasing goroutines without limits is not.

This guide builds a runnable static-HTML actor using Go’s standard library. It bounds queued and active work, applies per-request timeouts, closes response bodies, handles HTTP failures, and emits JSON Lines. The right worker count, retry policy, per-host rate limit, parser, and rendering approach depend on the crawl and are not universal constants.

1. Define the actor’s workload

Before tuning, decide what a job means and what success means. A job might contain a URL and an identifier; a result might contain the final URL, status, title, extracted fields, and an error. Define whether redirects are acceptable, which status codes are useful, what output rate is needed, and which resource limits matter.

Also distinguish static HTML from JavaScript-rendered pages. An HTTP client downloads the response body; it does not execute page JavaScript. If the required data appears only after client-side rendering, use an appropriate browser-rendering system or an API that captures rendered pages. Do not hide this difference behind a larger Go worker pool.

2. Build a bounded fetch-and-parse pipeline

The example uses a fixed number of workers and a bounded jobs channel. Producers block when the queue is full, providing backpressure instead of allowing an input burst to create unbounded pending work. One shared client and transport serve all goroutines. Workers parse a title without an external dependency so the example stays runnable with the standard library.

A bounded queue and fixed worker pool keep intake bursts from creating unbounded concurrent work.
A bounded queue and fixed worker pool keep intake bursts from creating unbounded concurrent work.
package main

import (
 "context"
 "encoding/json"
 "flag"
 "fmt"
 "io"
 "net/http"
 "os"
 "strings"
 "sync"
 "time"
)

type Result struct {
 URL string `json:"url"`
 Status int `json:"status,omitempty"`
 Title string `json:"title,omitempty"`
 Bytes int `json:"bytes,omitempty"`
 Error string `json:"error,omitempty"`
}

func titleFromHTML(s string) string {
 lower := strings.ToLower(s)
 start := strings.Index(lower, "<title")
 if start < 0 { return "" }
 endOpen := strings.Index(s[start:], ">")
 if endOpen < 0 { return "" }
 from := start + endOpen + 1
 end := strings.Index(strings.ToLower(s[from:]), "</title>")
 if end < 0 { return "" }
 return strings.TrimSpace(s[from:from+end])
}

func fetch(ctx context.Context, client *http.Client, rawURL string, maxBody int64) Result {
 result := Result{URL: rawURL}
 req, err := http.NewRequestWithContext(ctx, http.MethodGet, rawURL, nil)
 if err != nil { result.Error = err.Error(); return result }
 req.Header.Set("User-Agent", "example-scraper/1.0 (+contact: ops@example.com)")
 resp, err := client.Do(req)
 if err != nil { result.Error = err.Error(); return result }
 defer resp.Body.Close()
 result.Status = resp.StatusCode
 if resp.StatusCode < 200 || resp.StatusCode >= 300 {
  result.Error = resp.Status
  return result
 }
 body, err := io.ReadAll(io.LimitReader(resp.Body, maxBody+1))
 if err != nil { result.Error = err.Error(); return result }
 if int64(len(body)) > maxBody { result.Error = "response body exceeds configured limit"; return result }
 result.Bytes = len(body)
 result.Title = titleFromHTML(string(body))
 return result
}

func main() {
 workers := flag.Int("workers", 8, "maximum concurrent fetch workers")
 queueSize := flag.Int("queue", 32, "maximum queued URLs")
 timeout := flag.Duration("timeout", 20*time.Second, "timeout per request")
 maxBody := flag.Int64("max-body", 4<<20, "maximum response bytes")
 flag.Parse()
 if *workers < 1 || *queueSize < 0 || *maxBody < 1 { fmt.Fprintln(os.Stderr, "workers and max-body must be positive; queue cannot be negative"); os.Exit(2) }

 tr := http.DefaultTransport.(*http.Transport).Clone()
 tr.MaxIdleConns = 100
 tr.MaxIdleConnsPerHost = 8
 tr.MaxConnsPerHost = 16
 tr.IdleConnTimeout = 90 * time.Second
 client := &http.Client{Transport: tr, Timeout: *timeout}
 defer tr.CloseIdleConnections()

 jobs := make(chan string, *queueSize)
 results := make(chan Result)
 var wg sync.WaitGroup
 for i := 0; i < *workers; i++ {
  wg.Add(1)
  go func() {
   defer wg.Done()
   for url := range jobs { results <- fetch(context.Background(), client, url, *maxBody) }
  }()
 }
 go func() { wg.Wait(); close(results) }()
 go func() {
  defer close(jobs)
  scanner := os.Stdin
  data, err := io.ReadAll(scanner)
  if err != nil { fmt.Fprintln(os.Stderr, err); return }
  for _, line := range strings.Split(string(data), "\n") {
   url := strings.TrimSpace(line)
   if url != "" { jobs <- url }
  }
 }()
 enc := json.NewEncoder(os.Stdout)
 for result := range results {
  if err := enc.Encode(result); err != nil { fmt.Fprintln(os.Stderr, err); os.Exit(1) }
 }
}

Save as main.go, then run go run main.go -workers 8 -queue 32 < urls.txt. Each non-empty input line is fetched and produces one JSON object on stdout. In a production actor, make intake cancellation-aware and stream jobs from the actual queue or scheduler rather than reading all stdin into memory. The sample reads stdin as a whole to keep the concurrency flow compact; for a large feed, use a scanner or streaming decoder and send each parsed URL as it arrives.

Make the extraction parser appropriate

The small title function demonstrates the boundary between fetch and extraction, but it is not an HTML parser: malformed markup, unusual encoding, entities, and title tags with edge-case syntax can defeat string scanning. For real extraction, choose an HTML parser or scraping library that fits the project, parse the response content type and character encoding as needed, and treat missing fields as normal data rather than necessarily as a failed fetch.

Keep network errors separate from extraction errors in the result schema. This makes it possible to distinguish unreachable pages, non-success HTTP responses, oversized bodies, parse failures, and valid pages with absent fields.

3. Reuse and tune the HTTP transport

The Go net/http documentation says clients and transports are safe for concurrent use and should generally be created once and reused for efficiency. A transport manages connection reuse, including keep-alive connections. Creating one client per URL defeats much of that reuse and adds needless setup.

The example clones the default transport and sets illustrative pool limits. These are starting values for a sample, not recommended universal settings. The default transport supports HTTP/2 where available. When crawling many hosts, idle connections can accumulate; adjust idle limits, per-host limits, and idle timeout for the host distribution and resource budget. CloseIdleConnections releases idle pooled connections when appropriate.

Setting What it influences How to choose
MaxIdleConns Total idle pooled connections Balance reuse against sockets and memory across hosts.
MaxIdleConnsPerHost Idle connections retained for a host Consider repeated work on a small host set versus a broad crawl.
MaxConnsPerHost Connections in dialing, active, or idle states for each host Use as a per-host resource ceiling; account for HTTP/2 multiplexing behavior.
IdleConnTimeout How long idle connections remain pooled Shorten if host churn or resource pressure outweighs reuse.
Client.Timeout Overall request lifetime, including redirects and body reading Set from the job’s latency budget; use request contexts for cancellation.

Consult the official net/http documentation for the exact behavior and defaults of the Go version you deploy. Pool settings interact with target latency, HTTP protocol, proxying, and crawl shape. Observe connection reuse and resource use before changing them.

4. Bound concurrency and respect the workload

Workers overlap network waits; they do not create bandwidth, speed up a slow remote server, or remove CPU and memory limits. Go’s performance guidance illustrates the network ceiling with a 100 Mbps connection already using more than 90 Mbps. That example explains saturation; it is not a scraper benchmark.

Start with a conservative worker count and queue size, then increase workers while observing useful records per second, request latency, status and error rates, CPU, memory, open connections, and bandwidth. Segment metrics by host: an aggregate can conceal one overloaded or slow domain among many healthy ones. If one host is the bottleneck, adding global workers may only increase queued connections and failures.

For production use, add a per-host concurrency or request-rate policy where the workload requires it. Apply the target’s published access rules and operational constraints. Retries should be bounded, limited to errors that may recover, and delayed with backoff; retrying every status or retrying immediately can multiply load. Preserve cancellation through the pipeline so a stopped crawl does not leave workers doing unnecessary requests.

5. Measure before optimizing

Use the same representative URL mix and output requirements when comparing configurations. Record both throughput and quality: a faster run that silently drops fields or raises failure rates is not an improvement. Change one variable at a time where practical, and repeat enough runs to account for network variability.

Go’s official performance guidance covers CPU, heap, blocking, and goroutine profiles. Use CPU profiles for hot computation, heap profiles for allocations and retained memory, and goroutine or blocking profiles for stalled work or excessive concurrency. Diagnostic tools can interfere with one another, so collect profiles for a specific question in isolation where practical. Protect a profiling endpoint according to the deployment environment.

import _ "net/http/pprof"

// In a separately protected diagnostics process or listener:
go func() {
 log.Println(http.ListenAndServe("127.0.0.1:6060", nil))
}()

// Capture a timed CPU profile:
go tool pprof http://127.0.0.1:6060/debug/pprof/profile?seconds=30

// Inspect a heap profile:
go tool pprof http://127.0.0.1:6060/debug/pprof/heap

The snippet needs log and net/http imports in a complete application. Do not expose profiling endpoints publicly without access controls. The net/http/pprof documentation describes profile handlers and go tool pprof workflows.

Consider profile-guided optimization last

Go’s PGO documentation reports around 2–14% improvement across benchmarks for a representative set of Go programs in the Go 1.22 documentation. This range is contextual and is not a prediction for a scraping actor. The Go Authors caution that microbenchmarks are often poor PGO inputs because they exercise too little of an application. Use a representative production profile, build with PGO according to the official PGO guide, and compare the result with the same workload.

6. Choose static fetching or rendered capture

Use the standard HTTP pipeline when the needed data is present in the returned HTML and the extraction can be done server-side. It is usually simpler to operate than a browser process. Use browser rendering when content requires JavaScript execution or when the deliverable is a visual screenshot or PDF. Browser capture adds rendering setup and resource considerations, so keep that path separate from ordinary HTML fetches.

Static HTML fetching and JavaScript rendering solve different capture requirements.
Static HTML fetching and JavaScript rendering solve different capture requirements.

For page screenshots, a dedicated API can remove browser installation and capture orchestration from the actor. ScreenshotNeo is a website screenshot API and MCP server; it returns PNG, JPEG, WebP, or PDF from a URL. Its capture options include full-page and element screenshots, waits, custom headers and cookies, and other page controls. See the ScreenshotNeo site for the product and the docs alongside its API example below.

Or skip the browser setup

Call ScreenshotNeo when the job needs a rendered screenshot instead of extracted HTML. Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. One thousand screenshots a month are free with no card; paid plans start at $5 for 3,000.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

The Node.js example uses Bun’s file-writing helper for brevity; with Node.js, write the response bytes using await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()))) in an async function. Store the API key in an environment variable in deployed code, rather than committing it. Review the ScreenshotNeo API documentation for request parameters and response headers. Create a free account at ScreenshotNeo sign-up to get 1,000 screenshots a month with no card.

7. Reliability, performance, and cost checklist

  • Reuse one client and transport per relevant configuration; close every response body.
  • Set an overall timeout and propagate cancellation from the job scheduler.
  • Bound both active workers and queued jobs; define what happens when intake outruns processing.
  • Limit response size and validate content types before parsing untrusted pages.
  • Record status, latency, bytes, extraction outcome, and errors by host.
  • Bound retries, add delay, and avoid multiplying traffic during an outage.
  • Set connection pool limits from observed host count and connection behavior.
  • Estimate operating cost from the resources actually consumed: network transfer, CPU, memory, browser execution if used, and any external API plan. This dossier provides no universal per-page cost or benchmark.

Reliability also depends on durable job acknowledgement and output handling. If a process can restart, decide whether jobs may be repeated, whether writes are idempotent, and when a job counts as complete. A fetch succeeding does not guarantee the result was durably stored.

8. Troubleshooting common failures

Symptom Likely cause Fix
Requests stall or exceed the deadline Slow host, blocked network operation, or overly strict timeout Inspect latency by host, use request contexts, and set a timeout from the job budget. Check goroutine and blocking profiles.
Too many open connections or file descriptors High host churn or overly generous idle pools Review transport idle limits and timeout; call CloseIdleConnections when retiring a client or crawl phase.
Throughput stops improving as workers rise Bandwidth, remote latency, per-host bottleneck, parser CPU, or memory pressure Compare host-level latency and errors with CPU, bandwidth, and profiles. Reduce concurrency if errors or resource pressure rise.
Some pages return errors despite reachable URLs Non-success status, redirects, TLS/DNS failures, remote throttling, or cancellation Log the status and wrapped error class. Handle redirects deliberately; use bounded backoff only for retryable conditions.
Titles or fields are missing String-based example parser, encoding mismatch, malformed markup, or JavaScript-only content Use an HTML parser and encoding-aware decoding; determine whether rendered content is required.
Memory grows during large responses Whole-body reads, unbounded queues, or retained results Keep body limits, stream large inputs, bound queues, and inspect heap profiles for retained references.
Jobs remain queued after shutdown Intake and workers do not share cancellation or shutdown coordination Propagate a root context, stop intake, close the jobs channel, and wait for workers to finish or their deadlines.

9. Frequently asked questions

Should each worker have its own HTTP client?

Usually no. A shared client and transport are safe for concurrent use. Separate clients make sense when requests need distinct proxy, TLS, cookie, or connection-pool configuration.

What is the best number of workers?

There is no workload-independent best number. Tune against the required useful output rate while watching per-host errors, latency, and resource use.

Does more concurrency make scraping faster?

Only while the actor is limited by parallelizable waiting and still has capacity. It cannot exceed a saturated network or remote service, and can worsen throttling or resource pressure.

Does Go’s PGO range apply to my actor?

No guaranteed gain follows from the published range. It summarizes benchmarks for representative Go programs, not this workload; profile and measure your own build.

Can the standard HTTP client capture a JavaScript page?

No. It fetches HTTP responses but does not run browser JavaScript. Use a rendering-capable approach when the target content depends on it.

Primary references