Web Scraping in Golang: Tutorial with Quick Start Examples
Learn Go web scraping with net/http, goquery, and Colly, including runnable code, crawler controls, troubleshooting, and clean screenshots.
Short answer: In Go, fetch HTML with net/http, parse it with goquery, and use Colly when you need link traversal, domain limits, callbacks, caching, cookies, concurrency, or robots.txt handling. Start with one page, validate status and parsing, then add bounded crawling.
What web scraping in Golang means
Scraping has two separate jobs: downloading a response and extracting data from its HTML. The standard library is enough for the first job. A selector parser makes the second job maintainable. A crawler framework coordinates many requests and links. Keeping these layers separate makes failures easier to diagnose and lets you replace one layer without rewriting the others.
Prerequisites and project setup
- Install a current Go toolchain.
- Create a module:
mkdir go-scraper cd go-scraper go mod init example.com/go-scraper - Add goquery when you need CSS selectors:
go get github.com/PuerkitoBio/goquery - Add Colly for crawling:
go get github.com/gocolly/colly/v2
Check the target site’s terms and /robots.txt before collecting data. Keep your scope and request rate small enough to avoid degrading the service.
Quick start: fetch one page with net/http
This complete program follows the safe response lifecycle: create the request, handle transport errors, close the body, check the status, and read the bytes.
package main
import (
"fmt"
"io"
"log"
"net/http"
"time"
)
func main() {
client := &http.Client{Timeout: 20 * time.Second}
resp, err := client.Get("https://example.com/")
if err != nil {
log.Fatal(err)
}
defer resp.Body.Close()
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
log.Fatalf("unexpected HTTP status: %s", resp.Status)
}
body, err := io.ReadAll(resp.Body)
if err != nil {
log.Fatal(err)
}
fmt.Printf("%s", body)
}
defer resp.Body.Close() matters: leaving bodies open can exhaust connections and file descriptors during a crawl. A non-2xx response is data about the target, not a successful page, so decide whether to retry, record, or skip it.
Parse HTML with goquery
Fetch first, then parse the response bytes. CSS selectors are easiest to maintain when they use stable semantic elements or classes. Test selectors against representative pages because missing nodes are normal.
package main
import (
"fmt"
"log"
"net/http"
"time"
"github.com/PuerkitoBio/goquery"
)
func main() {
client := &http.Client{Timeout: 20 * time.Second}
resp, err := client.Get("https://example.com/")
if err != nil { log.Fatal(err) }
defer resp.Body.Close()
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
log.Fatalf("unexpected HTTP status: %s", resp.Status)
}
doc, err := goquery.NewDocumentFromReader(resp.Body)
if err != nil { log.Fatal(err) }
title := doc.Find("title").First().Text()
fmt.Println("title:", title)
doc.Find("a[href]").Each(func(_ int, s *goquery.Selection) {
href, ok := s.Attr("href")
if ok { fmt.Println("link:", href) }
})
}
Use Text() for visible text and Attr() for attributes. Normalize whitespace and resolve relative URLs before storing links. Treat absent selectors as empty results and validate required fields explicitly.
Build a multi-page crawler with Colly
Colly supplies a collector and callbacks for traversal. This example restricts the crawl to one host, prints page titles, resolves links, and avoids duplicate visits.
package main
import (
"fmt"
"log"
"github.com/gocolly/colly/v2"
)
func main() {
c := colly.NewCollector(
colly.AllowedDomains("example.com"),
colly.MaxDepth(2),
)
c.OnHTML("title", func(e *colly.HTMLElement) {
fmt.Printf("%s => %s\n", e.Request.URL, e.Text)
})
c.OnHTML("a[href]", func(e *colly.HTMLElement) {
link := e.Request.AbsoluteURL(e.Attr("href"))
if link != "" {
if err := e.Request.Visit(link); err != nil {
fmt.Println("visit skipped:", err)
}
}
})
c.OnRequest(func(r *colly.Request) {
fmt.Println("visiting", r.URL.String())
})
c.OnError(func(r *colly.Response, err error) {
fmt.Printf("%s failed: %v (status %d)\n", r.Request.URL, err, r.StatusCode)
})
if err := c.Visit("https://example.com/"); err != nil {
log.Fatal(err)
}
}
Install Colly with go get github.com/gocolly/colly/v2. Add URL pattern checks when a site contains calendars, search parameters, or other unbounded link generators.
Choosing net/http, goquery, or Colly
| Need | Recommended stack | Why |
|---|---|---|
| One page or a small API-like extraction | net/http | Smallest dependency surface and explicit control. |
| Selectors, text, attributes, and links | net/http + goquery | Separates transport from HTML parsing. |
| Link traversal and crawl rules | Colly | Collector callbacks, allowed domains, depth, caching, cookies, concurrency, and robots.txt support. |
There is no authoritative apples-to-apples benchmark here that proves one stack is universally faster. Measure your target, selector work, network latency, and concurrency together.
Useful Colly configuration
Scope and politeness
c := colly.NewCollector(
colly.AllowedDomains("example.com"),
colly.MaxDepth(3),
colly.IgnoreRobotsTxt(false),
)
c.Limit(&colly.LimitRule{
DomainGlob: "example.com/*",
Parallelism: 2,
Delay: 500 * time.Millisecond,
})
Keep parallelism and delay bounded. Domain restrictions are a safety boundary, not permission to crawl aggressively. Add an allowlist of paths when possible.
Headers, cookies, and identity
c.OnRequest(func(r *colly.Request) {
r.Headers.Set("User-Agent", "research-bot/1.0 (+https://example.com/contact)")
r.Headers.Set("Accept", "text/html")
})
c.SetCookies("example.com", []*http.Cookie{{Name: "session", Value: "value"}})
Use an honest user agent and only send cookies or authorization you are allowed to use. Do not hard-code secrets in source.
Retries and status handling
c := colly.NewCollector()
c.SetRequestTimeout(20 * time.Second)
c.OnResponse(func(r *colly.Response) {
if r.StatusCode < 200 || r.StatusCode >= 300 {
fmt.Println("non-success:", r.StatusCode, r.Request.URL)
}
})
c.OnError(func(r *colly.Response, err error) {
// Record the URL and error, then retry only bounded, transient failures.
fmt.Println("error:", r.Request.URL, err)
})
Retry timeouts and selected 5xx responses with exponential backoff and a maximum attempt count. Avoid retrying permanent 4xx responses or malformed URLs.
Caching and repeatable runs
Colly documents HTTP caching and response controls. Cache during development so selector changes do not repeatedly hit the origin. In production, record the crawl timestamp, URL, status, content hash, and parser version so runs can be audited.
Handling JavaScript-rendered pages
net/http and Colly receive the server response; they do not execute browser JavaScript. If the required content appears only after scripts run, first look for a documented JSON endpoint or embedded data. Browser automation or a hosted rendering service is an advanced branch: it adds startup time, resource use, and more failure modes. Keep it separate from the simple HTTP crawler and enforce the same domain, rate, and terms rules.
Equivalent quick starts in other languages
cURL
curl --fail --location --max-time 20 https://example.com/
Python
import requests
r = requests.get("https://example.com/", timeout=20)
r.raise_for_status()
print(r.text)
Node.js
const res = await fetch('https://example.com/');
if (!res.ok) throw new Error(`HTTP ${res.status}`);
console.log(await res.text());
These examples fetch HTML only. Parsing and browser rendering remain separate steps in each language.
Edge cases to design for
- Redirects: record the final URL and reject redirects that leave your allowlist.
- Relative links: resolve against the page URL; discard fragments when deduplicating.
- Huge responses: enforce a maximum body size before parsing.
- Encoding: preserve the declared charset and normalize text before storage.
- Missing or changing markup: treat selectors as optional, validate required fields, and keep parser metrics.
- Pagination: set a page limit and stop when the next link repeats.
- Forms and sessions: model state explicitly and protect credentials.
- Bot checks and CAPTCHAs: do not attempt to bypass access controls; stop or use an authorized browser workflow.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
context deadline exceeded |
Timeout is too short or the origin is slow. | Set a realistic client timeout, log timings, and retry transient failures with a cap. |
| Body read errors or connection exhaustion | Response bodies are not closed. | Defer resp.Body.Close() immediately after a successful request. |
| Empty selector results | Markup changed, selector is wrong, or content is JavaScript-rendered. | Save a sample response, inspect the HTML, update selectors, or use a browser-capable path. |
| 403 or 429 responses | Access policy or rate limit. | Check terms and robots.txt, slow down, identify your client, and stop when access is denied. |
| Colly revisits too many URLs | Unbounded query parameters or missing domain/path rules. | Use AllowedDomains, depth and path filters, canonicalize URLs, and cap pages. |
| Non-UTF text | Response charset differs from UTF-8. | Detect the declared charset and convert before parsing or indexing. |
Performance, reliability, and cost
- Performance: network latency usually dominates. Reuse an
http.Client, keep responses bounded, cache during development, and add concurrency only after measuring. - Reliability: record URL, status, duration, attempt, parser version, and error. Use bounded retries, idempotent storage, checkpoints, and graceful cancellation.
- Cost: self-hosted Go uses your compute, bandwidth, and storage. Browser rendering costs more resources than plain HTTP. Caching and incremental crawls reduce both load and spend.
- Responsible operation: honor robots.txt and terms, limit scope and rate, and do not bypass authentication, CAPTCHAs, or other access controls.
Or skip the browser setup
For a clean screenshot or PDF rather than raw HTML, ScreenshotNeo provides a single API request. Cookie banners, newsletter popups, and chat widgets are removed before the shot; bot checks, blank pages, timeouts, failed loads, and cache hits are never billed. Its MCP server lets AI agents take screenshots, and 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000. See the ScreenshotNeo API docs for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Create a free ScreenshotNeo account and start with 1,000 screenshots per month at no charge and no card.
FAQ
Is Go good for web scraping?
Yes. net/http, goquery, and Colly cover one-page extraction through structured multi-page crawls.
Do I need Colly for one URL?
No. net/http plus a parser is simpler and keeps every step visible.
Can Colly execute JavaScript?
Colly is an HTTP crawler. JavaScript-only content needs an endpoint or browser-capable workflow.
How do I avoid scraping too much?
Use domain and path allowlists, depth and page caps, robots.txt checks, delays, bounded concurrency, and a clear stop condition.


