How to Use a Go SDK for Web Scraping APIs
Install a provider’s Go SDK, authenticate safely, make a context-aware scrape request, and handle API and network errors in production.

A Go SDK for a web scraping API is a provider-specific wrapper around that provider’s HTTP API. To use one, install its documented Go module, configure its API key outside source control, pass a caller-owned context.Context to a scrape method, and inspect both the returned response and any error. There is no universal Go scraping SDK, import path, key variable, or response format.
This guide uses webscrape.ai’s Go SDK for a concrete single-page example. Its package reference requires Go 1.22 or newer. Check the provider’s current documentation before copying the version, installation command, or request fields: SDK requirements and API behavior can change.
1. Choose an SDK for the operation you need
Start with the task, then check whether the provider’s SDK supports it. A single-page scrape, a multi-page crawl, a batch request, and structured data extraction are different operations. Their availability and response handling vary by provider.
| Check | What to confirm |
|---|---|
| Go compatibility | Minimum Go version, module path, active release, and import name. |
| Authentication | How to supply a key, whether an environment variable is supported, and how to rotate credentials. |
| Operation and output | Whether you need one page, a crawl, batch processing, HTML, markdown, or structured extraction. |
| Request lifecycle | Whether a call returns directly or starts a job that must be polled; how context cancellation is handled. |
| Errors and service limits | How to identify authentication failures, throttling, invalid requests, quotas, and provider-specific error details. |
For example, webscrape.ai documents a context-based Scrape call. Webclaw documents a different Go module and API key setup; its SDK overview lists Go 1.21+ and operations that include scrape, crawl, map, batch, and extract. Those are provider-specific details, not shared Go conventions. Compare providers using current official documentation; the available source material does not support a speed, success-rate, or price ranking.
2. Install the module and configure a key
For the webscrape.ai example, the documented installation command is:
go get github.com/webscrape-ai/webscrape-ai/sdk/go
The import path ends in /go, so the README uses the explicit import name webscrape. Set the key in your shell or deployment configuration rather than committing it to a Go file:
export WEBSCRAPE_API_KEY="your-provider-key"
The package documentation says New() reads WEBSCRAPE_API_KEY and returns ErrNoAPIKey if no key is available. It also documents creating a client with an explicit key. These rules belong to this SDK. Do not assume another provider uses the same environment variable or accepts the same key format.
In production, provide secrets through your hosting platform’s secret manager or protected environment configuration. Keep keys out of source control, logs, error messages sent to users, and URLs. Limit access to the service that needs the credential and rotate it using the provider’s documented process.
3. Make a context-aware scrape request
This complete example loads the key through the documented New() behavior, gives the request a deadline, asks for cleaned output and extracted links, and writes the returned HTML to a file. Save it as main.go, set the environment variable, and run go run . from a module directory.

package main
import (
"context"
"fmt"
"log"
"os"
"time"
webscrape "github.com/webscrape-ai/webscrape-ai/sdk/go"
)
func main() {
client, err := webscrape.New()
if err != nil {
log.Fatalf("create scraping client: %v", err)
}
ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
defer cancel()
resp, err := client.Scrape(ctx, &webscrape.ScrapeRequest{
WebsiteURL: "https://example.com",
Clean: webscrape.Bool(true),
ExtractLinks: webscrape.Bool(true),
})
if err != nil {
log.Fatalf("scrape page: %v", err)
}
if resp == nil || resp.Data == nil || resp.Data.HTML == nil {
log.Fatal("scrape response did not contain HTML")
}
if err := os.WriteFile("page.html", []byte(*resp.Data.HTML), 0600); err != nil {
log.Fatalf("write page.html: %v", err)
}
fmt.Printf("saved %d bytes of HTML to page.html\n", len(*resp.Data.HTML))
}
The context.WithTimeout deadline is an application choice, not a claim about the SDK’s default timeout. Set a deadline appropriate to your caller and workload. The response check avoids dereferencing optional data when the provider returns a valid response without the field your code expected.
4. Understand optional fields and responses
The webscrape.ai README describes optional request fields as pointer types with omitempty. That lets the SDK distinguish “not provided” from an explicitly supplied false or zero value. Its helper functions, including Bool, Int, and String, construct pointers for optional values. For example, webscrape.Bool(true) explicitly enables the documented option; leaving Clean out leaves it unset.
Choose request fields deliberately:
- Request only the output you need, and read the matching documented response field.
- Do not assume a missing optional response field is an empty string; it may be absent or represented as a nil pointer.
- Check the provider’s schema for the meaning of “clean,” link extraction, markdown, or other options before relying on them.
- Store or transform the returned content according to your application’s requirements and the provider’s terms.
The SDK source also describes response credit fields. Treat those as provider-specific usage metadata and verify their current meaning in the package documentation. The research available for this guide does not establish a comparable quota or pricing figure.
5. Handle errors by category
A non-nil Go error can represent a context deadline, a connection problem, an SDK error, or a provider response indicating that the request was invalid or unauthorized. These cases need different responses. Start by preserving the error context, then use the SDK’s documented error types or helpers to classify it.
resp, err := client.Scrape(ctx, req)
if err != nil {
if errors.Is(err, context.DeadlineExceeded) {
// The caller's deadline expired. Decide whether a later retry is useful.
return fmt.Errorf("scrape deadline exceeded: %w", err)
}
if errors.Is(err, context.Canceled) {
// The caller no longer needs the result; normally stop this work.
return fmt.Errorf("scrape canceled: %w", err)
}
// Inspect provider-specific API error details here, if exposed by the SDK.
return fmt.Errorf("scrape request failed: %w", err)
}
This fragment assumes it is inside a function returning an error and that errors and fmt are imported. Do not retry every error identically. A canceled request should generally stop because its caller no longer wants the result. A malformed request or invalid key will not be repaired by an immediate retry. For throttling or temporary service errors, follow the provider’s documented retry guidance and any supplied retry metadata.
Webclaw’s repository describes typed API errors and helper predicates for conditions such as rate limiting, authentication errors, and not-found responses. Those helpers are Webclaw-specific. Use the selected SDK’s own documented error inspection rather than copying another client’s predicates.
6. Add production controls
Bound time and concurrency
Pass a request-scoped context from the caller when possible. A deadline prevents a handler or worker from waiting indefinitely, while cancellation lets upstream work stop when the result is no longer needed. For many URLs, cap concurrent requests instead of launching an unbounded goroutine per page. Choose the cap using the provider’s current limits and your own memory and latency budgets.
Retry selectively
Retry only failures that may be temporary, and use bounded attempts with backoff. Avoid retrying authentication failures, invalid URLs, or other permanent request errors without first changing the cause. Respect provider guidance for throttling and quotas. If a timed-out request may have completed remotely, verify whether the provider supports an idempotency key or documents duplicate-request behavior before retrying side-effecting operations.
Make outputs observable without leaking secrets
Record useful operational details such as operation name, elapsed time, result status, and a provider request identifier if the SDK exposes one. Do not log API keys or unrestricted scraped content by default. Track error categories separately so a spike in timeouts is distinguishable from a bad deployment secret or rate limiting.
Confirm cost and terms before scaling
SDK installation does not tell you the provider’s billing unit, rate limit, retention policy, or allowed use. Confirm current pricing, quotas, concurrency limits, and service terms in official provider documentation before scheduling large jobs. Estimate volume from the number of requested pages and any retries, then monitor the provider’s usage reporting. Do not infer a provider’s price or scrape coverage from its SDK features.
7. Common problems and fixes
| Symptom | Likely cause | What to do |
|---|---|---|
go get cannot find the module |
The module path was copied incorrectly, the version is unavailable, or a proxy/network issue blocked retrieval. | Copy the install path from the provider’s current package documentation; check the module path and Go module settings. |
| The import does not compile | The import path or package name differs from the example, or the installed SDK version changed. | Use the documented import path and inspect the package’s current examples. In this example, the explicit alias is webscrape. |
ErrNoAPIKey at client creation |
The documented environment variable is unset in the process that runs the program. | Set WEBSCRAPE_API_KEY in that process’s environment, or use the provider’s documented explicit-key constructor. |
| Authentication or permission error | The key may be incorrect, revoked, restricted, or associated with insufficient access. | Verify the key and account permissions in the provider’s dashboard; replace the secret securely. Avoid printing it during debugging. |
| Deadline exceeded | The request took longer than the caller’s context deadline, or the network/provider did not respond in time. | Check connectivity and provider status, inspect latency, and choose a suitable deadline. Retry only if the error and operation make retry appropriate. |
| Rate limit or quota error | Request volume exceeded a provider-specific limit or available usage. | Reduce concurrency, apply documented backoff, and check current quotas and usage. Do not assume limits from another SDK apply. |
| HTML is nil or unexpectedly empty | The response did not include the requested field, the page produced no content, or the response schema differs from the code’s assumption. | Check the full documented response structure and request options; handle absent fields explicitly and inspect provider error details. |
| A retry duplicates work | The first request may have completed despite a client-side timeout. | Check the API’s documented idempotency and job semantics before retrying; persist job identifiers when the workflow provides them. |
8. Alternatives for page screenshots
A scraping SDK returns page content through the provider’s scraping operation. If your actual task is to capture a visual screenshot, a screenshot API is a more direct fit. ScreenshotNeo is a website screenshot API and MCP server: one GET request can return a PNG, JPEG, WebP, or PDF. Its screenshot features include full-page capture, CSS selector capture, device presets, custom CSS and JavaScript, and configurable waits. It does not replace a scraping SDK when you need HTML or structured extraction.
9. Or skip the browser setup
For a screenshot rather than scraped HTML, call ScreenshotNeo’s API directly. See the ScreenshotNeo API documentation for the current request options. The following Go example makes a GET request with the standard library, checks the HTTP status, and saves the response body. Set SCREENSHOTNEO_API_KEY in your environment first.

package main
import (
"context"
"fmt"
"io"
"net/http"
"net/url"
"os"
"time"
)
func main() {
key := os.Getenv("SCREENSHOTNEO_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "SCREENSHOTNEO_API_KEY is required")
os.Exit(1)
}
endpoint, _ := url.Parse("https://api.screenshotneo.com/v1/shot")
query := endpoint.Query()
query.Set("access_key", key)
query.Set("url", "https://stripe.com")
endpoint.RawQuery = query.Encode()
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
defer cancel()
req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint.String(), nil)
if err != nil {
fmt.Fprintln(os.Stderr, "create request:", err)
os.Exit(1)
}
res, err := http.DefaultClient.Do(req)
if err != nil {
fmt.Fprintln(os.Stderr, "send request:", err)
os.Exit(1)
}
defer res.Body.Close()
if res.StatusCode < 200 || res.StatusCode >= 300 {
body, _ := io.ReadAll(res.Body)
fmt.Fprintf(os.Stderr, "ScreenshotNeo returned %s: %s\n", res.Status, body)
os.Exit(1)
}
f, err := os.Create("shot.webp")
if err != nil {
fmt.Fprintln(os.Stderr, "create output:", err)
os.Exit(1)
}
if _, err := io.Copy(f, res.Body); err != nil {
_ = f.Close()
fmt.Fprintln(os.Stderr, "save screenshot:", err)
os.Exit(1)
}
if err := f.Close(); err != nil {
fmt.Fprintln(os.Stderr, "close output:", err)
os.Exit(1)
}
}
Equivalent requests using the documented endpoint and parameters:
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots a month without a card; paid plans start at $5 for 3,000 shots. Sign up free for ScreenshotNeo.
10. FAQ
Is there one standard Go SDK for scraping websites?
No. Each provider publishes its own module, methods, authentication setup, request schema, and error behavior. Select the SDK that documents the operation and output you need.
Should I use context.Background() for every scrape?
Use it as a root context when appropriate, then derive a context with a deadline or pass through a caller’s context. That gives your application control over cancellation and wait time.
Can a screenshot endpoint return scraped HTML?
A screenshot endpoint is intended to return an image or PDF. Use a scraping API when your application needs page HTML, markdown, or structured extraction.
How do I know whether an SDK supports my Go version?
Check the current package documentation and module metadata before installing. The examples here are version-specific statements from the cited provider materials, not guarantees for every release.


