Retrying Failed Requests in Go: Safe Backoff, Deadlines, and Idempotency
Build bounded, cancelable retries in Go without duplicating writes, losing request bodies, or amplifying an outage.

Retrying a failed HTTP request in Go is an application policy, not a switch on http.Client. The safe pattern is to classify the failure, confirm that repeating the operation is allowed, rebuild the request body, wait with capped exponential backoff and jitter, and stop when the caller’s context or attempt budget ends.
Go’s http.Transport performs only narrow transport-level retries. It does not generally retry 5xx responses or arbitrary Client.Do errors. A production retry loop must therefore define its own attempt limit, deadline, status policy, body replay mechanism, and observability.
1. Decide whether the operation is safe to repeat
Start with the operation’s effect, not just its HTTP method. RFC 9110 describes GET, HEAD, OPTIONS, and TRACE as idempotent by intended semantics. Repeating an idempotent operation should have the same intended effect as one request, although the remote system can still change between attempts.
| Operation | Default retry position | Required check |
|---|---|---|
| GET or HEAD read | Usually retry transient transport failures and selected 5xx/429 responses | Confirm the endpoint is safe and responses are not charged per attempt in a harmful way |
| PUT or DELETE | Often retryable by HTTP semantics | Verify the service’s actual behavior and authorization side effects |
| POST create or action | Do not blindly retry | Use a server-supported idempotency key or deduplication token |
| Payment, job submission, email, or mutation | Retry only with an explicit contract | Reuse the same logical operation ID on every attempt |
A lost response creates an ambiguous outcome: the server may have completed the write even though the client saw a timeout. A retry without deduplication can create two records, charge twice, or trigger an action twice. For a non-idempotent write, send an idempotency key that the remote service documents and preserve that same key across attempts. An Idempotency-Key or X-Idempotency-Key header also lets Go’s transport recognize a request as eligible for some narrow connection-level retry cases, but it does not replace the server’s deduplication behavior.
2. Classify failures instead of retrying everything
Retry only failures that are plausibly temporary. A useful policy separates transport errors, response statuses, and permanent client errors.

- Usually transient: connection reset, temporary DNS or network failure, timeout before a response, HTTP 408, 429, and selected 500, 502, 503, or 504 responses.
- Usually permanent: malformed URL, invalid request data, 400, 401, 403, 404, and most 422 responses. Retrying these consumes time and load without changing the cause.
- Ambiguous: a timeout while sending a write, a connection drop after the server may have received the body, and any operation with external side effects. Resolve these with an idempotency contract or reconciliation query.
When the server sends Retry-After, honor it for the statuses where the service defines it, especially 429 and 503. The value can be a number of seconds or an HTTP date. Cap the resulting wait by your own maximum delay and by the caller’s remaining deadline. See the HTTP semantics definition of Retry-After.
3. A complete bounded retry loop in Go
The following example accepts a request factory, so every attempt gets a fresh request and body. It uses a total context budget, a maximum attempt count, capped exponential backoff, full jitter, response classification, and bounded parsing of Retry-After.
package main
import (
"context"
"errors"
"fmt"
"io"
"math/rand/v2"
"net/http"
"net/url"
"strconv"
"strings"
"time"
)
type RetryPolicy struct {
MaxAttempts int
BaseDelay time.Duration
MaxDelay time.Duration
}
func (p RetryPolicy) normalized() RetryPolicy {
if p.MaxAttempts < 1 { p.MaxAttempts = 1 }
if p.BaseDelay <= 0 { p.BaseDelay = 200 * time.Millisecond }
if p.MaxDelay <= 0 { p.MaxDelay = 5 * time.Second }
return p
}
func retryableStatus(code int) bool {
switch code {
case http.StatusRequestTimeout, http.StatusTooManyRequests,
http.StatusInternalServerError, http.StatusBadGateway,
http.StatusServiceUnavailable, http.StatusGatewayTimeout:
return true
default:
return false
}
}
func retryAfter(h string, now time.Time) time.Duration {
h = strings.TrimSpace(h)
if h == "" { return 0 }
if seconds, err := strconv.Atoi(h); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second
}
if when, err := http.ParseTime(h); err == nil && when.After(now) {
return time.Until(when)
}
return 0
}
func waitFor(ctx context.Context, d time.Duration) error {
timer := time.NewTimer(d)
defer timer.Stop()
select {
case <-ctx.Done():
return ctx.Err()
case <-timer.C:
return nil
}
}
// doWithRetry calls makeRequest once per attempt. The factory must create a
// fresh body reader and request, and must reuse one idempotency key for a write.
func doWithRetry(ctx context.Context, client *http.Client, policy RetryPolicy,
makeRequest func(context.Context) (*http.Request, error)) (*http.Response, error) {
policy = policy.normalized()
var lastErr error
for attempt := 1; attempt <= policy.MaxAttempts; attempt++ {
if err := ctx.Err(); err != nil { return nil, err }
req, err := makeRequest(ctx)
if err != nil { return nil, err } // request construction is not transient
resp, err := client.Do(req)
if err == nil && !retryableStatus(resp.StatusCode) {
return resp, nil
}
if err != nil {
lastErr = err
} else {
lastErr = fmt.Errorf("HTTP status %s", resp.Status)
// Drain a bounded amount before closing to improve connection reuse.
_, _ = io.CopyN(io.Discard, resp.Body, 32<<10)
_ = resp.Body.Close()
}
if attempt == policy.MaxAttempts { break }
delay := policy.BaseDelay * time.Duration(1<<(attempt-1))
if delay > policy.MaxDelay { delay = policy.MaxDelay }
if err == nil {
if serverDelay := retryAfter(resp.Header.Get("Retry-After"), time.Now()); serverDelay > delay {
delay = serverDelay
if delay > policy.MaxDelay { delay = policy.MaxDelay }
}
}
// Full jitter: choose uniformly from [0, delay].
if delay > 0 { delay = time.Duration(rand.Int64N(int64(delay) + 1)) }
if err := waitFor(ctx, delay); err != nil { return nil, err }
}
return nil, lastErr
}
func main() {
ctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
defer cancel()
client := &http.Client{Timeout: 8 * time.Second}
operationID := "order-123-attempt-group" // one value for every retry
response, err := doWithRetry(ctx, client, RetryPolicy{
MaxAttempts: 4,
BaseDelay: 250 * time.Millisecond,
MaxDelay: 3 * time.Second,
}, func(ctx context.Context) (*http.Request, error) {
// Recreate the body on every attempt. For this example the body is empty.
req, err := http.NewRequestWithContext(ctx, http.MethodPost,
"https://api.example.com/orders", strings.NewReader(`{"sku":"book-1"}`))
if err != nil { return nil, err }
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", operationID)
return req, nil
})
if err != nil { panic(err) }
defer response.Body.Close()
fmt.Println(response.Status)
}
var _ = errors.Is
var _ = url.URL{}
In a real package, remove the two blank-reference lines and unused imports if you do not need them. The important structure is the request factory. Calling client.Do repeatedly with the same already-consumed io.Reader does not replay the body.
Why the context is the total budget
Create the parent context at the operation boundary. Each request derives from it, and the backoff wait selects on ctx.Done(). This prevents a caller that has stopped waiting from being held hostage by another attempt. Go’s net/http documentation states that an outgoing request context controls obtaining a connection, sending the request, and reading response headers and body. http.Client.Timeout additionally covers redirects and reading the response body; a zero value means no client timeout.
Use either a context deadline, Client.Timeout, or both with clear ownership. A per-request client timeout can protect one attempt, while the parent context limits the entire retry operation. Ensure the sum of attempts, response reads, and waits fits the parent deadline.
4. Replaying request bodies correctly
A request body is a stream. A file, pipe, socket, or one-shot reader cannot automatically be sent again. Choose one of these approaches:
- Keep the payload in memory and create
bytes.NewReader(payload)for each attempt. - Open or seek a file separately for each attempt.
- Build the request with a body and set
Request.GetBodywhen you control the request and want transport-level replay support. - Use a retry library that documents body rewind behavior, such as HashiCorp go-retryablehttp.
payload := []byte(`{"name":"Ada"}`)
makeRequest := func(ctx context.Context) (*http.Request, error) {
body := bytes.NewReader(payload)
req, err := http.NewRequestWithContext(ctx, http.MethodPost,
endpoint, body)
if err != nil { return nil, err }
req.GetBody = func() (io.ReadCloser, error) {
return io.NopCloser(bytes.NewReader(payload)), nil
}
req.Header.Set("Idempotency-Key", operationID)
return req, nil
}
Transport retries are narrower than this loop. According to the Transport documentation, Go considers a retry only for certain reused connections and idempotent requests, with no body or a defined GetBody. Do not interpret that behavior as application retries for response statuses.
5. Backoff, jitter, and overload control
Exponential backoff increases the pause after each failure. A cap prevents one request from sleeping indefinitely. Jitter prevents thousands of clients that failed together from sending again at the same instant. There is no universal numeric configuration: begin with a small attempt count and delay, then evaluate the remote service’s limits and your caller’s latency budget.
| Control | Purpose |
|---|---|
| Maximum attempts | Bounds work and makes worst-case latency predictable |
| Maximum delay | Prevents very large exponential sleeps |
| Full or equal jitter | Spreads retries over time |
| Overall context deadline | Stops all attempts when the caller’s budget is gone |
| Retry-After handling | Follows server-specific overload guidance |
Unlimited retries can turn an outage into a larger outage and can inflate usage charges. AWS’s Go SDK guidance also warns about unlimited retries and provides configurable maximum attempts and rate limiting. If an SDK already retries, calculate the combined worst case before adding an outer loop.
6. Observability and response handling
Record the operation name, attempt number, elapsed time, selected delay, status or error class, and final outcome. Avoid logging secrets, authorization headers, cookies, and full request bodies. Include the idempotency key only if it is safe in your logging policy.
Always close response bodies. If a response will be discarded before another attempt, drain only a bounded number of bytes, then close it. Never read an unbounded error page merely to preserve connection pooling. Return the final response when the status is accepted by your policy; return the last error after the attempt budget is exhausted.
7. When a retry library is a better fit
A direct loop is often clearest for one API with a small, explicit policy. A package can help when many clients need consistent classification, backoff, logging, and body rewind. Compare options on these axes:
- Policy control: can you classify this service’s statuses and errors?
- Request replay: can bodies be recreated and operation identity preserved?
- Cancellation: does waiting stop on context completion?
- Load behavior: are attempts, delay, jitter, and
Retry-Afterbounded? - Integration: will it stack with retries already inside an SDK?
go-retryablehttp documents retry checks, exponential backoff, and request-body rewind support. Review its current module version and customization points before adopting it. For AWS, understand the AWS SDK for Go v2 retryer already in use before wrapping calls, because two policies can multiply attempts and delay.
8. Common errors and fixes
| Symptom | Cause | Fix |
|---|---|---|
| POST creates duplicates | The write was retried without deduplication | Use a documented idempotency key and reuse it for every attempt, or do not retry |
| Second attempt has an empty body | The original reader was consumed | Rebuild the reader or implement GetBody |
| Retries continue after request cancellation | Backoff used time.Sleep |
Wait with a timer that selects on ctx.Done() |
| 400 or 401 is retried repeatedly | Status policy retries all non-2xx responses | Allowlist transient statuses and classify permanent errors |
| Traffic spikes every few seconds | All clients use fixed delay or identical exponential delays | Add jitter and honor server rate limits |
| Connections are not reused | Response bodies are left open | Close every body; drain only a bounded amount when discarding |
| Requests take far longer than expected | Per-attempt timeouts were added without a total deadline | Set a parent context deadline and cap attempts and waits |
| Retry-After causes an excessive sleep | Server value was trusted without a cap | Bound it by the policy maximum and remaining context time |
9. Testing a retry policy without causing an outage
Use an httptest.Server to return a controlled sequence: connection failure, 503, 429 with Retry-After, then success. Assert the attempt count, body on every attempt, idempotency header, cancellation during backoff, and response-body closure. Add a test where the context expires between attempts. Keep timing tests tolerant of scheduler variation; inject a clock or wait function if exact backoff assertions are important.
10. Or skip the browser setup
If the operation you need is capturing a webpage rather than calling your own HTTP service, ScreenshotNeo provides a single request that returns PNG, JPEG, WebP, or PDF. The equivalent call is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for request parameters and response headers. Before capture, cookie and consent banners, newsletter popups, and chat widgets are removed. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the page verdict and billing result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.
11. Performance, reliability, and cost notes
- Latency: retries add the delay plus another network round trip. Keep the maximum attempts low for interactive requests and reserve longer budgets for background work.
- Throughput: jitter and server-provided delays reduce synchronized load. Use a client-side rate limiter when many goroutines share one API.
- Reliability: retries help only when failures are transient. Circuit breakers, queues, and reconciliation are often better for sustained outages or ambiguous writes.
- Cost: every attempt can consume compute, quota, or money at the remote service. Calculate worst-case attempts, including retries inside SDKs and proxies.
- Concurrency: a retry loop in each goroutine can multiply traffic quickly. Bound concurrent operations and collect retry metrics by endpoint.
FAQ
Does Go automatically retry HTTP requests?
http.Transport retries only a narrow set of network failures under documented connection reuse, idempotency, and body replay conditions. It does not provide a general application retry policy.
Should I retry every 500 response?
No. Retry only statuses your service treats as transient, with an attempt cap and deadline. Some 500 responses represent permanent application failures.
Is POST always unsafe to retry?
POST is not automatically safe. Retry it only when the remote service provides a working idempotency or deduplication contract and you reuse the same operation identity.
How many attempts should I allow?
There is no universal number. Choose a finite limit that fits the caller’s deadline, remote rate limits, and the value of the operation, then measure the result.
Can I use time.Sleep between attempts?
It works only when cancellation is irrelevant. Production code should use a timer selected with the context so a canceled caller stops immediately.
Should I use a retry package?
Use a package when shared policy, body rewind, and consistent instrumentation outweigh an extra dependency. A small direct loop is appropriate when the service-specific rules are clearer in local code.


