Common Questions About Web Scraping and Guzzle in PHP
Learn how to scrape websites with Guzzle in PHP: headers, cookies, redirects, retries, JavaScript limits, troubleshooting, and a browser-rendering alternative.

Direct answer: Guzzle is an HTTP client, so it is an excellent request layer for PHP scrapers that can work with the HTML or JSON returned by a server. Create a GuzzleHttp\Client, pass explicit request options such as headers, query parameters, cookies, timeouts, redirects, and authentication, then parse the response. Guzzle does not execute JavaScript or render a browser DOM. If the values you need appear only after client-side code runs, add a browser-rendering service or automation layer.
What Guzzle does in a scraper
Guzzle provides synchronous and asynchronous HTTP requests, PSR-7 request and response messages, streams, and middleware. Its client is configured with defaults, while each call can override options. The official Quickstart and request-options reference describe this model.

| Need | Guzzle approach |
|---|---|
| Fetch HTML or JSON | request(), get(), or post() |
| Set a User-Agent or Accept header | headers request option |
| Add query-string values | query option |
| Keep a login/session | CookieJar, FileCookieJar, or SessionCookieJar |
| Follow or inspect redirects | allow_redirects options |
| Handle failures | http_errors, exceptions, status checks, and bounded retries |
| Run JavaScript | Use a browser automation or rendering layer; Guzzle alone is not a browser |
Install Guzzle and create a client
Start a PHP project with Composer:
composer require guzzlehttp/guzzle
A minimal scraper can then fetch a page and inspect its status, headers, and body:
<?php
require __DIR__ . '/vendor/autoload.php';
use GuzzleHttp\Client;
$client = new Client([
'timeout' => 20,
'connect_timeout' => 10,
'http_errors' => false,
]);
$response = $client->request('GET', 'https://example.com');
printf("Status: %d\n", $response->getStatusCode());
printf("Content-Type: %s\n", $response->getHeaderLine('Content-Type'));
echo $response->getBody()->getContents();
Client defaults are set at construction and are effectively immutable. Create another client when a different base URI, timeout, or handler configuration is required. Keep request-specific options beside the request so a scraper’s behavior is easy to audit.
Send scraper-friendly requests
Use a descriptive User-Agent and an Accept header that matches the content you can parse. Put URL parameters in query; Guzzle encodes them correctly.
<?php
require __DIR__ . '/vendor/autoload.php';
use GuzzleHttp\Client;
$client = new Client([
'base_uri' => 'https://example.com',
'timeout' => 30,
]);
$response = $client->get('/search', [
'headers' => [
'User-Agent' => 'CatalogResearchBot/1.0 (+https://your-domain.example/bot-info)',
'Accept' => 'text/html,application/xhtml+xml',
'Accept-Language' => 'en-US,en;q=0.8',
],
'query' => [
'q' => 'guzzle',
'page' => 2,
],
]);
$html = (string) $response->getBody();
Respect the target site’s terms, robots guidance, authentication rules, and rate limits. A clear User-Agent also gives operators a way to contact you.
Parse HTML and JSON safely
Guzzle returns bytes; parsing is your responsibility. For JSON APIs, decode and validate the result before reading fields:
<?php
$data = json_decode((string) $response->getBody(), true, 512, JSON_THROW_ON_ERROR);
$items = $data['items'] ?? [];
foreach ($items as $item) {
if (isset($item['id'], $item['name'])) {
printf("%s: %s\n", $item['id'], $item['name']);
}
}
For HTML, use a DOM parser such as PHP’s DOMDocument and DOMXPath, or a maintained third-party parser. Check the response’s Content-Type and status first; an error page can be valid HTML but still contain none of the data you expected.
Headers, authentication, and request bodies
Headers can be supplied per request or as client defaults. Basic authentication uses the auth option. Form, JSON, and raw bodies use different options:
<?php
$response = $client->post('/api/items', [
'auth' => ['api-user', 'api-password'],
'headers' => ['Accept' => 'application/json'],
'json' => ['name' => 'sample', 'active' => true],
]);
$formResponse = $client->post('/login', [
'form_params' => ['email' => 'you@example.com', 'password' => 'secret'],
]);
Do not log authorization headers, passwords, session cookies, or full URLs containing secrets. For large uploads or downloads, use streams and process incrementally instead of loading the entire body into memory.
Keep cookies between requests
Cookies require a cookie jar and cookie middleware. The documented default handler stack includes that middleware; a custom handler stack without it makes the cookies option ineffective. See Guzzle’s cookie options and handler documentation.
<?php
require __DIR__ . '/vendor/autoload.php';
use GuzzleHttp\Client;
use GuzzleHttp\Cookie\CookieJar;
$jar = new CookieJar();
$client = new Client([
'cookies' => $jar,
'timeout' => 30,
]);
$client->post('https://example.com/login', [
'form_params' => ['email' => 'you@example.com', 'password' => 'secret'],
]);
$account = $client->get('https://example.com/account');
echo $account->getBody()->getContents();
Use FileCookieJar when a session must survive a process restart, and SessionCookieJar when PHP session storage is appropriate. Protect cookie files because they may grant account access.
Redirects and redirect diagnostics
Guzzle follows normal redirects automatically, up to five by default. Disable following when you need to inspect a 3xx response:
$response = $client->get('https://example.com/old-page', [
'allow_redirects' => false,
]);
$redirected = $client->get('https://example.com/start', [
'allow_redirects' => [
'max' => 10,
'strict' => true,
'protocols' => ['https'],
'track_redirects' => true,
],
]);
$history = $redirected->getHeaderLine('X-Guzzle-Redirect-History');
$statuses = $redirected->getHeaderLine('X-Guzzle-Redirect-Status-History');
Use on_redirect for an application callback. Redirect history contains intermediate locations and statuses; the initial URI and final status are not included in those history values. Restricting protocols helps prevent an HTTP downgrade or an unexpected scheme.
Timeouts, status codes, and bounded retries
Set both a connection timeout and an overall timeout. Decide whether HTTP errors should throw. With the default error middleware, responses at or above 400 can raise a Guzzle exception; setting http_errors => false lets your code classify every response explicitly.
<?php
use GuzzleHttp\Exception\ConnectException;
use GuzzleHttp\Exception\RequestException;
$attempts = 0;
while (true) {
try {
$attempts++;
$response = $client->get($url, [
'timeout' => 30,
'connect_timeout' => 8,
'http_errors' => false,
]);
$status = $response->getStatusCode();
if ($status === 429 || $status >= 500) {
if ($attempts < 3) {
$retryAfter = (int) $response->getHeaderLine('Retry-After');
sleep(max(1, min($retryAfter ?: 2 ** $attempts, 30)));
continue;
}
}
if ($status >= 400) {
throw new RuntimeException("HTTP failure: {$status}");
}
$body = (string) $response->getBody();
break;
} catch (ConnectException $e) {
if ($attempts >= 3) throw $e;
usleep(250000 * $attempts);
} catch (RequestException $e) {
throw $e;
}
}
Retry only transient failures such as connection resets, 429, and selected 5xx responses. Use a cap and backoff; never retry a non-idempotent operation blindly. Log the URL, method, attempt, status, and timing while redacting secrets.
Handler stacks and transports
Guzzle can use cURL, PHP streams, sockets, or other handlers. The handler stack controls middleware for cookies, redirects, body preparation, and HTTP errors. If an option appears to do nothing after you provide a custom handler, inspect that stack and add the middleware you need. HandlerStack::create() builds a stack with the normal defaults.

Transport choice changes TLS support, proxy behavior, connection reuse, and observability, but it does not turn Guzzle into a JavaScript browser. Use Guzzle for direct HTTP or API calls wherever possible; reserve browser automation for pages whose content is created after scripts execute.
When Guzzle is not enough
If a plain GET returns a shell page while a real browser shows products, comments, or prices, inspect the HTML and network calls. The data may come from an API you can call directly, or it may require JavaScript execution, cookies set by scripts, interaction, or a bot challenge. Guzzle’s documented feature set covers HTTP transport and response handling, not DOM rendering or JavaScript execution.
Performance, reliability, and cost considerations
- Reuse one client so connections and configuration can be reused.
- Set explicit timeouts and cap concurrency; high parallelism can trigger rate limits or exhaust file descriptors.
- Prefer an official JSON endpoint over repeatedly downloading and parsing large HTML documents.
- Stream large bodies to disk and avoid retaining every page in memory.
- Cache responses when freshness requirements allow it, and use conditional requests such as ETag or Last-Modified where supported.
- Record status, final URL, content type, byte count, latency, and retry count for diagnosis.
- Costs are normally your PHP host, bandwidth, storage, and any proxy or browser service you add. Guzzle itself is open-source software; the target service may impose its own limits.
Troubleshooting checklist
| Symptom | Likely cause | Fix |
|---|---|---|
| 403 response | Access policy, missing headers, or bot protection | Verify permission, send an honest User-Agent, slow requests, and avoid bypassing controls. |
| 429 response | Rate limit exceeded | Honor Retry-After, reduce concurrency, add bounded backoff, and cache results. |
| Cookies are ignored | No cookie jar or missing cookie middleware | Pass a CookieJar and retain the default HandlerStack middleware. |
| Redirect option has no effect | Custom handler stack omitted redirect middleware | Build from HandlerStack::create() or add redirect middleware. |
| HTML lacks visible data | Data is inserted by JavaScript | Find the underlying API or use a browser-rendering layer. |
| SSL or connection timeout | Network, DNS, TLS, or an overloaded origin | Check DNS and certificates, separate connect and total timeouts, then retry transient failures. |
| Parser reports malformed content | Error page, compressed/encoded response, or truncated body | Check status and Content-Type, read the complete stream, and verify encoding before parsing. |
Or skip the browser setup
When your goal is a clean screenshot or PDF rather than parsed HTTP data, ScreenshotNeo provides a single website screenshot API request. It accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. It also offers an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for all options. This PHP example saves a WebP image:
<?php
require __DIR__ . '/vendor/autoload.php';
use GuzzleHttp\Client;
$client = new Client(['timeout' => 90]);
$response = $client->get('https://api.screenshotneo.com/v1/shot', [
'query' => [
'access_key' => 'YOUR_API_KEY',
'url' => 'https://stripe.com',
],
]);
file_put_contents('shot.webp', $response->getBody()->getContents());
Equivalent requests:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo supports full-page captures with lazy images loaded, CSS-selector element shots, dark mode, device presets, custom viewports, retina scale, PDFs, custom CSS and JavaScript, clicks, selector or network-idle waits, blocking rules, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, selectable cache TTLs, signed links, async jobs with signed webhooks, bulk capture for up to 100 URLs per call, usage reporting, and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which can simplify migration.
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is available on every plan, and yearly billing gives two months free. Create a free ScreenshotNeo account.
FAQ
Can Guzzle scrape a site behind a login?
Yes, when the login flow is ordinary HTTP. Submit the form or API request with a CookieJar, then reuse that jar for authenticated pages. Multi-factor prompts, JavaScript challenges, and WebAuthn generally require browser automation.
How do I limit a scraper’s concurrency?
Use Guzzle’s promise APIs with a small pool or queue, and keep the number of in-flight requests below the target’s documented limit. Add backoff for 429 responses.
Why does my final URL differ from the URL I requested?
Redirects are enabled by default. Disable them to inspect the 3xx response, or enable track_redirects and read the history headers.
Should I parse HTML with regular expressions?
Use a DOM or HTML parser. Regular expressions are brittle when attributes, whitespace, nesting, or markup change.
Can I use Guzzle and ScreenshotNeo together?
Yes. Use Guzzle for APIs and server-rendered pages, and call ScreenshotNeo when you need a rendered screenshot or PDF without maintaining browser infrastructure.


