ScreenshotNeo

BlogGuides

Web Scraping with PHP: Detailed Examples and Code

Learn PHP web scraping with cURL, DOMDocument, XPath, DomCrawler, pagination, retries, JavaScript limits, and production troubleshooting.

By the ScreenshotNeo team1 October 20268 min read

Direct answer: PHP web scraping usually follows five steps: request the page with cURL or an HTTP client, check both the transport result and HTTP status, parse the returned HTML, select stable fields with XPath or CSS selectors, and validate the extracted data. A normal HTTP request does not execute the page’s client-side JavaScript, so use browser automation only when the required content is absent from the server response and you have permission to access it.

What PHP web scraping does

A scraper downloads an HTTP response and processes it as data. PHP’s cURL extension handles HTTP and HTTPS requests; DOMDocument, XPath, and Symfony DomCrawler help you navigate the returned markup.

There are two separate failure classes:

  • Transport failure: DNS, TLS, connection, timeout, or libcurl errors. curl_exec() returns false.
  • HTTP failure: the server returns a status such as 404, 403, 429, or 500. The request can still be a successful cURL execution, so inspect the HTTP status separately.

See the PHP cURL execution documentation and cURL handle documentation.

1. Fetch HTML with PHP cURL

Install PHP with the cURL extension enabled. This standalone example follows redirects, sets explicit timeouts, identifies the client honestly, and rejects unexpected statuses.

<?php
declare(strict_types=1);

$url = 'https://example.com/';
$ch = curl_init($url);

if ($ch === false) {
    throw new RuntimeException('Could not initialize cURL');
}

curl_setopt_array($ch, [
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_FOLLOWLOCATION => true,
    CURLOPT_MAXREDIRS => 5,
    CURLOPT_CONNECTTIMEOUT => 10,
    CURLOPT_TIMEOUT => 30,
    CURLOPT_USERAGENT => 'ExampleResearchBot/1.0 (contact: admin@example.com)',
    CURLOPT_HTTPHEADER => [
        'Accept: text/html,application/xhtml+xml',
    ],
]);

$html = curl_exec($ch);
$status = curl_getinfo($ch, CURLINFO_RESPONSE_CODE);
$error = curl_error($ch);
curl_close($ch);

if ($html === false) {
    throw new RuntimeException("Request failed: {$error}");
}
if ($status < 200 || $status >= 300) {
    throw new RuntimeException("Unexpected HTTP status: {$status}");
}

echo $html;

Use a real contact address where appropriate. Do not disable TLS verification to hide certificate problems.

2. Parse HTML with DOMDocument and XPath

DOMDocument::loadHTML() is convenient for many pages, but its tree-building rules are not the same as an HTML5 browser parser. Suppress and inspect parser warnings deliberately, then test selectors against saved response fixtures.

<?php
$dom = new DOMDocument();
libxml_use_internal_errors(true);
$loaded = $dom->loadHTML($html);
$parseErrors = libxml_get_errors();
libxml_clear_errors();

if ($loaded === false) {
    throw new RuntimeException('Could not parse HTML');
}

$xpath = new DOMXPath($dom);
$headings = $xpath->query('//article//h2');

if ($headings === false) {
    throw new RuntimeException('Invalid XPath expression');
}

foreach ($headings as $heading) {
    $value = trim($heading->textContent);
    if ($value !== '') {
        echo $value, PHP_EOL;
    }
}

For HTML5-conforming parsing, PHP’s manual points to Dom\HTMLDocument::createFromString() and createFromFile(), added in PHP 8.4. Do not call those APIs on older runtimes; check your deployment version first. Read the DOMDocument::loadHTML manual.

3. Extract stable fields

Prefer semantic containers and attributes that describe the data. Avoid selectors tied to generated class names or a page’s visual layout.

<?php
$records = [];
$nodes = $xpath->query('//article[contains(@class, "product")]');

foreach ($nodes as $node) {
    $titleNode = $xpath->query('.//h2 | .//h3', $node)->item(0);
    $priceNode = $xpath->query('.//*[contains(@class, "price")]', $node)->item(0);
    $linkNode = $xpath->query('.//a[@href]', $node)->item(0);

    $href = $linkNode instanceof DOMElement ? $linkNode->getAttribute('href') : null;
    $records[] = [
        'title' => $titleNode ? trim($titleNode->textContent) : null,
        'price' => $priceNode ? trim($priceNode->textContent) : null,
        'url' => $href,
    ];
}

foreach ($records as $record) {
    if ($record['title'] === null) {
        continue;
    }
    echo json_encode($record, JSON_UNESCAPED_SLASHES), PHP_EOL;
}

Normalize whitespace, validate required fields, preserve the source URL, and log records that fail validation instead of silently emitting incomplete data.

4. Use Symfony DomCrawler for convenient traversal

In a Composer or Symfony application, DomCrawler provides a navigation layer for HTML and XML. Install it with:

composer require symfony/dom-crawler symfony/css-selector

Load Composer’s autoloader, then use CSS selectors or XPath:

<?php
require __DIR__ . '/vendor/autoload.php';

use Symfony\Component\DomCrawler\Crawler;

$crawler = new Crawler($html, $url);

$titles = $crawler->filter('article h2, article h3')->each(
    static fn (Crawler $node): string => trim($node->text())
);

$links = $crawler->filter('article a[href]')->each(
    static fn (Crawler $node): array => [
        'text' => trim($node->text()),
        'href' => $node->attr('href'),
    ]
);

print_r($titles);
print_r($links);

DomCrawler is for navigation, not general DOM manipulation or re-dumping. Its parser may correct malformed markup, so inspect unexpected selections. The Symfony DomCrawler documentation covers traversal, XPath, and CSS selectors.

5. Fetch and crawl with Symfony HttpClient

Symfony also provides an HTTP browser integration. Keep the transport and crawler responsibilities explicit:

composer require symfony/http-client symfony/browser-kit symfony/dom-crawler
<?php
require __DIR__ . '/vendor/autoload.php';

use Symfony\Component\BrowserKit\HttpBrowser;
use Symfony\Component\HttpClient\HttpClient;

$client = HttpClient::create([
    'timeout' => 30,
    'headers' => ['User-Agent' => 'ExampleResearchBot/1.0 (contact: admin@example.com)'],
]);
$browser = new HttpBrowser($client);
$crawler = $browser->request('GET', 'https://example.com/');

foreach ($crawler->filter('article h2') as $node) {
    echo trim($node->textContent), PHP_EOL;
}

BrowserKit’s testing-oriented client and an external HTTP browser are not interchangeable in every configuration. Instantiate the client you actually need and follow Symfony’s BrowserKit documentation.

6. Pagination, retries, and pacing

Only request pages the site permits you to crawl. Follow links from the response, cap the page count, deduplicate URLs, and stop when a next link is absent.

<?php
function fetch(string $url, int $attempts = 3): string
{
    for ($attempt = 1; $attempt <= $attempts; $attempt++) {
        $ch = curl_init($url);
        curl_setopt_array($ch, [
            CURLOPT_RETURNTRANSFER => true,
            CURLOPT_FOLLOWLOCATION => true,
            CURLOPT_CONNECTTIMEOUT => 10,
            CURLOPT_TIMEOUT => 30,
            CURLOPT_USERAGENT => 'ExampleResearchBot/1.0 (contact: admin@example.com)',
        ]);
        $body = curl_exec($ch);
        $status = curl_getinfo($ch, CURLINFO_RESPONSE_CODE);
        $error = curl_error($ch);
        curl_close($ch);

        if ($body !== false && $status >= 200 && $status < 300) {
            return $body;
        }
        if ($status === 403 || $status === 429) {
            throw new RuntimeException("Access denied or throttled: HTTP {$status}");
        }
        if ($attempt === $attempts) {
            throw new RuntimeException("Fetch failed: HTTP {$status}; {$error}");
        }
        sleep(2 ** ($attempt - 1));
    }
    throw new LogicException('Unreachable');
}

$base = 'https://example.com';
$next = $base . '/products?page=1';
$seen = [];
$page = 0;
while ($next !== null && $page++ < 20) {
    if (isset($seen[$next])) break;
    $seen[$next] = true;
    $html = fetch($next);
    // Parse and persist this page here.
    $next = null; // Set to a validated, normalized next-page URL when present.
    usleep(500000); // Choose a conservative delay for the target.
}

Retry transient network failures and selected 5xx responses with exponential backoff. Do not retry authentication failures, access denials, or rate limits indefinitely. Cache responses when freshness allows it.

7. Static HTML versus JavaScript-rendered pages

cURL and Symfony HTTP clients receive the server response; they do not run browser JavaScript. If the HTML contains an empty app root and the data arrives through a client-side API, inspect whether a documented, permitted endpoint can provide the data. Use browser automation only when necessary and authorized. Do not bypass authentication, CAPTCHAs, bot checks, paywalls, or technical access controls.

8. Responsible crawling checklist

  • Review the site’s terms and applicable permissions.
  • Read and honor /robots.txt where relevant. RFC 9309 describes it as a requested crawler protocol and says, “These rules are not a form of access authorization.” Read the IETF RFC 9309.
  • Use an honest User-Agent and contact address.
  • Request only the pages and fields needed.
  • Use conservative concurrency, timeouts, caching, and backoff.
  • Stop when the site returns access-denied or throttling responses.
  • Avoid collecting personal or sensitive information without a valid basis.

9. Common errors and fixes

Symptom Likely cause Fix
curl_exec() returns false DNS, TLS, connection, or timeout failure Log curl_error(), verify DNS and certificates, and increase timeouts only when justified.
HTTP 404/403/429 Missing resource, denied request, or throttling Check the URL and permissions; slow down; stop rather than trying to bypass controls.
Empty selector result Selector mismatch or content rendered by JavaScript Save the response, inspect its HTML, verify the selector, and determine whether the data exists server-side.
Malformed or shifted DOM loadHTML() is not an HTML5 browser parser Use PHP 8.4’s HTML5 APIs where available, or account for parser behavior and test fixtures.
SSL certificate error Outdated CA bundle, hostname mismatch, or interception Fix the trust store or network configuration. Keep TLS verification enabled.
Memory grows during a crawl Keeping every HTML document or DOM in memory Process one page at a time, persist results, and release references.
Duplicate records Pagination or canonical URL variations Normalize URLs and deduplicate by a stable ID or canonical URL.

10. Performance, reliability, and cost

  • Measure first: record request duration, status, response size, parse time, and extraction counts. The dossier contains no reliable cross-library benchmark, so avoid universal speed claims.
  • Control concurrency: more workers can increase throttling and failures. Start conservatively and observe responses.
  • Cache safely: cache pages when permitted and when freshness requirements allow; use a bounded cache and clear invalidation rules.
  • Bound work: set connection and total timeouts, maximum redirects, page limits, response-size limits, and retry counts.
  • Persist checkpoints: store processed URLs and extracted records so a restart does not repeat the entire crawl.
  • Protect data: avoid logging cookies, Authorization headers, or sensitive page content.

Or skip the browser setup

If your goal is a clean screenshot or PDF rather than raw HTML extraction, ScreenshotNeo provides a single website screenshot API request. See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Can PHP scrape a page without JavaScript?

Yes, when the needed data is present in the HTTP response. PHP’s HTTP clients do not execute browser JavaScript.

Should I use XPath or CSS selectors?

Native DOM APIs provide XPath. DomCrawler supports XPath and CSS selectors when Symfony CssSelector is installed. Choose the syntax your project team can maintain.

Is robots.txt permission to scrape?

No. RFC 9309 says robots rules are not access authorization. Review terms, permissions, and applicable requirements separately.

When should I choose DomCrawler?

Use it when you already have a Composer or Symfony project and want concise traversal and selector APIs. Native DOM APIs keep dependencies smaller.

Where can I read more?

The php|architect publisher sample for Web Scraping with PHP, 2nd edition, includes DOM interoperability and Symfony library material for further study.