Data Extraction in PHP: Parse XML, HTML, JSON, CSV, Requests, and SQL Safely
Choose the right PHP approach for XML, HTML, JSON, CSV, request input, or database rows—with runnable examples, validation guidance, and troubleshooting.

In PHP, the right way to extract data depends on its source and shape: use DOMDocument for XML or HTML that benefits from tree navigation, XMLReader for sequential XML processing, json_decode() for JSON, and fgetcsv() for CSV. Treat request values as untrusted until validated, and use PDO placeholders when storing or querying extracted values. Parsing retrieves structure; it does not prove the data is valid, safe, or suitable for your application.
This guide covers those common formats, how to choose between them, and how to handle errors and scale. If the source is a public web page, PHP’s HTML parsers may not behave like a modern browser; a browser-based screenshot can be useful when the desired output is a visual capture rather than parsed fields.
1. Choose the extraction method
| Input | Good starting point | Use it when |
|---|---|---|
| XML | DOMDocument |
You need a navigable document tree. |
| Large XML | XMLReader |
You can process nodes sequentially without building a complete tree. |
| HTML | DOMDocument or a version-appropriate HTML5 parser |
You need to query elements; first account for parser behavior and the input’s origin. |
| JSON | json_decode() |
The source is a JSON string or API response. |
| CSV | fgetcsv() |
You need records and fields, including quoted separators. |
| HTTP request | filter_input() plus validation |
You need a request value and must check its expected format. |
| SQL result | PDO fetch methods | You need rows returned by a database query. |
Before writing a parser, answer four questions: Is the input trusted? Is it too large to keep in memory? Do you need a complete tree or just selected values? What validation does the application require after parsing?
2. Extract XML with DOMDocument
DOMDocument::load() loads XML from a file and returns a success boolean. DOM is convenient when relationships between nodes matter or you need to query the same document in several ways. Check the result and treat malformed or inaccessible input as an error.

<?php
$path = __DIR__ . '/catalog.xml';
$doc = new DOMDocument();
if (!$doc->load($path)) {
throw new RuntimeException('Could not load XML file');
}
$xpath = new DOMXPath($doc);
foreach ($xpath->query('/catalog/item') as $item) {
$id = $item->getAttribute('id');
$nameNode = $xpath->query('./name', $item)->item(0);
$name = $nameNode?->textContent ?? '';
echo htmlspecialchars($id . ': ' . $name, ENT_QUOTES | ENT_SUBSTITUTE, 'UTF-8'), PHP_EOL;
}
?>
Use XPath to select nodes by structure rather than relying on incidental whitespace or position. For data from an untrusted source, review PHP and libxml’s current XML security guidance before enabling parser flags or resolving external resources. Do not assume that a successful parse means the values meet your application’s schema or business rules.
3. Stream large XML with XMLReader
XMLReader is a forward-only pull parser: your code advances node by node. It is a natural fit when a file is too large for a whole-document tree operation or when you only need selected records. Its contents are represented internally as UTF-8 under libxml.
<?php
$reader = new XMLReader();
if (!$reader->open(__DIR__ . '/large-catalog.xml')) {
throw new RuntimeException('Could not open XML file');
}
try {
while ($reader->read()) {
if ($reader->nodeType !== XMLReader::ELEMENT || $reader->name !== 'item') {
continue;
}
$id = $reader->getAttribute('id') ?? '';
$xml = $reader->readOuterXml();
if ($xml === '') {
continue;
}
$item = new DOMDocument();
if (!$item->loadXML($xml)) {
// Record or skip malformed item according to your import policy.
continue;
}
$name = $item->getElementsByTagName('name')->item(0)?->textContent ?? '';
processItem($id, $name);
}
} finally {
$reader->close();
}
?>
Define a policy for malformed records: stop the import, skip a record with a logged error, or quarantine it for review. Streaming avoids materializing the entire source tree, but it does not automatically make downstream work constant-memory: retaining every extracted record or issuing unbounded database writes can still consume resources.
4. Extract HTML carefully
The legacy DOMDocument::loadHTML() and loadHTMLFile() APIs use libxml’s HTML parser. PHP’s HTML parsing RFC describes that parser as supporting HTML through 4.01 and explains the mismatch with modern HTML5 parsing rules. The RFC records work on a new HTML5 parser as implemented; check the PHP version and API available in the runtime you deploy before choosing it.
<?php
$html = file_get_contents(__DIR__ . '/page.html');
if ($html === false) {
throw new RuntimeException('Could not read HTML');
}
$doc = new DOMDocument();
$previous = libxml_use_internal_errors(true);
try {
if (!$doc->loadHTML($html)) {
throw new RuntimeException('Could not parse HTML');
}
$xpath = new DOMXPath($doc);
foreach ($xpath->query('//h1') as $heading) {
echo htmlspecialchars($heading->textContent, ENT_QUOTES | ENT_SUBSTITUTE, 'UTF-8'), PHP_EOL;
}
} finally {
libxml_clear_errors();
libxml_use_internal_errors($previous);
}
?>
This example parses a local file’s contents, not a JavaScript-rendered page. HTML returned by a server may omit content that appears after scripts run in a browser. If you need rendered layout or a visual record, browser capture is a different task from extracting semantic fields. For externally supplied HTML, avoid treating parsed text as trusted output: escape it for its destination, such as HTML text or an attribute.
5. Decode JSON and handle failures
json_decode() accepts a UTF-8 JSON string and returns a PHP value. Choose objects or associative arrays intentionally. Using JSON_THROW_ON_ERROR turns decode failures into exceptions instead of making a returned null ambiguous with a valid JSON null.
<?php
$json = file_get_contents(__DIR__ . '/response.json');
if ($json === false) {
throw new RuntimeException('Could not read JSON');
}
try {
$data = json_decode($json, associative: true, depth: 512, flags: JSON_THROW_ON_ERROR);
} catch (JsonException $e) {
throw new RuntimeException('Invalid JSON: ' . $e->getMessage(), previous: $e);
}
if (!is_array($data) || !isset($data['items']) || !is_array($data['items'])) {
throw new UnexpectedValueException('Expected an items array');
}
foreach ($data['items'] as $row) {
if (!is_array($row) || !isset($row['id'], $row['name'])) {
continue; // Replace with a strict validation or quarantine policy as needed.
}
processItem((string) $row['id'], (string) $row['name']);
}
?>
Use JSON_BIGINT_AS_STRING if large integer values must not be converted to imprecise floating-point numbers. Raise the nesting depth only when the data contract requires it; unbounded nesting is not a substitute for input limits. Decode errors and schema errors are different: valid JSON can still have the wrong shape or missing fields.
6. Read CSV rows, including quoted fields
Use fgetcsv(), not explode(',', $line), because quoted fields can contain commas and line breaks. The separator, enclosure, and escape parameters each have byte-level constraints. PHP 8.4 deprecates relying on the default escape value; pass it explicitly. For interoperable CSV, the manual recommends the empty string to disable PHP’s proprietary escape mechanism.
<?php
$handle = fopen(__DIR__ . '/users.csv', 'r');
if ($handle === false) {
throw new RuntimeException('Could not open CSV');
}
try {
$headers = fgetcsv($handle, null, ',', '"', '');
if ($headers === false) {
throw new RuntimeException('CSV is empty or unreadable');
}
$headers = array_map(static fn($v) => trim((string) $v), $headers);
while (($fields = fgetcsv($handle, null, ',', '"', '')) !== false) {
// Blank lines can produce a one-element array containing null.
if ($fields === [null]) {
continue;
}
if (count($fields) !== count($headers)) {
// Log or quarantine the row instead of silently misaligning columns.
continue;
}
$row = array_combine($headers, $fields);
if ($row === false) {
continue;
}
processRow($row);
}
} finally {
fclose($handle);
}
?>
For a non-comma delimiter, provide the actual single-byte separator. Locale settings can affect parsing of some one-byte encodings, so normalize or account for the source encoding when imports behave differently across environments. Treat headers as input too: reject duplicate or unexpected names if the mapping must be dependable.
7. Validate request input before using it
filter_input() reads the original raw value supplied by the SAPI. Its default FILTER_DEFAULT is an alias of FILTER_UNSAFE_RAW; it does not validate or sanitize by default. Choose validation rules for the field’s expected format, then separately encode data for its output context.
<?php
$page = filter_input(INPUT_GET, 'page', FILTER_VALIDATE_INT, [
'options' => ['min_range' => 1, 'max_range' => 500],
]);
if ($page === false || $page === null) {
http_response_code(400);
exit('Invalid page number');
}
$email = filter_input(INPUT_POST, 'email', FILTER_VALIDATE_EMAIL);
if ($email === false || $email === null) {
http_response_code(400);
exit('Invalid email address');
}
// Escape at the HTML output boundary, even after validation.
echo htmlspecialchars($email, ENT_QUOTES | ENT_SUBSTITUTE, 'UTF-8');
?>
A validation result of false means a supplied value failed the filter; null can mean the variable was absent or unavailable. Decide whether absence is an error or whether the field is optional. For optional fields, handle the missing case explicitly rather than treating it as a valid empty value.
8. Persist extracted values with PDO placeholders
Keep data values out of SQL text. PDO supports named or question-mark markers, but one statement must use one marker style. A marker stands for a complete value, not a column name, keyword, or arbitrary SQL fragment.
<?php
$pdo = new PDO($dsn, $dbUser, $dbPassword, [
PDO::ATTR_ERRMODE => PDO::ERRMODE_EXCEPTION,
]);
$stmt = $pdo->prepare(
'INSERT INTO catalog_items (external_id, name) VALUES (:external_id, :name)'
);
foreach ($items as $item) {
$stmt->execute([
'external_id' => $item['id'],
'name' => $item['name'],
]);
}
?>
For large imports, use transactions and bounded batches so you can roll back a failed batch and avoid unbounded memory use. Database driver behavior matters: PDO_MYSQL documents emulated prepares as its default. Confirm the driver and prepare settings for your deployment rather than assuming every driver uses server-side prepares in the same way. Placeholders do not replace validation, authorization, or schema constraints.
9. Extract a webpage visually with ScreenshotNeo
If your goal is a screenshot of a public page rather than structured fields from its markup, a browser capture API can handle browser rendering without wiring up a browser in your PHP application. ScreenshotNeo is a website screenshot API and MCP server. Its one-call endpoint returns an image or PDF; the options include full-page capture, element selection, viewport and device presets, wait conditions, custom CSS and JavaScript, and request controls. See the API documentation.

Or skip the browser setup
<?php
$url = 'https://stripe.com';
$query = http_build_query([
'access_key' => 'YOUR_API_KEY',
'url' => $url,
]);
$ch = curl_init('https://api.screenshotneo.com/v1/shot?' . $query);
curl_setopt_array($ch, [
CURLOPT_RETURNTRANSFER => true,
CURLOPT_TIMEOUT => 90,
]);
$body = curl_exec($ch);
$status = curl_getinfo($ch, CURLINFO_RESPONSE_CODE);
if ($body === false || $status < 200 || $status >= 300) {
throw new RuntimeException('Screenshot request failed: ' . curl_error($ch));
}
file_put_contents(__DIR__ . '/shot.webp', $body);
curl_close($ch);
?>
With ScreenshotNeo, cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed, and response headers report the page verdict and billing status. Its MCP server gives AI agents screenshot, page-info, and PDF capture tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free account and get 1,000 screenshots a month with no card.
10. Performance, reliability, and cost
- Memory: A DOM tree and a decoded JSON structure occupy memory proportional to the materialized document. Prefer XMLReader or row-by-row CSV reads for large inputs, and avoid collecting all rows before processing.
- Time: Parsing cost depends on input size, structure, and downstream work. Avoid claiming a fixed speed from the format alone; measure with representative files in the target runtime.
- Reliability: Set explicit policies for malformed documents, missing fields, duplicate CSV headers, partial imports, retries, and logging. Preserve the source or a record identifier so failures can be traced.
- Database writes: Use transactions and sensible batch sizes. Handle unique-key conflicts and retry only failures that are safe to retry.
- Web capture: Browser rendering includes network and page behavior beyond PHP parsing. Set a suitable request timeout and preserve response status or diagnostic headers when available. ScreenshotNeo’s stated billing rule is that only clean shots are billed; its responses identify verdict and billing status.
- Cost: Local PHP parsers have no per-capture API charge, but consume application compute and maintenance time. ScreenshotNeo offers a free allowance of 1,000 shots monthly; paid tiers start at $5 for 3,000, with yearly billing giving two months free. Every listed feature is on every plan.
11. Troubleshooting common failures
| Symptom | Likely cause | What to do |
|---|---|---|
json_decode() returns null or throws |
Malformed JSON, invalid UTF-8, or nesting beyond the configured depth. | Catch JsonException, inspect the source bytes and error message, and validate the payload shape separately. |
| XML load returns false | Bad path, unreadable file, or malformed XML. | Check file permissions and path; report parser errors in a controlled log and reject or quarantine the document. |
| HTML selectors find no elements | Parser differences, invalid markup, selector mismatch, or content created by JavaScript after initial HTML. | Inspect the raw response and parsed tree; check the PHP parser available and whether browser rendering is required. |
| CSV columns shift or merge | Using a string split instead of a CSV parser, wrong separator/enclosure, or a nonstandard escape convention. | Use fgetcsv() with explicit parameters that match the producer; test quoted separators and multiline fields. |
| CSV import has an extra blank record | A blank line parses as a one-field array containing null. | Detect [null] and apply an explicit skip or reject policy. |
| Request filter rejects a value unexpectedly | Missing input, unexpected type/format, or filter defaults assumed to validate. | Distinguish null from false, choose the expected validation filter, and handle optional fields explicitly. |
| PDO says an identifier cannot be bound | Placeholders bind values only, not table or column names. | Keep identifiers in a strict allowlist and bind data values. Never splice arbitrary request text into SQL. |
| Screenshot is blank or timed out | The page may require more load time, challenge a bot, or return no useful content. | Check response status and page-verdict/billing headers; adjust supported wait settings if appropriate or inspect the page manually. |
12. Practical checklist
- Identify the actual format from the response or file, not its filename alone.
- Select tree parsing for convenient navigation and streaming for sequential processing of large inputs.
- Check parser return values and convert malformed input into visible, controlled errors.
- Validate the extracted structure and field constraints before business logic or persistence.
- Escape values at the output boundary and bind values in SQL.
- Set memory, time, batch, and failure limits that match the workload.
- For visual webpage capture, choose a browser-based capture workflow instead of treating legacy HTML parsing as a browser.
13. FAQ
Does PHP have one function that extracts data from every format?
No. XML, HTML, JSON, CSV, request values, and SQL rows have different structures and failure modes, so choose the matching parser or interface.
Should extracted values be sanitized before storing them?
Validate against the field’s expected type and constraints, use parameterized SQL, and apply output encoding when displaying values. “Sanitize once” is not a universal safety rule.
Can a parsed HTML document include content rendered by JavaScript?
Not if you only parse the server-returned HTML string. A browser must execute the page scripts to capture content that is created client-side.
When is JSON better than CSV for an import?
Use the format the producer supplies and that the data contract defines. JSON represents nested objects and arrays; CSV is a row-and-field format. Neither removes the need to validate fields.


