Export Specific PDF Pages in PHP with Guzzle
Download a PDF with Guzzle, validate its pages, and export a new PDF containing only the pages you select with FPDI.
Direct answer: Guzzle downloads the PDF; it does not understand PDF pages. Use Guzzle to fetch the document, then use FPDI with FPDF, TCPDF, or tFPDF to import the requested pages into a newly generated PDF. Page numbers in FPDI’s documented workflow are 1-based.
How the workflow works
- Install Guzzle, FPDF, and FPDI with Composer.
- Request the source PDF with a timeout and streaming enabled.
- Check the HTTP status and content type before parsing.
- Save the response to a temporary file.
- Use
setSourceFile()to discover the page count. - Validate and import each requested page with
importPage(). - Create a new output PDF and remove the temporary file.
Guzzle is an HTTP client, as its maintainers describe in the official repository. FPDI is a set of PHP classes for reading existing PDF pages and using them as templates in FPDF, according to the official FPDI repository.
Install the dependencies
composer require guzzlehttp/guzzle setasign/fpdf setasign/fpdi
The package names and Composer setup follow the Guzzle and FPDI documentation. Verify the versions supported by your PHP runtime before deploying.
Complete PHP example
<?php
require __DIR__ . '/vendor/autoload.php';
use GuzzleHttp\Client;
use GuzzleHttp\Exception\GuzzleException;
use setasign\Fpdi\Fpdi;
$sourceUrl = 'https://example.com/source.pdf';
$outputPath = __DIR__ . '/selected-pages.pdf';
$requestedPages = [1, 3, 5];
$tmpPath = tempnam(sys_get_temp_dir(), 'pdf_');
if ($tmpPath === false) {
throw new RuntimeException('Could not create a temporary file.');
}
try {
$client = new Client([
'timeout' => 30,
'connect_timeout' => 10,
'allow_redirects' => [
'max' => 5,
'strict' => true,
],
'http_errors' => false,
]);
$response = $client->request('GET', $sourceUrl, [
'stream' => true,
'headers' => [
'Accept' => 'application/pdf',
],
]);
$status = $response->getStatusCode();
if ($status < 200 || $status >= 300) {
throw new RuntimeException("PDF download failed with HTTP status {$status}.");
}
$contentType = strtolower($response->getHeaderLine('Content-Type'));
if ($contentType !== '' && !str_contains($contentType, 'application/pdf')) {
throw new RuntimeException("Expected application/pdf, received {$contentType}.");
}
$body = $response->getBody();
$handle = fopen($tmpPath, 'wb');
if ($handle === false) {
throw new RuntimeException('Could not open the temporary file.');
}
while (!$body->eof()) {
$chunk = $body->read(1024 * 1024);
if ($chunk === '') {
break;
}
fwrite($handle, $chunk);
}
fclose($handle);
$pdf = new Fpdi();
$pageCount = $pdf->setSourceFile($tmpPath);
$pages = array_values(array_unique(array_map('intval', $requestedPages)));
sort($pages);
if ($pages === []) {
throw new InvalidArgumentException('Select at least one page.');
}
foreach ($pages as $pageNumber) {
if ($pageNumber < 1 || $pageNumber > $pageCount) {
throw new OutOfRangeException(
"Page {$pageNumber} is outside the document range 1-{$pageCount}."
);
}
$template = $pdf->importPage($pageNumber);
$size = $pdf->getTemplateSize($template);
$pdf->AddPage($size['orientation'], [$size['width'], $size['height']]);
$pdf->useTemplate($template);
}
$pdf->Output('F', $outputPath);
echo "Wrote {$outputPath}" . PHP_EOL;
} catch (GuzzleException $e) {
throw new RuntimeException('The PDF request failed: ' . $e->getMessage(), 0, $e);
} finally {
if (is_file($tmpPath)) {
unlink($tmpPath);
}
}
Return the selected PDF from a web endpoint
When this runs inside a controller, write the generated file to a response stream and set the PDF content type. Avoid exposing arbitrary filesystem paths.
$pdf->Output('I', 'selected-pages.pdf');
For a framework response, use its streamed-download helper where available and set Content-Disposition: attachment when the browser should download the file.
Select ranges, remove duplicates, and preserve order
FPDI imports one page at a time. Convert user input such as 1-3,7 into validated integers before calling importPage():
function parsePageList(string $input, int $pageCount): array
{
$pages = [];
foreach (explode(',', $input) as $part) {
$part = trim($part);
if ($part === '') {
continue;
}
if (preg_match('/^(\d+)-(\d+)$/', $part, $m)) {
$start = (int) $m[1];
$end = (int) $m[2];
if ($start > $end) {
throw new InvalidArgumentException('Range start must not exceed its end.');
}
for ($page = $start; $page <= $end; $page++) {
$pages[] = $page;
}
continue;
}
if (!ctype_digit($part)) {
throw new InvalidArgumentException('Invalid page expression.');
}
$pages[] = (int) $part;
}
$pages = array_values(array_unique($pages));
foreach ($pages as $page) {
if ($page < 1 || $page > $pageCount) {
throw new OutOfRangeException("Page {$page} is outside 1-{$pageCount}.");
}
}
return $pages;
}
Keep the input order if the output order matters. Sorting is useful when the product’s contract says pages must be ascending; otherwise, do not sort automatically.
Important options and edge cases
Streaming versus memory
stream => true lets you copy the PSR-7 response body to disk in chunks. This avoids holding the entire download in PHP memory. A temporary file still needs enough disk space for the source document and should be deleted in a finally block.
HTTP validation
Use http_errors => false when you want to inspect status codes yourself. Reject redirects to unexpected hosts, HTML error pages, and responses with a content type that is not PDF. A content type check is useful but not sufficient; the PDF parser remains the final validity check.
Encrypted, malformed, or unusual PDFs
FPDI creates a new document by importing pages. It is therefore selective re-creation, not in-place editing. Interactive forms, annotations, bookmarks, encryption, signatures, embedded files, and other document-level features may not survive. A signed source document will not retain a valid signature after re-creation. Test representative files when those features matter.
Duplicate pages and page numbering
FPDI’s documented page numbers are 1-based. Decide whether duplicate selections should produce duplicate output pages or be collapsed. Make that behavior explicit in your API.
Remote URLs and SSRF
Never pass an unrestricted user-supplied URL directly to a server-side downloader. Allow only https if possible, block loopback and private network ranges, limit redirects, cap download size, and apply an allowlist when the source hosts are known.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| HTTP 403 or 401 | The source requires authentication or rejects the request. | Use the required authorization headers or cookies, and confirm that automated downloads are permitted. |
| HTML saved as a PDF | The URL returned an error page or login page. | Check status, content type, redirects, and the first bytes before calling FPDI. |
| Page number out of range | The request uses a number above the imported page count. | Call setSourceFile() first and validate every number against its result. |
| Parser error or “Unable to find PDF trailer” | The download is truncated, malformed, encrypted, or not a PDF. | Confirm the complete response was written, retry transient downloads, and test the source with a PDF validator. |
| Out-of-memory failure | The whole file or many large pages are being held in memory. | Stream to a temporary file, process one source at a time, and enforce a file-size limit. |
| Output looks different | Imported pages were re-rendered into a new document. | Check fonts, transparency, annotations, and page dimensions; test with the actual files. |
| Temporary files remain | An exception or early return skipped cleanup. | Put cleanup in finally and use a scheduled janitor for abandoned files. |
Performance, reliability, and cost
- Network time is often dominated by downloading the source. Use connection and total timeouts, and retry only transient failures with bounded backoff.
- Disk streaming reduces PHP memory pressure, while FPDI still needs to parse and generate the selected pages.
- Limit source size, page count, concurrent jobs, and temporary-storage quotas at the application boundary.
- Cache downloads only when the source is safe to retain and its freshness rules are understood.
- Log the source host, status, selected page count, elapsed time, and parser errors without logging credentials or sensitive PDF contents.
- There is no authoritative benchmark for speed, memory, or maximum page count in the cited documentation. Measure representative files in your own deployment.
Or skip the browser setup
If the PDF or source page begins as a website, ScreenshotNeo can capture it through one HTTP request and supports PDF output, paper size, margins, landscape mode, and page ranges. Its browser handles cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server also lets Claude, Cursor, and other MCP clients use take_screenshot, get_page_info, and capture_pdf.
See the ScreenshotNeo API documentation for PDF parameters and current options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo includes 1,000 screenshots each month on its free plan with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Can Guzzle split a PDF by itself?
No. Guzzle transports the bytes. A PDF library such as FPDI performs page import and output generation.
Does this edit the original PDF?
No. FPDI writes a new document from imported pages.
Are FPDI page numbers zero-based?
No. The documented importPage() workflow uses 1-based page numbers.
Can I preserve digital signatures?
Do not assume so. Re-creating a PDF generally invalidates signatures and may change annotations, forms, bookmarks, or encryption.
Should I load the entire response with getContents()?
Only for documents small enough for your memory budget. For untrusted or large files, stream the body to a bounded temporary file.


