ScreenshotNeo

BlogHTML to image & PDF

How to Parse PDF Files in PHP

Extract PDF text, read pages, handle coordinates and passwords, or import existing pages into new PDFs with practical PHP examples.

By the ScreenshotNeo team1 October 20267 min read

For ordinary text extraction in PHP, install smalot/pdfparser with Composer, call parseFile() for a path or parseContent() for PDF bytes, then read the result with getText(). Use FPDI when the job is importing existing PDF pages into a newly generated PDF. FPDI does not edit the source document in place.

Choose the right PHP PDF tool

Need Recommended path What it does
Extract searchable text Smalot PdfParser Reads text objects from local files or in-memory PDF bytes.
Extract text from one page Smalot PdfParser Use getPages() and call getText() on the selected page.
Use coordinates or transformation matrices Smalot PdfParser Use getDataTm() for layout-aware processing.
Assemble a new PDF from existing pages FPDI Imports pages into FPDF, TCPDF or tFPDF output.
Encrypted or password-protected input FPDI PDF-Parser Adds parser support for difficult input; OpenSSL is required for encrypted files.
Commercial maintained extraction component SetaPDF-Extractor Pure-PHP extraction of text, words and coordinates.

Install a PDF parser with Composer

composer require smalot/pdfparser

For page import with FPDF:

composer require setasign/fpdf setasign/fpdi

With TCPDF, use the TCPDF package and the FPDI integration documented by Setasign. FPDI v2 requires PHP above 7.2 and Zlib. If encrypted or password-protected PDFs must be parsed through FPDI, install FPDI PDF-Parser and ensure OpenSSL is available.

Extract all text from a local PDF

This complete script reads a file, validates its size, catches parser errors and writes extracted text to standard output.

<?php
declare(strict_types=1);

require __DIR__ . '/vendor/autoload.php';

use Smalot\PdfParser\Parser;
use Throwable;

$path = $argv[1] ?? null;
if ($path === null) {
    fwrite(STDERR, "Usage: php extract.php document.pdf\n");
    exit(1);
}

if (!is_file($path) || !is_readable($path)) {
    throw new RuntimeException('PDF file does not exist or is not readable.');
}

$maxBytes = 25 * 1024 * 1024;
$size = filesize($path);
if ($size === false || $size > $maxBytes) {
    throw new RuntimeException('PDF exceeds the configured 25 MB limit.');
}

try {
    $parser = new Parser();
    $pdf = $parser->parseFile($path);
    echo $pdf->getText();
} catch (Throwable $e) {
    fwrite(STDERR, "Could not parse PDF: {$e->getMessage()}\n");
    exit(2);
}

The documented workflow is to create a parser object and point it to a file. getText() returns the document text as a string.

Parse PDF bytes from an upload or HTTP response

Use parseContent() when the PDF is already in memory. The example below handles a PHP upload without writing the original bytes to a permanent location.

<?php
declare(strict_types=1);

require __DIR__ . '/vendor/autoload.php';

use Smalot\PdfParser\Parser;
use Throwable;

if (!isset($_FILES['pdf']) || $_FILES['pdf']['error'] !== UPLOAD_ERR_OK) {
    http_response_code(400);
    exit('Upload a PDF file.');
}

$tmp = $_FILES['pdf']['tmp_name'];
$size = (int) $_FILES['pdf']['size'];
if ($size < 1 || $size > 25 * 1024 * 1024) {
    http_response_code(413);
    exit('The PDF must be between 1 byte and 25 MB.');
}

$bytes = file_get_contents($tmp);
if ($bytes === false || substr($bytes, 0, 5) !== '%PDF-') {
    http_response_code(400);
    exit('The uploaded file is not recognized as a PDF.');
}

try {
    $pdf = (new Parser())->parseContent($bytes);
    header('Content-Type: text/plain; charset=utf-8');
    echo $pdf->getText();
} catch (Throwable $e) {
    http_response_code(422);
    echo 'PDF parsing failed.';
}

Read a single page

<?php
require __DIR__ . '/vendor/autoload.php';

use Smalot\PdfParser\Parser;

$pdf = (new Parser())->parseFile(__DIR__ . '/document.pdf');
$pages = $pdf->getPages();
$pageNumber = 1;

if (!isset($pages[$pageNumber - 1])) {
    throw new OutOfBoundsException('Page does not exist.');
}

echo $pages[$pageNumber - 1]->getText();

getPages() is zero-indexed, so page 1 is element 0. For a quick extraction limit, the documentation also demonstrates $pdf->getText(5); treat that argument as an application-level limit and confirm behavior against the installed library version.

Extract text with coordinates

Plain text can lose columns and visual order. Smalot exposes getDataTm(), whose transformation matrix includes x and y positions. This lets you retain words with their locations and then group them into rows or columns.

<?php
require __DIR__ . '/vendor/autoload.php';

use Smalot\PdfParser\Parser;

$pdf = (new Parser())->parseFile(__DIR__ . '/invoice.pdf');
foreach ($pdf->getPages() as $pageIndex => $page) {
    foreach ($page->getDataTm() as $item) {
        // Inspect the installed library's returned structure before production use.
        $text = $item[0] ?? '';
        $matrix = $item[1] ?? [];
        $x = $matrix[4] ?? null;
        $y = $matrix[5] ?? null;
        if ($text !== '') {
            printf("page=%d x=%s y=%s text=%s\n", $pageIndex + 1, (string) $x, (string) $y, trim($text));
        }
    }
}

PDF reading order varies by producer. Validate row grouping, coordinate units and rotated text against representative invoices or forms before relying on it.

Import existing PDF pages into a new PDF with FPDI

FPDI extends a PDF generator so you can import each source page and write a new document. The source file remains unchanged.

<?php
declare(strict_types=1);

require __DIR__ . '/vendor/autoload.php';

use setasign\Fpdi\Fpdi;

$source = __DIR__ . '/source.pdf';
$output = __DIR__ . '/copy.pdf';

$pdf = new Fpdi();
$pageCount = $pdf->setSourceFile($source);

for ($pageNo = 1; $pageNo <= $pageCount; $pageNo++) {
    $templateId = $pdf->importPage($pageNo);
    $size = $pdf->getTemplateSize($templateId);
    $pdf->AddPage($size['orientation'], [$size['width'], $size['height']]);
    $pdf->useTemplate($templateId);
}

$pdf->Output('F', $output);
echo "Wrote {$output}\n";

setSourceFile() returns the source page count. Use the TCPDF-specific FPDI class when your generator is TCPDF, as shown in the FPDI manual.

Or skip the browser setup

If you need a visual capture of a PDF viewer or a page that hosts a PDF, ScreenshotNeo returns an image or PDF from one GET request. See the ScreenshotNeo API docs.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests; r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90); open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups and chat widgets are removed before the shot. Bot checks, blank pages and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Encrypted and password-protected PDFs

Do not assume a normal parser can open every protected document. FPDI PDF-Parser requires PHP above 7.2, Zlib and OpenSSL for encrypted or password-protected input, and your application still needs the correct password.

composer require setasign/fpdi setasign/fpdi-pdf-parser

Handle wrong passwords and unsupported encryption as expected input failures. Catch the package exception, avoid logging passwords, and return a useful application error. Parsing and writing can be CPU- and memory-intensive because a PDF may contain thousands of objects; configure execution and memory limits for your workload.

Scanned PDFs and OCR

A scanned PDF may contain only raster images. Smalot PdfParser and FPDI read PDF structures and text objects; they do not guarantee OCR for image-only pages. Detect empty extraction, render or rasterize the relevant pages, and send those images through an OCR system when searchable text is required.

Production checklist

  • Install dependencies with Composer and commit composer.lock.
  • Check PHP, Zlib and, for encrypted FPDI parsing, OpenSSL in every deployment environment.
  • Bound upload size, request size, execution time and memory.
  • Validate file signatures and keep uploaded files outside public directories.
  • Test ordinary, multi-page, compressed, malformed, scanned and protected PDFs.
  • Catch parser exceptions and return stable application-level errors.
  • Keep coordinate-based parsing fixtures for every document layout you support.

Troubleshooting

Symptom Likely cause Fix
Class not found Composer autoloader is missing or the package is not installed. Run Composer in the deployment build and require vendor/autoload.php.
Empty string from getText() The PDF is scanned, uses unusual encoding, or contains no text objects. Inspect a page directly; if it is image-only, use OCR. Test another producer’s PDF.
Only some pages extract Malformed objects or page-specific encoding. Process pages individually, record the failing page, and test with a current package version.
Upload causes memory exhaustion The whole file and parser object exceed the PHP memory limit. Lower the upload limit, increase memory for a controlled worker, or process jobs asynchronously.
FPDI cannot open an encrypted file FPDI PDF-Parser or OpenSSL is missing, or the password is wrong. Install the parser extension, verify OpenSSL, supply the correct password and catch the exception.
Imported page is cropped or stretched The destination page size does not match the imported template. Use getTemplateSize() and pass its orientation and dimensions to AddPage().
Text order is wrong PDF drawing order does not match reading order. Use getDataTm(), group by y position, sort by x position and validate against fixtures.
Timeouts on large files Parsing or writing many PDF objects is expensive. Set realistic limits, queue large jobs and report progress or a retryable failure.

Performance, reliability and cost

Parsing cost depends on page count, embedded objects, compression, fonts and images. Large or malformed PDFs can consume substantial CPU and memory. Measure representative files in your own PHP version and deployment rather than assuming page count alone predicts runtime.

For reliable services, isolate parsing workers, enforce size and time limits, make retries idempotent, and retain the original file long enough to reproduce failures under your retention policy. Do not treat a successful parser call as proof that a scanned document contains usable text.

Smalot PdfParser and FPDI are open-source Composer dependencies. A commercial component such as SetaPDF-Extractor may be appropriate when maintained support, word-level and coordinate extraction, encryption handling or broader document operations justify its license cost.

FAQ

Can PHP read a PDF without converting it first?

Yes. Smalot PdfParser reads PDF text objects directly from a path or byte string. Image-only pages still require OCR.

Should I use FPDI to extract text?

Use FPDI for importing pages into a new PDF. Use a text parser when the output is text or coordinates.

Can FPDI edit the original file?

No. Its normal workflow imports source pages and writes a separate generated PDF.

What should I do when layout matters?

Read coordinate data with getDataTm(), then build and test your own row and column rules for each document layout.

Is a password-protected PDF always supported?

Support depends on the encryption, correct password and installed parser components. Use FPDI PDF-Parser with OpenSSL and handle failures explicitly.