Install and Use a PHP PDF Parser with Composer
Install smalot/pdfparser with Composer, extract text and metadata in PHP, handle limits, troubleshoot errors, and automate PDF capture with ScreenshotNeo.
Direct answer: In your PHP project, run composer require smalot/pdfparser. Then include Composer’s autoloader, create \Smalot\PdfParser\Parser, parse a file with parseFile(), and read its text with getText().
composer require smalot/pdfparser
<?php
require __DIR__ . '/vendor/autoload.php';
$parser = new \\Smalot\\PdfParser\\Parser();
$pdf = $parser->parseFile(__DIR__ . '/document.pdf');
$text = $pdf->getText();
echo $text;
The package’s documented example uses this same sequence. See the project documentation and Composer’s dependency guide for the primary references.
1. Check your PHP project before installing
Run these commands from the directory containing your application:
php -v
composer --version
php -m | grep -E 'iconv|zlib|mbstring'
smalot/pdfparser declares PHP >=7.1, the iconv and zlib extensions, and symfony/polyfill-mbstring ^1.18. Composer checks PHP and extensions as platform requirements, so check the PHP binary used by your web server or worker, not only the one in your interactive shell.
2. Install the parser with Composer
- Change to the application root.
- Run
composer require smalot/pdfparser. - Commit both
composer.jsonandcomposer.lockfor an application.
cd /path/to/your-project
composer require smalot/pdfparser
git add composer.json composer.lock
git commit -m "Add PDF parser"
composer require updates the dependency declaration and lockfile. On deployment, use composer install so the exact versions recorded in the lockfile are installed. Use composer update when you intentionally want Composer to resolve newer versions within your constraints and rewrite the lockfile.
3. Extract text from a local PDF
<?php
declare(strict_types=1);
require __DIR__ . '/vendor/autoload.php';
use Smalot\\PdfParser\\Parser;
$path = __DIR__ . '/document.pdf';
if (!is_file($path) || !is_readable($path)) {
throw new RuntimeException("PDF is missing or unreadable: {$path}");
}
$parser = new Parser();
$pdf = $parser->parseFile($path);
$text = $pdf->getText();
echo $text;
Keep uploaded files outside a public directory, validate their size and type before parsing, and remove temporary files after processing. Parsing untrusted PDFs can consume substantial memory, so use a queue or worker with process limits for large batches.
Read pages individually
<?php
require __DIR__ . '/vendor/autoload.php';
use Smalot\\PdfParser\\Parser;
$pdf = (new Parser())->parseFile(__DIR__ . '/document.pdf');
foreach ($pdf->getPages() as $number => $page) {
printf("--- Page %d ---%s", $number + 1, PHP_EOL);
echo $page->getText();
echo PHP_EOL;
}
The README documents ordered-page text extraction. Page-level processing is useful when indexing, displaying previews, or attaching errors to a source page.
Read document metadata
<?php
require __DIR__ . '/vendor/autoload.php';
use Smalot\\PdfParser\\Parser;
$pdf = (new Parser())->parseFile(__DIR__ . '/document.pdf');
$details = $pdf->getDetails();
foreach ($details as $key => $value) {
if (is_scalar($value)) {
echo $key . ': ' . $value . PHP_EOL;
}
}
Metadata fields depend on what the PDF contains. Treat them as optional input and normalize values before storing them.
4. Parse an uploaded PDF safely
<?php
declare(strict_types=1);
require __DIR__ . '/vendor/autoload.php';
use Smalot\\PdfParser\\Parser;
if ($_SERVER['REQUEST_METHOD'] !== 'POST' || !isset($_FILES['pdf'])) {
http_response_code(400);
exit('Upload a PDF with the pdf field.');
}
$file = $_FILES['pdf'];
if ($file['error'] !== UPLOAD_ERR_OK) {
throw new RuntimeException('Upload failed with code ' . $file['error']);
}
$maxBytes = 20 * 1024 * 1024;
if ($file['size'] > $maxBytes) {
throw new RuntimeException('PDF exceeds the 20 MB application limit.');
}
$mime = (new finfo(FILEINFO_MIME_TYPE))->file($file['tmp_name']);
if ($mime !== 'application/pdf') {
throw new RuntimeException('The uploaded file is not identified as a PDF.');
}
$parser = new Parser();
$pdf = $parser->parseFile($file['tmp_name']);
header('Content-Type: text/plain; charset=utf-8');
echo $pdf->getText();
MIME detection is only a first check. For a public upload service, add authentication, rate limits, storage quotas, malware scanning and a job timeout appropriate to your environment.
5. What the library supports
- PDF headers and objects.
- Document metadata.
- Text extraction from ordered pages.
- Compressed PDFs.
- MAC OS Roman text.
- Hexadecimal and octal encoded text.
- Custom parser configuration documented by the project.
The reviewed documentation does not claim OCR. A scanned PDF containing only page images may produce little or no text. The README also states that secured documents and PDF form data are unsupported.
6. Composer versioning and deployment
Do not hard-code a package version unless you have reviewed compatibility and deliberately chosen a constraint. Available Packagist snapshots showed conflicting stable and beta listings, so check Packagist immediately before publishing or pin a version after your own review.
# Development or CI installation with the committed lockfile
composer install --no-interaction --prefer-dist
# Deliberately resolve updates, then review the diff
composer update smalot/pdfparser
The package is licensed under LGPL-3.0. Review how that license fits your distribution model. Its README describes maintenance as limited: it remains compatible with supported PHP versions, but there is no active feature development and pull requests may not be reviewed promptly.
7. Common errors and fixes
| Error or symptom | Likely cause | Fix |
|---|---|---|
Class "Smalot\\PdfParser\\Parser" not found |
Composer’s autoloader was not included, or the command ran in another directory. | Require __DIR__ . '/vendor/autoload.php' and run Composer in the project root. |
| Composer reports a PHP or extension conflict | The runtime is below PHP 7.1 or lacks iconv/zlib. |
Enable the extensions for the runtime that executes the application, then rerun Composer. |
| Text is empty | The PDF may be scanned images, encrypted, malformed, or use text encoding your file needs you to inspect. | Try a text-based PDF, inspect the security settings, and do not assume this library performs OCR. |
| Only some characters are wrong | Font encoding or unusual character maps. | Test representative files and preserve the original PDF. Normalize extracted text only after checking the expected encoding. |
| Parsing times out or exhausts memory | Large or complex PDFs require more resources. | Set upload limits, process asynchronously, cap worker memory, and record the file and page that failed. |
| Deployment works locally but fails in production | Production did not install the lockfile or uses a different PHP binary/extensions. | Run composer install in the build, deploy vendor/ or install it during release, and compare php -v and php -m. |
8. Performance, reliability and cost
- Performance: parsing cost depends on file size, page count, compression and object complexity. Measure with your real PDFs; the research does not establish an independent benchmark.
- Memory: avoid loading unbounded user files in a web request. Queue large files and set explicit worker limits.
- Reliability: keep the original upload, log parser exceptions, make jobs retryable, and record parser and PHP versions with each result.
- Cost: the package is installed through Composer rather than a per-document API charge. Your infrastructure still incurs storage, CPU and queue costs.
- Security: treat PDFs as untrusted input, keep them private, validate size and type, and remove temporary files according to your retention policy.
9. Or skip the browser setup
If your workflow starts with a web page that must become a PDF before PHP processes it, ScreenshotNeo can capture the page through one request. Its PDF endpoint supports paper size, margins, landscape mode and page ranges. Cookie banners, newsletter popups and chat widgets are removed before the capture; bot checks, blank pages and failed loads are not billed. The service also provides an MCP server so AI agents can call capture_pdf and related tools.
Read the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-d format=pdf \
-o page.pdf
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={
"access_key": "YOUR_API_KEY",
"url": "https://stripe.com",
"format": "pdf",
},
timeout=90,
)
r.raise_for_status()
open("page.pdf", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com',
format: 'pdf'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo request failed: ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('page.pdf', buffer);
ScreenshotNeo includes response headers identifying the page verdict and whether the request was billed. It offers 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
10. FAQ
Can this parser extract text from a scanned PDF?
Not reliably. The reviewed documentation does not claim OCR, so scanned image-only pages may return no useful text.
Does it support password-protected PDFs?
The README says secured documents are unsupported.
Should I commit the vendor directory?
For most applications, commit composer.json and composer.lock, then run composer install during deployment according to your build process.
How do I keep page boundaries?
Iterate over $pdf->getPages() and process each page separately instead of using only the document-level getText().
What should I evaluate before adopting it?
Check your PHP version, required extensions, representative PDFs, encrypted-document and form requirements, LGPL-3.0 compatibility, and the project’s limited-maintenance status.


