How to Parse PDFs in Laravel
Extract PDF text, page content, and metadata in Laravel with Smalot PDFParser, storage-safe code, troubleshooting, and production guidance.
To parse a PDF in Laravel, store the uploaded file with Laravel’s filesystem, install smalot/pdfparser, then call parseFile() and getText(). You can also parse bytes with parseContent(), read one page at a time, and inspect available metadata.
This approach works well for ordinary, text-based PDFs. Smalot PDFParser documents that secured documents and form-data extraction are not supported, and its documentation does not promise OCR, reliable table reconstruction, or perfect visual-layout preservation. Test representative files before making extraction quality a product requirement.
What you will build
The example below accepts an uploaded PDF, stores it on a private disk, extracts its text, and returns a JSON response. The same parser can read a local path, a byte string, individual pages, and document details.
1. Install the PDF parser
composer require smalot/pdfparser
See the Smalot PDFParser documentation for the package API. Laravel’s filesystem disks and upload storage are documented in the Laravel filesystem guide.
2. Add an upload route
use App\\Http\\Controllers\\PdfController;
use Illuminate\\Support\\Facades\\Route;
Route::post('/pdfs/parse', [PdfController::class, 'parse']);
3. Store and parse an uploaded PDF
Create app/Http/Controllers/PdfController.php:
<?php
namespace App\\Http\\Controllers;
use Illuminate\\Http\\Request;
use Smalot\\PdfParser\\Parser;
class PdfController extends Controller
{
public function parse(Request $request)
{
// Adjust the size limit to your application and infrastructure.
$validated = $request->validate([
'pdf' => ['required', 'file', 'mimetypes:application/pdf', 'max:51200'],
]);
// The default disk can be local or another configured Laravel disk.
$storedPath = $validated['pdf']->store('private/pdfs');
$absolutePath = storage_path('app/' . $storedPath);
$parser = new Parser();
$pdf = $parser->parseFile($absolutePath);
return response()->json([
'path' => $storedPath,
'text' => $pdf->getText(),
'details' => $pdf->getDetails(),
'pages' => count($pdf->getPages()),
]);
}
}
Send a multipart request with a field named pdf. Keep the stored file private unless the application explicitly needs public access.
4. Parse the uploaded bytes without a local path
Some disks, including remote object storage, may not expose a normal local filesystem path. In that case, read the object and pass its bytes to parseContent():
use Illuminate\\Support\\Facades\\Storage;
use Smalot\\PdfParser\\Parser;
$storedPath = $request->file('pdf')->store('private/pdfs', 's3');
$bytes = Storage::disk('s3')->get($storedPath);
$pdf = (new Parser())->parseContent($bytes);
$text = $pdf->getText();
For large documents, check the memory behavior of your selected parser and storage adapter before loading the entire object into PHP memory. Laravel also provides file retrieval and stream APIs; choose the path or byte workflow that matches the parser and disk you use.
5. Extract text from a particular page
$parser = new \\Smalot\\PdfParser\\Parser();
$pdf = $parser->parseFile($absolutePath);
$pages = $pdf->getPages();
$firstPageText = $pages[0]->getText();
foreach ($pages as $number => $page) {
echo "Page " . ($number + 1) . "\\n";
echo $page->getText();
}
getPages() returns page objects indexed from zero. Guard the index when the input may be empty or malformed:
$pages = $pdf->getPages();
$pageNumber = 3; // human-facing page 3
if (!isset($pages[$pageNumber - 1])) {
abort(422, 'The requested page does not exist.');
}
$text = $pages[$pageNumber - 1]->getText();
6. Read document metadata
$details = $pdf->getDetails();
return response()->json([
'title' => $details['Title'] ?? null,
'author' => $details['Author'] ?? null,
'creation_date' => $details['CreationDate'] ?? null,
'all_details' => $details,
]);
Metadata fields vary by file. Treat the returned array as optional data rather than assuming every PDF contains title, author, dates, or producer information.
7. A complete form and request example
<form action="{{ url('/pdfs/parse') }}" method="post" enctype="multipart/form-data">
@csrf
<label for="pdf">PDF file</label>
<input id="pdf" type="file" name="pdf" accept="application/pdf" required>
<button type="submit">Parse PDF</button>
</form>
curl -X POST https://example.test/pdfs/parse \
-H 'Accept: application/json' \
-F 'pdf=@./document.pdf'
8. Move parsing into a queued job
Parsing can consume noticeable CPU and memory for long documents. For a web application, store the upload, dispatch a job, and return a job identifier instead of keeping the request open.
use App\\Jobs\\ParsePdf;
$storedPath = $request->file('pdf')->store('private/pdfs');
ParsePdf::dispatch($storedPath);
return response()->json(['path' => $storedPath], 202);
<?php
namespace App\\Jobs;
use Illuminate\\Bus\\Queueable;
use Illuminate\\Contracts\\Queue\\ShouldQueue;
use Illuminate\\Foundation\\Bus\\Dispatchable;
use Illuminate\\Queue\\InteractsWithQueue;
use Illuminate\\Queue\\SerializesModels;
use Illuminate\\Support\\Facades\\Storage;
use Smalot\\PdfParser\\Parser;
class ParsePdf implements ShouldQueue
{
use Dispatchable, InteractsWithQueue, Queueable, SerializesModels;
public function __construct(public string $storedPath) {}
public function handle(): void
{
$disk = Storage::disk('local');
$absolutePath = $disk->path($this->storedPath);
$pdf = (new Parser())->parseFile($absolutePath);
$text = $pdf->getText();
// Persist $text and any selected metadata in your application.
}
}
Configure queue retries and failed-job handling for malformed, encrypted, or unexpectedly large files. Do not place raw PDF text in logs when it may contain confidential information.
9. Validate and protect uploads
- Accept only the file types your application needs and enforce a size limit at Laravel, the web server, and the reverse proxy.
- Keep uploaded PDFs on a private disk unless public delivery is required.
- Use generated storage names instead of trusting the client filename.
- Do not render extracted text as unescaped HTML. Escape it or encode it before displaying it.
- Apply authorization checks before returning stored files or extracted content.
- Consider malware scanning and retention rules for untrusted uploads.
10. Understand extraction quality
PDFs store positioned drawing instructions rather than a universal document model. Text may be returned in an order that differs from the visual reading order; columns, headers, footers, ligatures, unusual encodings, and decorative text can affect results. A PDF containing only scanned page images has no text layer for this basic parser to extract.
Tables are especially layout-sensitive. If a requirement depends on row and column boundaries, evaluate the result against real samples and consider a specialized extraction or OCR pipeline. The Smalot documentation does not establish dependable OCR, exact table reconstruction, or a universal extraction success rate.
11. Handle unsupported and difficult PDFs
Encrypted or password-protected files
Smalot PDFParser documentation identifies secured documents as unsupported. Detect parser exceptions, report that the file must be unlocked by an authorized user, and do not attempt to bypass a password.
AcroForm and other form data
Form-data extraction is listed as unsupported. If your application needs submitted field values, choose a library that explicitly supports the form technology used by your documents and verify it with sample files.
Scanned PDFs
Image-only pages require OCR. This parser’s basic text extraction path should not be presented as an OCR solution.
Remote files
Use Storage::get() with parseContent() when the configured disk does not provide a local path. For large objects, avoid unnecessary copies and measure memory usage in the queue worker.
12. Error handling pattern
use Illuminate\\Http\\JsonResponse;
use Throwable;
try {
$pdf = (new \\Smalot\\PdfParser\\Parser())->parseFile($absolutePath);
} catch (Throwable $e) {
report($e);
return response()->json([
'message' => 'The PDF could not be parsed.',
], 422);
}
Return a generic message to clients and keep detailed exception information in protected application logs. Add a correlation ID if operators need to locate a failed job without exposing file contents.
13. Choosing between PHP parsers
PrinsFrank PDFParser is another PHP option. Its maintainers describe it as low-memory, MIT licensed, and independent of external tools; those are project claims rather than independent benchmark results. Compare current PHP compatibility, license, maintenance activity, supported PDF features, memory behavior, and extraction quality using representative documents. The available evidence does not establish a speed or accuracy winner.
14. Performance, reliability, and cost notes
- Performance: store once, parse once, and persist the extracted result if it will be queried repeatedly. Queue long-running work and set worker memory limits appropriate to your corpus.
- Memory:
parseContent(file_get_contents(...))creates an in-memory representation of the input. Prefer a file path where supported, and measure remote-storage byte parsing before accepting large uploads. - Reliability: treat malformed, encrypted, image-only, and unusual-encoding PDFs as expected input cases. Record status and retry only transient storage or infrastructure failures.
- Cost: Laravel and the parser do not define a universal per-document price. Your costs come from storage, queue workers, OCR or other external services, and operational limits.
Or skip the browser setup
If your workflow also needs a clean image or PDF capture of a web page, ScreenshotNeo provides a website screenshot API. It is separate from server-side PDF text extraction, but can remove the browser automation work from a capture step.
One request returns a PNG, JPEG, WebP, or PDF:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for request options. Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 shots.
Create a free ScreenshotNeo account.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
Class "Smalot\\PdfParser\\Parser" not found |
Composer dependency is missing or autoload files are stale. | Run composer require smalot/pdfparser and deploy the updated vendor directory or run composer dump-autoload. |
| File cannot be opened | The path points to the wrong disk or the worker cannot read it. | Use the configured disk’s path method, verify permissions, and confirm the file exists before parsing. |
| Empty text | The PDF may be scanned, contain unusual encodings, or have no usable text layer. | Inspect the document visually and route image-only files to an OCR workflow. |
| Gar garbled or reordered text | PDF positioning and font encoding do not map cleanly to reading order. | Test the actual corpus, normalize output where safe, and avoid promising layout fidelity. |
| Encrypted-file exception | The document is secured. | Request an authorized unlocked copy or use a library that explicitly supports the required security model. |
| Worker runs out of memory | The parser and PDF size exceed the worker’s memory budget. | Queue the work on a suitably sized worker, avoid duplicate byte copies, enforce upload limits, and measure representative files. |
| Metadata keys are missing | PDF metadata is optional and varies by producer. | Use null-safe lookups and treat getDetails() as an untrusted, variable map. |
FAQ
Can Laravel parse a PDF without a package?
Laravel handles HTTP uploads and storage, but it does not provide a general PDF text parser. Add a package such as Smalot PDFParser or another library selected for your document requirements.
Can I parse only one page?
Parse the document, call getPages(), and read the selected page object’s getText(). The page array is zero-indexed.
Does this extract images from a PDF?
The documented workflow here extracts text and metadata. It does not establish a complete image-extraction workflow.
Should I store extracted text or parse on every request?
For repeated searches or display, parse once in a job and store the result with a reference to the original file and parser version.
Is a remote S3 PDF supported?
Use Laravel’s disk API to retrieve bytes and pass them to parseContent(), or use a local path when the disk provides one. Check memory use for large files.


