How to save a webpage screenshot as a searchable PDF with Playwright
Create a searchable PDF from a Playwright webpage screenshot with OCR, or use Playwright’s native PDF export when print rendering is acceptable.
To save a webpage screenshot as a searchable PDF with Playwright, capture the page as an image, convert that image to a PDF, then add a text layer with OCR. Playwright’s page.screenshot() captures pixels; it does not make the resulting PDF searchable by itself. If you do not need screenshot-faithful pixels, use Playwright’s page.pdf() instead: it creates a PDF from the page and uses print CSS by default. Playwright Page API · OCRmyPDF documentation.
Choose the right PDF workflow
| Workflow | What it preserves | Searchable text | Use it when |
|---|---|---|---|
| Screenshot → image PDF → OCR | The captured pixels and screen appearance | Added by OCR; accuracy depends on the source image and OCR | The screenshot’s visual appearance matters most |
Playwright page.pdf() |
Browser-rendered PDF output | Text is generally represented by the page’s PDF rendering | A print-style document is acceptable |
These are different outputs. Playwright documents screenshot capture and PDF rendering as separate features. OCRmyPDF adds a text layer to image PDFs so text can be selected, searched, and copied. A screenshot is not automatically OCR-processed.
Prerequisites
- Node.js and a Playwright project with its browser installed.
- Python 3 and Pillow for the image-to-PDF conversion in this example.
- OCRmyPDF installed and available as the
ocrmypdfcommand. Follow its documentation for installation requirements on your operating system.
Install the JavaScript dependencies:
npm init -y
npm install playwright
npx playwright install chromium
Install Pillow in a Python environment:
python -m pip install Pillow
Install OCRmyPDF using the instructions for your platform in the OCRmyPDF documentation. Its dependencies and install method vary by operating system.
Capture a full-page screenshot and make it searchable
- Navigate to the page and wait for a page state suitable for capture.
- Capture the viewport or the full scrollable page as a PNG.
- Convert the PNG to a one-page image PDF.
- Run OCRmyPDF to add a searchable text layer.
- Inspect the PDF visually and search for text that is clearly visible in the source.
1. Capture with Playwright
Save as capture.mjs. Replace the example URL with the page you are authorized to capture.
import { chromium } from 'playwright';
const url = 'https://example.com';
const browser = await chromium.launch();
try {
const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });
await page.goto(url, { waitUntil: 'networkidle', timeout: 60000 });
await page.screenshot({ path: 'page.png', fullPage: true });
} finally {
await browser.close();
}
Run it with:
node capture.mjs
fullPage: true captures the full scrollable page; omit it to capture only the current viewport. The official Playwright screenshots guide documents saving screenshots to a path and capturing buffers.
2. Convert the screenshot to an image PDF
Save as image_to_pdf.py. Pillow converts the PNG to an RGB PDF, which is suitable for the OCR step:
from PIL import Image
with Image.open('page.png') as image:
image.convert('RGB').save('page-image.pdf', 'PDF', resolution=150.0)
Run the conversion and then OCR:
python image_to_pdf.py
ocrmypdf page-image.pdf page-searchable.pdf
The output, page-searchable.pdf, retains the screenshot as its page image and includes OCR text for searching and selection. OCR quality depends on legibility, contrast, resolution, and the content being recognized. Review the result before relying on extracted text.
One-command OCR alternative
OCRmyPDF can also create a PDF from image files as part of its image-processing workflow. Check the installed version’s command-line documentation for supported input formats and options. The explicit conversion above keeps the image-to-PDF step visible and makes the input to OCR clear.
When native Playwright PDF output is better
Use page.pdf() when the goal is a browser-rendered document rather than a PDF that reproduces the screenshot’s pixels. PDF output uses print CSS by default. To render screen styles, emulate screen media before exporting:
import { chromium } from 'playwright';
const browser = await chromium.launch();
try {
const page = await browser.newPage();
await page.goto('https://example.com', { waitUntil: 'networkidle', timeout: 60000 });
await page.emulateMedia({ media: 'screen' });
await page.pdf({ path: 'page.pdf', format: 'A4', printBackground: true });
} finally {
await browser.close();
}
For the default print rendering, omit emulateMedia. The printBackground setting includes background graphics in the PDF. Choose this route if selectable page text and print layout matter more than identical screenshot pixels. Verify the rendered PDF for the target site; page-specific print styles can change layout or hide content.
Capture options and page-state choices
| Choice | Effect | Practical guidance |
|---|---|---|
| Viewport screenshot | Captures the current visible viewport | Use when only the visible screen is needed. |
fullPage: true |
Captures the full scrollable page | Use for long articles; very tall pages create large images and may be slow to OCR. |
| Navigation wait condition | Controls when capture begins | networkidle can be unsuitable for pages with ongoing requests. Use a different wait condition or wait for a meaningful selector when appropriate. |
| Navigation timeout | Bounds how long navigation waits | Set a timeout suitable for the site and handle slow or unavailable pages. |
| PNG screenshot | Lossless image input for OCR | Usually a practical choice for text; make sure the source is legible at its capture size. |
| PDF page layout | Determines how an image is placed on a PDF page | The Pillow example creates a page from the image. For a fixed paper size, scale or tile the image deliberately; scaling a long page down can make text difficult to read. |
For dynamic pages, waiting only for navigation may capture before important content appears. Wait for a content selector, or use a deliberate delay if the page has a known rendering delay. Lazy-loaded images may not appear until scrolled into view; full-page capture alone should not be treated as proof that every site has loaded all its content. Inspect the image before processing it.
cURL, Python, and Node.js alternatives
Playwright’s browser automation APIs are available in multiple languages. The Node.js examples above show the complete workflow. Here are equivalent capture examples for Python and the Playwright CLI. The OCR and conversion steps still follow afterward.
Python Playwright
Install the package and browser:
python -m pip install playwright
python -m playwright install chromium
Save as capture.py:
import asyncio
from playwright.async_api import async_playwright
async def main():
async with async_playwright() as playwright:
browser = await playwright.chromium.launch()
try:
page = await browser.new_page(viewport={"width": 1440, "height": 1000})
await page.goto("https://example.com", wait_until="networkidle", timeout=60000)
await page.screenshot(path="page.png", full_page=True)
finally:
await browser.close()
asyncio.run(main())
python capture.py
python image_to_pdf.py
ocrmypdf page-image.pdf page-searchable.pdf
Playwright CLI
If you have the Playwright command-line tooling available, a basic screenshot can be captured from a shell. Consult the installed CLI’s help for exact options available in your version:
npx playwright screenshot --full-page https://example.com page.png
Then convert and OCR the saved PNG using the same steps above. CLI behavior and options can vary by Playwright version; use the API code when you need explicit waits, viewport setup, error handling, or repeatable application logic.
cURL
cURL can download an existing PDF, but it cannot render a webpage with Playwright or turn a screenshot into searchable text. The workflow requires browser automation plus image conversion and OCR. If an application already exposes a PDF URL, a direct download looks like this:
curl -L 'https://example.com/document.pdf' -o document.pdf
This downloads the server-provided document; it does not capture the webpage or run OCR.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. It returns a PNG, JPEG, WebP, or PDF with one GET request. For a screenshot-faithful PDF that needs searchable text, capture the page with ScreenshotNeo and then run OCRmyPDF on the resulting image PDF workflow as appropriate; a screenshot image alone does not become searchable automatically. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo accepts cookie and consent banners before capture and removes 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing, with the page verdict and billing status reported in response headers. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up free for 1,000 screenshots a month, with no card required.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Screenshot is blank or missing content | The capture started before the page rendered, navigation failed, or important content loads later. | Check the page URL and browser output; wait for a relevant selector or suitable page state; inspect the PNG before converting it. |
networkidle never arrives |
The page keeps network requests open, such as polling or streaming. | Use a different navigation wait condition and explicitly wait for the content you need. |
| Images or sections are absent | They may be lazy-loaded or dependent on interaction. | Inspect the page and scroll or interact as needed before capturing; validate the screenshot rather than assuming all content loaded. |
ocrmypdf command not found |
OCRmyPDF is not installed or its executable is not on the shell path. | Install it using the documentation for your platform and activate the environment where it is installed. |
| PDF has no searchable text | The OCR step was skipped, failed, or the wrong output file was checked. | Run OCRmyPDF on the image PDF, inspect its exit status, and search the generated output file. |
| Search finds incorrect or garbled text | Small text, low resolution, poor contrast, unusual fonts, or complex layouts reduce recognition quality. | Capture at a readable viewport or higher device scale, use a clear source image, and verify important text manually. |
| Native PDF looks different from screenshot | page.pdf() uses print CSS by default and may follow print-specific layout rules. |
Use screenshot-to-image-PDF when pixel appearance is the priority, or call page.emulateMedia({ media: 'screen' }) before page.pdf(). |
| Full-page image is huge or PDF text is tiny | A long screenshot was squeezed onto one fixed-size page. | Use an image-sized page, split the capture into sections, or choose native PDF pagination if it suits the requirement. |
Performance, reliability, and cost
A full-page screenshot consumes more memory and takes more processing than a viewport image, especially for long pages. OCR adds a separate processing step, and large images can take longer and produce larger files. Capture only the area needed, use an appropriate viewport, and avoid unnecessary repeated captures. Keep timeouts bounded and treat navigation, conversion, and OCR failures as separate steps so you can identify where a run stopped.
OCR is a recognition process, not a guarantee of exact text. Small text, complex layouts, and low contrast can affect the searchable layer. For records where text accuracy matters, preserve the original screenshot and review extracted content. Native page.pdf() avoids the OCR stage when its browser-rendered output is acceptable, but its print CSS behavior may change the appearance.
Playwright and OCRmyPDF are software tools; this workflow has no per-capture service price specified here. Account for the compute, storage, and processing time in the environment where you run it. ScreenshotNeo offers 1,000 shots a month free with no card, then paid tiers of $5 for 3,000, $15 for 15,000, $39 for 60,000, $99 for 250,000, and $249 for 1,000,000; yearly billing gives two months free, and every feature is on every plan. OCR processing, if needed for searchable screenshot PDFs, remains a separate step.
FAQ
Does page.screenshot() create a PDF?
No. It creates an image. Convert that image to PDF, then run OCR to add searchable text.
Is page.pdf() a screenshot PDF?
No. It renders the page as a PDF using print CSS by default. Emulate screen media first if screen styles are required.
Can OCR recover text that is not visible in the screenshot?
No. OCR reads the pixels in the captured image; content absent from the image cannot be recovered from that screenshot.
Should I keep the image PDF?
Keeping the pre-OCR image PDF or original PNG gives you a visual source to compare if the OCR layer needs review or regeneration.


