ScreenshotNeo

BlogHow-to

How to save a webpage screenshot as a searchable PDF with Playwright

Create a searchable PDF from a Playwright webpage screenshot with OCR, or use Playwright’s native PDF export when print rendering is acceptable.

By the ScreenshotNeo team4 October 20269 min read

To save a webpage screenshot as a searchable PDF with Playwright, capture the page as an image, convert that image to a PDF, then add a text layer with OCR. Playwright’s page.screenshot() captures pixels; it does not make the resulting PDF searchable by itself. If you do not need screenshot-faithful pixels, use Playwright’s page.pdf() instead: it creates a PDF from the page and uses print CSS by default. Playwright Page API · OCRmyPDF documentation.

Choose the right PDF workflow

Workflow What it preserves Searchable text Use it when
Screenshot → image PDF → OCR The captured pixels and screen appearance Added by OCR; accuracy depends on the source image and OCR The screenshot’s visual appearance matters most
Playwright page.pdf() Browser-rendered PDF output Text is generally represented by the page’s PDF rendering A print-style document is acceptable

These are different outputs. Playwright documents screenshot capture and PDF rendering as separate features. OCRmyPDF adds a text layer to image PDFs so text can be selected, searched, and copied. A screenshot is not automatically OCR-processed.

Prerequisites

  • Node.js and a Playwright project with its browser installed.
  • Python 3 and Pillow for the image-to-PDF conversion in this example.
  • OCRmyPDF installed and available as the ocrmypdf command. Follow its documentation for installation requirements on your operating system.

Install the JavaScript dependencies:

npm init -y
npm install playwright
npx playwright install chromium

Install Pillow in a Python environment:

python -m pip install Pillow

Install OCRmyPDF using the instructions for your platform in the OCRmyPDF documentation. Its dependencies and install method vary by operating system.

Capture a full-page screenshot and make it searchable

  1. Navigate to the page and wait for a page state suitable for capture.
  2. Capture the viewport or the full scrollable page as a PNG.
  3. Convert the PNG to a one-page image PDF.
  4. Run OCRmyPDF to add a searchable text layer.
  5. Inspect the PDF visually and search for text that is clearly visible in the source.

1. Capture with Playwright

Save as capture.mjs. Replace the example URL with the page you are authorized to capture.

import { chromium } from 'playwright';

const url = 'https://example.com';
const browser = await chromium.launch();
try {
  const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });
  await page.goto(url, { waitUntil: 'networkidle', timeout: 60000 });
  await page.screenshot({ path: 'page.png', fullPage: true });
} finally {
  await browser.close();
}

Run it with:

node capture.mjs

fullPage: true captures the full scrollable page; omit it to capture only the current viewport. The official Playwright screenshots guide documents saving screenshots to a path and capturing buffers.

2. Convert the screenshot to an image PDF

Save as image_to_pdf.py. Pillow converts the PNG to an RGB PDF, which is suitable for the OCR step:

from PIL import Image

with Image.open('page.png') as image:
    image.convert('RGB').save('page-image.pdf', 'PDF', resolution=150.0)

Run the conversion and then OCR:

python image_to_pdf.py
ocrmypdf page-image.pdf page-searchable.pdf

The output, page-searchable.pdf, retains the screenshot as its page image and includes OCR text for searching and selection. OCR quality depends on legibility, contrast, resolution, and the content being recognized. Review the result before relying on extracted text.

One-command OCR alternative

OCRmyPDF can also create a PDF from image files as part of its image-processing workflow. Check the installed version’s command-line documentation for supported input formats and options. The explicit conversion above keeps the image-to-PDF step visible and makes the input to OCR clear.

When native Playwright PDF output is better

Use page.pdf() when the goal is a browser-rendered document rather than a PDF that reproduces the screenshot’s pixels. PDF output uses print CSS by default. To render screen styles, emulate screen media before exporting:

import { chromium } from 'playwright';

const browser = await chromium.launch();
try {
  const page = await browser.newPage();
  await page.goto('https://example.com', { waitUntil: 'networkidle', timeout: 60000 });
  await page.emulateMedia({ media: 'screen' });
  await page.pdf({ path: 'page.pdf', format: 'A4', printBackground: true });
} finally {
  await browser.close();
}

For the default print rendering, omit emulateMedia. The printBackground setting includes background graphics in the PDF. Choose this route if selectable page text and print layout matter more than identical screenshot pixels. Verify the rendered PDF for the target site; page-specific print styles can change layout or hide content.

Capture options and page-state choices

Choice Effect Practical guidance
Viewport screenshot Captures the current visible viewport Use when only the visible screen is needed.
fullPage: true Captures the full scrollable page Use for long articles; very tall pages create large images and may be slow to OCR.
Navigation wait condition Controls when capture begins networkidle can be unsuitable for pages with ongoing requests. Use a different wait condition or wait for a meaningful selector when appropriate.
Navigation timeout Bounds how long navigation waits Set a timeout suitable for the site and handle slow or unavailable pages.
PNG screenshot Lossless image input for OCR Usually a practical choice for text; make sure the source is legible at its capture size.
PDF page layout Determines how an image is placed on a PDF page The Pillow example creates a page from the image. For a fixed paper size, scale or tile the image deliberately; scaling a long page down can make text difficult to read.

For dynamic pages, waiting only for navigation may capture before important content appears. Wait for a content selector, or use a deliberate delay if the page has a known rendering delay. Lazy-loaded images may not appear until scrolled into view; full-page capture alone should not be treated as proof that every site has loaded all its content. Inspect the image before processing it.

cURL, Python, and Node.js alternatives

Playwright’s browser automation APIs are available in multiple languages. The Node.js examples above show the complete workflow. Here are equivalent capture examples for Python and the Playwright CLI. The OCR and conversion steps still follow afterward.

Python Playwright

Install the package and browser:

python -m pip install playwright
python -m playwright install chromium

Save as capture.py:

import asyncio
from playwright.async_api import async_playwright

async def main():
    async with async_playwright() as playwright:
        browser = await playwright.chromium.launch()
        try:
            page = await browser.new_page(viewport={"width": 1440, "height": 1000})
            await page.goto("https://example.com", wait_until="networkidle", timeout=60000)
            await page.screenshot(path="page.png", full_page=True)
        finally:
            await browser.close()

asyncio.run(main())
python capture.py
python image_to_pdf.py
ocrmypdf page-image.pdf page-searchable.pdf

Playwright CLI

If you have the Playwright command-line tooling available, a basic screenshot can be captured from a shell. Consult the installed CLI’s help for exact options available in your version:

npx playwright screenshot --full-page https://example.com page.png

Then convert and OCR the saved PNG using the same steps above. CLI behavior and options can vary by Playwright version; use the API code when you need explicit waits, viewport setup, error handling, or repeatable application logic.

cURL

cURL can download an existing PDF, but it cannot render a webpage with Playwright or turn a screenshot into searchable text. The workflow requires browser automation plus image conversion and OCR. If an application already exposes a PDF URL, a direct download looks like this:

curl -L 'https://example.com/document.pdf' -o document.pdf

This downloads the server-provided document; it does not capture the webpage or run OCR.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. It returns a PNG, JPEG, WebP, or PDF with one GET request. For a screenshot-faithful PDF that needs searchable text, capture the page with ScreenshotNeo and then run OCRmyPDF on the resulting image PDF workflow as appropriate; a screenshot image alone does not become searchable automatically. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo accepts cookie and consent banners before capture and removes 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing, with the page verdict and billing status reported in response headers. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up free for 1,000 screenshots a month, with no card required.

Troubleshooting

Symptom Likely cause Fix
Screenshot is blank or missing content The capture started before the page rendered, navigation failed, or important content loads later. Check the page URL and browser output; wait for a relevant selector or suitable page state; inspect the PNG before converting it.
networkidle never arrives The page keeps network requests open, such as polling or streaming. Use a different navigation wait condition and explicitly wait for the content you need.
Images or sections are absent They may be lazy-loaded or dependent on interaction. Inspect the page and scroll or interact as needed before capturing; validate the screenshot rather than assuming all content loaded.
ocrmypdf command not found OCRmyPDF is not installed or its executable is not on the shell path. Install it using the documentation for your platform and activate the environment where it is installed.
PDF has no searchable text The OCR step was skipped, failed, or the wrong output file was checked. Run OCRmyPDF on the image PDF, inspect its exit status, and search the generated output file.
Search finds incorrect or garbled text Small text, low resolution, poor contrast, unusual fonts, or complex layouts reduce recognition quality. Capture at a readable viewport or higher device scale, use a clear source image, and verify important text manually.
Native PDF looks different from screenshot page.pdf() uses print CSS by default and may follow print-specific layout rules. Use screenshot-to-image-PDF when pixel appearance is the priority, or call page.emulateMedia({ media: 'screen' }) before page.pdf().
Full-page image is huge or PDF text is tiny A long screenshot was squeezed onto one fixed-size page. Use an image-sized page, split the capture into sections, or choose native PDF pagination if it suits the requirement.

Performance, reliability, and cost

A full-page screenshot consumes more memory and takes more processing than a viewport image, especially for long pages. OCR adds a separate processing step, and large images can take longer and produce larger files. Capture only the area needed, use an appropriate viewport, and avoid unnecessary repeated captures. Keep timeouts bounded and treat navigation, conversion, and OCR failures as separate steps so you can identify where a run stopped.

OCR is a recognition process, not a guarantee of exact text. Small text, complex layouts, and low contrast can affect the searchable layer. For records where text accuracy matters, preserve the original screenshot and review extracted content. Native page.pdf() avoids the OCR stage when its browser-rendered output is acceptable, but its print CSS behavior may change the appearance.

Playwright and OCRmyPDF are software tools; this workflow has no per-capture service price specified here. Account for the compute, storage, and processing time in the environment where you run it. ScreenshotNeo offers 1,000 shots a month free with no card, then paid tiers of $5 for 3,000, $15 for 15,000, $39 for 60,000, $99 for 250,000, and $249 for 1,000,000; yearly billing gives two months free, and every feature is on every plan. OCR processing, if needed for searchable screenshot PDFs, remains a separate step.

FAQ

Does page.screenshot() create a PDF?

No. It creates an image. Convert that image to PDF, then run OCR to add searchable text.

Is page.pdf() a screenshot PDF?

No. It renders the page as a PDF using print CSS by default. Emulate screen media first if screen styles are required.

Can OCR recover text that is not visible in the screenshot?

No. OCR reads the pixels in the captured image; content absent from the image cannot be recovered from that screenshot.

Should I keep the image PDF?

Keeping the pre-OCR image PDF or original PNG gives you a visual source to compare if the OCR layer needs review or regeneration.