ScreenshotNeo

BlogHTML to image & PDF

How to capture a website as a PDF with BrowserCat

Connect Playwright to BrowserCat’s cloud Chromium, prepare the page, and save a PDF with the print settings you need.

By the ScreenshotNeo team4 October 20267 min read

To save a website as a PDF with BrowserCat, connect Playwright to BrowserCat’s cloud Chromium session, navigate to the page, wait for the content you need, then call Playwright’s page.pdf(). BrowserCat provides the remote browser connection; Playwright provides the PDF-generation method and its print settings. The reviewed BrowserCat quick start demonstrates screenshots rather than a dedicated PDF endpoint, so this is Playwright PDF generation running in BrowserCat’s browser session.

1. Set up Playwright and connect to BrowserCat

BrowserCat’s quick start uses the WebSocket endpoint wss://api.browsercat.com/connect and an Api-Key connection header. The session runs on Chromium in the cloud. Install Playwright in your project and use a Playwright version whose chromium.connect() API accepts the connection options shown below. Keep your API key in an environment variable rather than in source code.

npm install playwright

Set the key in your shell before running the script:

export BROWSERCAT_API_KEY="your_browsercat_api_key"

On Windows PowerShell, use $env:BROWSERCAT_API_KEY="your_browsercat_api_key".

2. Generate a PDF with Playwright

Save this as capture-pdf.js. Replace the example URL with a page you are authorized to access. The readiness selector is an example: choose a selector that indicates the actual content you need is present, or remove that wait if it does not apply.

const { chromium } = require('playwright');

async function main() {
  const targetUrl = 'https://example.com';
  const browser = await chromium.connect('wss://api.browsercat.com/connect', {
    headers: { 'Api-Key': process.env.BROWSERCAT_API_KEY },
  });

  try {
    const page = await browser.newPage();
    await page.goto(targetUrl, { waitUntil: 'domcontentloaded' });

    // Replace with a page-specific readiness condition when needed.
    // await page.locator('main').waitFor({ state: 'visible', timeout: 15000 });

    await page.pdf({
      path: 'page.pdf',
      format: 'A4',
      printBackground: true,
      margin: { top: '12mm', right: '12mm', bottom: '12mm', left: '12mm' },
    });
  } finally {
    await browser.close();
  }
}

if (!process.env.BROWSERCAT_API_KEY) {
  throw new Error('Set BROWSERCAT_API_KEY before running this script.');
}

main().catch((error) => {
  console.error(error);
  process.exitCode = 1;
});

Run it with node capture-pdf.js. The example’s URL, selector, and output filename are placeholders. BrowserCat’s documented quick-start flow uses a remote connection in place of a local browser launch, followed by navigation, page work, and closing the browser connection.

3. Choose the PDF rendering behavior

Playwright’s page.pdf() renders with print CSS media by default. A page may therefore look different in the PDF than it does in a normal browser tab. If you want screen media styles instead, call page.emulateMedia({ media: 'screen' }) before page.pdf(). This changes the media mode used for rendering; it does not guarantee that every website’s screen layout will paginate as intended.

await page.emulateMedia({ media: 'screen' });
await page.pdf({ path: 'page.pdf', format: 'A4', printBackground: true });

Set print options according to the document you need. Playwright’s documented options include:

Option What it controls Documented behavior
format Named paper size Letter is the default. Named sizes include A4 and Legal. If set, format takes priority over explicit width and height.
width, height Paper dimensions Use dimensions when a named format is not appropriate. Consider preferCSSPageSize if the page declares its own print size.
margin Space around printed content Defaults to zero. Set top, right, bottom, and left margins with supported length values.
landscape Page orientation Portrait is the default; set true for landscape.
printBackground Background graphics and colors Defaults to false. Enable it if the page’s background graphics are part of the intended document.
scale Content size Defaults to 1 and accepts values from 0.1 to 2.
pageRanges Pages to include Select page ranges when needed; an empty range prints all pages.
preferCSSPageSize Whether CSS page size takes precedence When true, the page’s @page size takes priority. Otherwise, content is scaled to fit the selected paper size.
displayHeaderFooter, templates Printed header and footer Templates can include date, title, URL, page number, and total page count using Playwright’s documented classes.
tagged Accessible tagged PDF output Available in the documented API and false by default.

For example, print selected pages in landscape on Letter paper with backgrounds:

await page.pdf({
  path: 'selected-pages.pdf',
  format: 'Letter',
  landscape: true,
  printBackground: true,
  pageRanges: '1-3',
  margin: { top: '0.5in', right: '0.5in', bottom: '0.5in', left: '0.5in' },
});

For exact print colors, Playwright’s API reference points to the CSS property -webkit-print-color-adjust. Printed colors may be modified by default. Header and footer templates have limitations: scripts in them are not evaluated, and page styles are not visible inside the templates.

4. Wait for the page state you actually need

A successful navigation does not necessarily mean the page’s final content is ready. Client-rendered sections, images, charts, and interactive elements may load after the initial document. Prefer a page-specific condition, such as waiting for the report heading to appear or completing a required interaction, over assuming one navigation event covers every site.

await page.goto(targetUrl, { waitUntil: 'domcontentloaded' });
await page.locator('[data-report-ready="true"]').waitFor({ state: 'visible', timeout: 15000 });
await page.pdf({ path: 'report.pdf', format: 'A4', printBackground: true });

The selector above is illustrative; use one that the target page actually exposes. If a site needs a click, form submission, or authenticated state, perform that work before generating the PDF. The BrowserCat quick start demonstrates navigation, interaction, and waiting for a load state in its capture workflow.

5. Other BrowserCat connection configuration

BrowserCat’s configuration overview describes options passed through a BrowserCat-Opts JSON header or a smaller set of query parameters. When the same setting appears in both places, header values take precedence. The reviewed configuration documentation identifies Chrome and Chromium as available browser choices; Firefox and WebKit are marked as roadmap items. Preferred region routing is also described as a roadmap item, so do not rely on explicit region selection as a currently documented capability.

For this PDF workflow, start with the documented connection endpoint and API-key header. Add only connection settings that your BrowserCat account and current documentation support. The source material does not establish a dedicated BrowserCat PDF option: paper geometry and print rendering are controlled with Playwright’s page.pdf().

6. Do-it-yourself alternatives and tradeoffs

With local Playwright, the browser runs in your own environment; with BrowserCat, the connection uses its cloud Chromium session. BrowserCat recommends Playwright as a starting route and also lists Puppeteer and other CDP clients, but this guide covers Playwright’s documented page.pdf() method. Choose based on where you want the browser session to run and the integration you already maintain. The available sources do not establish a performance, output-quality, or cost advantage for either route.

7. Troubleshooting

Symptom Likely cause What to check
Connection fails before a page opens Missing or invalid API key, incorrect header, or endpoint mismatch Confirm BROWSERCAT_API_KEY is set and passed as Api-Key, and use wss://api.browsercat.com/connect.
PDF call errors or is unavailable The connected page or installed Playwright API does not match the documented method/version Check that the page is a Playwright Page and consult the Page API for the installed version’s page.pdf() options.
PDF is blank or missing a section Capture began before client-rendered content was ready, or a required interaction was skipped Wait for a page-specific selector or complete the interaction before calling page.pdf().
PDF looks different from the browser page.pdf() uses print media by default, or the site has print-specific CSS Inspect print styles; use emulateMedia({ media: 'screen' }) if screen media is the intended rendering target.
Colors or background artwork are absent Background printing is disabled, or print color adjustment changes colors Set printBackground: true and use -webkit-print-color-adjust in page CSS when exact print colors are needed.
Content is clipped or unexpectedly scaled Paper size, CSS @page rules, margins, or scale do not match the content Review format versus width/height, margins, preferCSSPageSize, and the scale range of 0.1 to 2.
Header/footer values are missing Template limitations or unsupported markup Use the documented date, title, URL, page number, and total-pages classes; do not rely on script execution or page styles inside templates.

8. Performance, reliability, and cost considerations

PDF generation includes page navigation, readiness work, and print rendering. A page-specific wait helps avoid both capturing too early and waiting on unrelated activity. Large or dynamic pages may need more preparation; the reviewed sources do not provide a benchmark or a guarantee that every site will render successfully. Always close the remote browser connection in a finally block so it is closed after either success or an error.

The supplied BrowserCat research does not establish service prices, usage limits, or comparative cost figures. Check BrowserCat’s current pricing and service documentation before estimating the cost of a production workflow. For reproducibility, record the target URL, chosen media mode, paper settings, and page-specific readiness condition alongside your job configuration.

Or skip the browser setup

If your goal is a PDF, BrowserCat and Playwright provide the cloud-browser workflow above. If you need a website screenshot or PDF through a screenshot API, ScreenshotNeo offers a one-request capture API and an MCP server. Its PDF options include paper size, margins, landscape orientation, and page ranges. The request below saves a PDF response; see the ScreenshotNeo API documentation for the API details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -d format=pdf -o page.pdf

ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Every feature is on every plan. Get started with the free ScreenshotNeo account.

FAQ

Does BrowserCat provide the PDF method in this guide?

The documented PDF call here is Playwright’s page.pdf(), running through BrowserCat’s cloud Chromium session. The reviewed BrowserCat quick start shows screenshot capture, not a dedicated PDF endpoint.

Will the PDF look exactly like the page on screen?

Not necessarily. PDF rendering uses print CSS by default. You can emulate screen media before generating the PDF, but site-specific layout and print rules still affect the result.

Can I create a PDF from only part of a page?

The settings covered here select PDF pages, not a webpage element. Prepare the page and its print styles to control what appears in the document.

Can I use Firefox or route to a preferred BrowserCat region?

The reviewed BrowserCat configuration overview marks Firefox, WebKit, and preferred region routing as roadmap items. Confirm current documentation before depending on them.