Convert a Webpage to PDF With Cookies Using Playwright
Use Playwright to load a webpage with the right cookies or saved login state, wait for its content, and create a PDF with the layout you need.
To convert a webpage that depends on cookies into a PDF with Playwright, add the required cookies to a browser context before navigating, wait until the page shows the content you need, then call page.pdf(). If authentication also depends on local storage or IndexedDB, restore the saved browser storage state instead of injecting cookies alone.
This guide uses Playwright’s JavaScript API. It covers direct cookie injection, saved authentication state, PDF layout options, common failures, and safe handling of session data. See the BrowserContext cookie API, the Page PDF API, and the authentication guide.
1. Install Playwright and its browser
Start in a new project directory. Playwright’s browser binaries are installed separately from the package:
npm init -y
npm install playwright
npx playwright install chromium
Chromium is used here because the example generates PDFs with page.pdf(). Check the API documentation for the Playwright version in your project if you need to confirm option availability.
2. Add cookies, open the page, and write the PDF
Create a context, add cookies with the scope the target site expects, and only then navigate. This complete Node.js script reads the cookie value from an environment variable so it is not written into the source file:
// save as capture-pdf.js
const { chromium } = require('playwright');
async function main() {
const cookieValue = process.env.SESSION_COOKIE;
if (!cookieValue) throw new Error('Set SESSION_COOKIE before running this script');
const browser = await chromium.launch();
let context;
try {
context = await browser.newContext();
await context.addCookies([
{
name: 'session',
value: cookieValue,
url: 'https://example.com',
httpOnly: true,
secure: true,
sameSite: 'Lax',
},
]);
const page = await context.newPage();
const response = await page.goto('https://example.com/account', {
waitUntil: 'domcontentloaded',
timeout: 30_000,
});
if (response && !response.ok()) {
throw new Error(`Navigation failed with HTTP ${response.status()}`);
}
// Replace this with a locator that proves the required content is ready.
await page.getByRole('heading', { name: 'Account' }).waitFor({ timeout: 15_000 });
await page.pdf({
path: 'page.pdf',
format: 'A4',
printBackground: true,
});
console.log('Wrote page.pdf');
} finally {
if (context) await context.close();
await browser.close();
}
}
main().catch((error) => {
console.error(error);
process.exitCode = 1;
});
Run it with a cookie value obtained through an authorized login or session flow:
SESSION_COOKIE='your-session-cookie-value' node capture-pdf.js
The name, value, host and attributes are placeholders. Use the actual cookie properties the site requires. A cookie can be scoped using url, or using both domain and path. For example, a host-only cookie might use url: 'https://app.example.com'; a domain cookie may use domain: '.example.com' and path: '/'. The cookie must be valid for the target URL and compatible with its security requirements.
3. Choose the right authentication state
Use direct cookies when cookies are sufficient
context.addCookies([...]) is useful when your authorized workflow already provides the cookie values and the site authenticates from cookies alone. Add them before opening the target page so the first navigation can send them.
Cookie attributes to check include:
nameandvalue: must match the cookie issued by the site.urlordomainpluspath: controls where the browser sends the cookie. Use one valid scope form.expires: expiration time, when the site uses a persistent cookie. An expired value will not authenticate.httpOnly: marks cookies unavailable to page JavaScript, as with many session cookies.secure: restricts sending to secure connections; use HTTPS for secure cookies.sameSite: use the site’s actual setting, represented by the API’s supported values.
Do not guess a production site’s cookie attributes. Use the browser context cookie API and the site’s authorized session setup as the source of truth.
Use saved storage state when login involves more than cookies
Some applications also rely on local storage or IndexedDB. Complete a permitted login flow once, save the context’s state, then create later contexts from that state:
// After completing the site's authorized login flow:
await page.context().storageState({ path: 'playwright/.auth/state.json' });
// In a later capture run:
const context = await browser.newContext({
storageState: 'playwright/.auth/state.json',
});
The auth guide explains that storage state can include cookies and local storage; IndexedDB can be included when requested by the API version you use. Standard storage state does not persist sessionStorage. If the application depends on sessionStorage, follow the guide’s custom initialization approach. Treat the state file as a credential: it may contain values that can impersonate the account. Keep it out of source control, restrict access, and rotate the session if it is exposed.
4. Wait for the page’s actual content
A successful navigation is not proof that a client-rendered page has finished loading its data. The example waits for an account heading. Choose a locator that only appears when the specific content needed in the PDF is ready, such as a report title, table, or account name:
await page.getByRole('heading', { name: 'Monthly report' }).waitFor();
// Then produce the PDF.
You can also wait for an appropriate selector or a deliberate delay where the application offers no reliable signal. Prefer a page-specific condition: an arbitrary sleep can be either too short to capture the content or unnecessarily long. Navigation options such as domcontentloaded describe document lifecycle events, not completion of every asynchronous application request.
5. Tune the PDF output
page.pdf() renders using print CSS media by default. The following options cover the common layout choices documented by Playwright:
| Option | Purpose | Example |
|---|---|---|
path |
Write the PDF to a file. Omit it to receive a PDF buffer. | path: 'page.pdf' |
format |
Choose a paper size. | format: 'A4' |
width, height |
Set explicit page dimensions when a named paper size is not appropriate. | width: '8.5in', height: '11in' |
margin |
Set page margins. | margin: { top: '12mm', right: '10mm', bottom: '12mm', left: '10mm' } |
printBackground |
Include background colors and images; it defaults to false. | printBackground: true |
preferCSSPageSize |
Let the document’s CSS @page size take priority over supplied paper dimensions. |
preferCSSPageSize: true |
pageRanges |
Limit output to selected PDF pages. | pageRanges: '1-3' |
scale |
Adjust the rendered page scale. | scale: 0.9 |
For example, combine a paper size, margins, background printing and CSS page sizing:
await page.pdf({
path: 'report.pdf',
format: 'A4',
margin: { top: '12mm', right: '10mm', bottom: '12mm', left: '10mm' },
printBackground: true,
preferCSSPageSize: true,
pageRanges: '1-5',
scale: 0.95,
});
When you need the screen layout instead of print layout, emulate screen media before generating the PDF:
await page.emulateMedia({ media: 'screen' });
await page.pdf({ path: 'screen-layout.pdf', printBackground: true });
For color fidelity in printed output, Playwright documents the CSS property -webkit-print-color-adjust. The page’s print styles still determine how content flows and which elements are hidden, so inspect the source site’s @media print and @page rules when output differs from expectations.
6. Troubleshoot missing authentication or PDF content
| Symptom | Likely cause | Fix |
|---|---|---|
| The PDF shows a login page | Cookie scope, expiry, or security attributes are wrong; or authentication needs local storage/IndexedDB too. | Verify the cookie’s domain, path, expiry, secure flag and value. If the app uses broader state, save and restore authenticated storage state. |
| The cookie seems present but is not sent | The cookie’s URL/domain does not match the destination, or the request is not using HTTPS for a secure cookie. | Set the correct URL or domain/path scope and navigate to the matching HTTPS origin. |
| Content is missing or partly rendered | The capture starts after navigation but before the application has finished rendering its data. | Wait for a meaningful content locator or another page-specific readiness signal before calling page.pdf(). |
| PDF layout differs from the browser | PDF generation uses print media by default, and print styles may change layout. | Use page.emulateMedia({ media: 'screen' }) for screen styles, or adjust the site’s print CSS and paper settings. |
| Backgrounds or colors are absent | Background printing is disabled by default, or print CSS adjusts colors. | Set printBackground: true; review -webkit-print-color-adjust for color behavior. |
| The saved-state run is signed out | The site depends on state that was not saved, such as sessionStorage, or the stored session expired. | Check which browser storage the app uses, initialize unsupported state as documented, or refresh the authorized login state. |
| Navigation times out | The page did not reach the selected lifecycle event within the timeout. | Confirm the URL and connectivity, select a suitable navigation condition, and wait separately for the content signal the capture requires. |
7. Reliability, performance, and cost
- Close resources: use
try/finallyso contexts and browsers close even when navigation or PDF writing fails. Reusing a browser process for several captures can avoid repeated launch overhead; keep each job’s context isolated when session data differs. - Wait precisely: a specific readiness locator makes output more consistent and avoids paying the time cost of a large fixed sleep on every run.
- Bound waits: set timeouts for navigation and content waits so a broken page does not hold a worker indefinitely. On failure, record the URL, error stage and HTTP status without logging cookie values.
- Control PDF size: page ranges can exclude unwanted pages. Paper size, scale and margins affect pagination and output dimensions. Enable backgrounds only when the design needs them.
- Budget for the full browser workflow: self-hosted Playwright work consumes the compute and storage of the environment running Chromium; account for browser startup, page rendering, PDF bytes and retention. The dossier provides no fixed runtime or cost benchmark, so measure with the pages and infrastructure you will actually use.
8. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. It returns a screenshot or PDF from one GET request. The browser-context workflow above remains the right fit when the page must use your own authenticated session cookies; do not send private session cookies to a third party unless that is explicitly appropriate for your use case.
For a public page, request a PDF directly. See the ScreenshotNeo API docs for the PDF parameters and response behavior:
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-d format=pdf \
-o page.pdf
Cookie banners are accepted and removed before capture, along with known newsletter popups and chat widgets; each step can be turned off. Bot checks, blank pages and failed loads are never billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan.
For an image response, the same API call works from Python or Node.js:
# Python image example
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
// Node.js image example
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned ${res.status}`);
await require('node:fs/promises').writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
For PDF output in those clients, request PDF using the documented API parameter and save the returned bytes with a .pdf filename. ScreenshotNeo is suited to public URL capture; the supplied product facts do not describe passing authenticated browser cookies. Create a free account for 1,000 screenshots a month, with no card required.
9. Frequently asked questions
Can Playwright create the PDF in memory?
Yes. Omit path from page.pdf() and use the returned PDF buffer in your application.
Will adding a cookie also restore local storage?
No. Cookie injection adds cookies to the context. Use saved storage state for supported broader authentication state, and handle sessionStorage separately if the application relies on it.
Does the saved authentication file expire?
The file itself remains until changed or removed, but the cookies or tokens it contains may expire or be revoked. Refresh it through the site’s authorized login flow when needed.
Can I save only selected pages?
Yes. Use pageRanges in the PDF options, such as '2-4', and confirm pagination matches the document version being captured.


