ScreenshotNeo

BlogHow-to

How to Screenshot a Website Behind a Login Form with Playwright in Python

Log in with Playwright, verify the authenticated page, save and reuse session state, and capture a viewport, full page, or element in Python.

By the ScreenshotNeo team4 October 20269 min read

To screenshot a page that requires a login, use Playwright to submit the site’s login form, wait for a reliable signed-in indicator, navigate to the protected page, and call page.screenshot(). For repeat runs, save the authenticated browser context with storage_state and load it in a later run. Treat that state file as a credential: anyone who can use it may be able to access the account. [Playwright authentication guide] [Playwright screenshots guide]

This walkthrough is for accounts and pages you are authorized to access. The example uses placeholder URLs, labels, and signed-in text: replace them with the site’s actual accessible form labels and a stable post-login signal.

1. Install Playwright and its browser

Create a project environment, install the Python package, then install the browser binary. Playwright’s Python documentation supports both synchronous and asynchronous APIs; the examples here use the synchronous API consistently. [Installation]

python -m venv .venv
source .venv/bin/activate
python -m pip install playwright
playwright install chromium

On Windows PowerShell, activate the environment with .venv\Scripts\Activate.ps1. Install the browser in the same environment that runs the script. If your project already uses Playwright, keep its pinned version and install the matching browser rather than changing versions just for this example.

2. Log in, confirm the session, and capture a page

Save the following as capture_after_login.py. Set LOGIN_URL and TARGET_URL to pages on the site you are authorized to access. Provide credentials through environment variables rather than putting real secrets in source code.

import os
from pathlib import Path
from playwright.sync_api import sync_playwright

LOGIN_URL = "https://example.com/login"
TARGET_URL = "https://example.com/account"
AUTH_FILE = Path("playwright/.auth/state.json")

username = os.environ["SITE_USERNAME"]
password = os.environ["SITE_PASSWORD"]

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    context = browser.new_context()
    page = context.new_page()

    page.goto(LOGIN_URL, wait_until="domcontentloaded")
    page.get_by_label("Username or email").fill(username)
    page.get_by_label("Password").fill(password)
    page.get_by_role("button", name="Sign in").click()

    # Replace with a stable signal that only appears after successful login.
    page.get_by_text("Account", exact=True).wait_for(state="visible")

    AUTH_FILE.parent.mkdir(parents=True, exist_ok=True)
    context.storage_state(path=str(AUTH_FILE))

    page.goto(TARGET_URL, wait_until="domcontentloaded")
    page.get_by_text("Account", exact=True).wait_for(state="visible")
    page.screenshot(path="account.png", full_page=True)

    context.close()
    browser.close()

Run it by setting the environment variables in your shell. For example, on macOS or Linux:

export SITE_USERNAME='your-test-user'
export SITE_PASSWORD='your-secret'
python capture_after_login.py

The form interaction uses accessible labels and a button role. Playwright recommends user-facing locators such as these; locator actions wait for elements to become actionable and retry assertions, which is generally more robust than selecting elements by fragile layout details. [Locators]

Why wait for a signed-in signal?

A successful click only means the button was clicked. It does not prove authentication succeeded. Wait for a site-specific element that appears only when signed in, such as an account menu or dashboard heading. If the site redirects after login, you can also wait for a known URL, but a visible authenticated-page element is a useful additional check.

Do not use a generic signal that also appears on the login page. A weak check can save an unauthenticated session and produce a screenshot of the login form while making the script look successful.

3. Reuse the saved login session

After one successful login, load the saved state when creating a new browser context. The state can include cookies and browser storage used by the application. It may expire, be revoked, or be bound to additional signals, so verify the authenticated state on every run. [Authentication] [Browser]

from pathlib import Path
from playwright.sync_api import sync_playwright

AUTH_FILE = Path("playwright/.auth/state.json")
TARGET_URL = "https://example.com/account"

if not AUTH_FILE.exists():
    raise SystemExit("No saved state. Run the login flow first.")

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    context = browser.new_context(storage_state=str(AUTH_FILE))
    page = context.new_page()
    page.goto(TARGET_URL, wait_until="domcontentloaded")

    # Replace with a reliable signed-in indicator for this application.
    page.get_by_text("Account", exact=True).wait_for(state="visible")
    page.screenshot(path="account.png", full_page=True)

    context.close()
    browser.close()

Generate the file only after a confirmed login. Store it outside public build artifacts, restrict access to it, and do not print its contents or commit it. Playwright explicitly warns that saved authentication state can contain cookies and headers that can impersonate the account, and recommends keeping authentication files out of version control. [Authentication]

# .gitignore
playwright/.auth/

Some authentication designs need extra investigation. Playwright documents storage state for cookies, local storage, and optionally IndexedDB; session storage is not persisted by the standard storage-state mechanism. Passkeys or app-specific session handling can also require a different permitted flow. Check the application’s design and your installed Playwright version before assuming the saved file contains everything needed. [BrowserContext] [Credentials]

4. Choose what to capture

Goal Example What it captures
Current viewport page.screenshot(path="view.png") The visible browser viewport.
Entire scrollable page page.screenshot(path="full.png", full_page=True) A full-page image extending beyond the viewport.
One element page.get_by_role("main").screenshot(path="main.png") The selected element’s visible bounds. A scrollable element’s screenshot includes only its visible portion.

These are different capture modes. Use a viewport image for a particular visible state, a full-page screenshot for a long document, and a locator screenshot for a component or region. Playwright’s screenshot API can also return image bytes if you omit the path and consume the result in memory. [Screenshots] [Locator]

Capture a specific element

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch()
    context = browser.new_context(storage_state="playwright/.auth/state.json")
    page = context.new_page()
    page.goto("https://example.com/account", wait_until="domcontentloaded")
    page.get_by_text("Account", exact=True).wait_for(state="visible")
    page.get_by_role("main").screenshot(path="main.png")
    browser.close()

For controlled element screenshots, Playwright’s locator screenshot API supports masking matching regions and disabling animations. Choose mask selectors carefully: masking covers their bounding boxes, so a broad selector can obscure more than intended. [Locator screenshot options]

Wait for content before capture

Modern pages can render the shell first and fetch protected content later. Wait for a meaningful page element before capturing instead of relying on a fixed sleep:

page.goto(TARGET_URL, wait_until="domcontentloaded")
page.get_by_role("heading", name="Account overview").wait_for(state="visible")
page.screenshot(path="account.png", full_page=True)

Replace the heading with an element that reliably indicates the content you need. A fixed delay may help with a known delayed animation, but it can be unnecessarily slow or still too short when the site is under load.

5. Authentication edge cases

  • MFA: A username and password example does not handle every multi-factor flow. Use the site’s permitted authentication path and wait for the completed signed-in state.
  • CAPTCHA or bot checks: Do not assume a form script can complete these checks. Follow the site’s access rules and an authorized workflow.
  • Federated identity: The sign-in button may redirect to an identity provider or open a popup. The example’s selectors and flow must be adapted to that site’s documented login process.
  • Session expiry: A previously saved state may stop working. Detect the signed-in indicator, then run the login flow again when necessary.
  • Multiple origins: A service may use more than one origin during authentication. Confirm the saved context contains the state needed by the protected page.
  • IndexedDB tokens: If the application stores authentication there, consult the installed Playwright version and storage-state API for the IndexedDB option.
  • Session storage: Standard storage-state reuse does not automatically preserve session storage. Handle it only if the application requires it, using a permitted, site-specific method.
  • API-based login: Playwright documents a supported pattern for using an API request context to establish state where the application offers a legitimate API authentication route. Do not use API calls to bypass the site’s access controls. [API testing]

6. Troubleshooting

Symptom Likely cause Fix
get_by_label() finds no field The page’s accessible label differs from the placeholder, or the field is inside a frame. Inspect the page’s accessible name and use the actual label. For a framed form, locate the relevant frame before finding controls.
Timeout waiting for the submit button or signed-in text The locator does not match, login failed, or the site needs another step. Check the actual accessible role/name and the visible error state. Wait for a real post-login signal and handle site-specific MFA or redirects.
Screenshot still shows the login page The click did not authenticate, the target redirects to login, or the state was saved before login completed. Wait for the authenticated indicator before saving state and again after navigating to the target. Confirm the target URL is protected by the same expected account.
Saved state works once, then fails The session expired, was revoked, or depends on storage not included in the state file. Reauthenticate through the permitted flow. Check whether the app depends on session storage, IndexedDB, or another method.
Screenshot is blank or missing page content The page shell loaded before the protected data, or the target element is not visible yet. Wait for a specific content element. Check the page URL and signed-in state before capture.
Full-page capture omits content in a scrollable panel A nested scroll container is not the same as the document’s full scrollable page. Capture the relevant locator or scroll the panel as required by the page, then capture the desired visible content.
Browser launch reports a missing executable The Playwright package is installed but its browser binary is not. Run playwright install chromium in the environment used for the script.
Authentication works locally but not in CI Different environment, expired state, missing browser dependencies, or an environment-specific sign-in step. Install the matching browser, create or securely provide fresh state in the CI environment, and verify the signed-in indicator before capture.

7. Reliability, performance, and cost

Reusing storage state avoids repeating the login form on every run, but it introduces a secret that needs lifecycle management. Keep state scoped to the right account and environment, refresh it when sessions expire, and avoid sharing it between unrelated jobs. A saved state is not a guarantee of permanent authentication.

For reliability, use semantic locators, wait for a stable authenticated-page signal, and wait for the specific content to appear. Keep login and capture as distinct steps so failures are easy to identify. Avoid adding arbitrary long sleeps as a substitute for verifying page state.

Capture cost is mainly the browser work and image size your own runtime must handle; a full-page image can be much larger than a viewport image. Prefer the smallest capture mode that answers the task, and close contexts and browsers when finished. This workflow uses your Playwright runtime; it has no ScreenshotNeo API charge.

8. Or skip the browser setup

If the target page is publicly accessible without authentication, ScreenshotNeo can return a website screenshot through one API request. It cannot sign into an account or capture private content that requires your login. For public pages, its API accepts a URL and returns an image or PDF. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, and failed loads are never billed, and response headers say which outcome occurred. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Every feature is on every plan. Sign up for 1,000 free screenshots a month, with no card.

9. FAQ

Can I save the screenshot as JPEG or WebP?

Playwright’s screenshot API supports image format options documented for the installed version. Check the current API documentation and release notes before relying on version-specific formats such as WebP. [Screenshots] [Release notes]

Can I take a screenshot without saving a file first?

Yes. The screenshot call can return image bytes, which you can pass to another library or upload step instead of writing to a path. [Screenshots]

Can ScreenshotNeo capture the private page after login?

This article’s ScreenshotNeo example is for publicly accessible pages. Use Playwright for pages that depend on your authenticated browser session.

References