ScreenshotNeo

BlogHow-to

How to Log In to a Webpage with Pyppeteer

Use Pyppeteer to fill a real login form, handle navigation safely, verify authentication, and troubleshoot redirects, MFA, cookies, and SPA flows.

By the ScreenshotNeo team1 October 20269 min read

Direct answer: open the login page, wait for the actual form controls, fill the username and password fields, submit the form, and verify authentication with a site-specific signal such as the final URL or an authenticated-only element. If the submit click causes navigation, await the click and navigation together with asyncio.gather; waiting for them separately can race.

Pyppeteer is an unofficial Python port of Puppeteer. Its project repository currently says it is unmaintained and recommends considering Playwright Python for new work. The procedure below remains useful for maintaining an existing Pyppeteer script, but confirm behavior against the version installed in your environment and the site you are authorized to access.

1. Install Pyppeteer and prepare Chromium

Create an isolated environment and install the package:

python -m venv .venv
source .venv/bin/activate       # Windows: .venv\\Scripts\\activate
python -m pip install --upgrade pip
python -m pip install pyppeteer

On first use, Pyppeteer downloads Chromium when it cannot find a suitable executable. In CI or a locked-down server, provide a browser executable explicitly with executablePath or install Chromium through your system image. Keep the browser version and the Pyppeteer package under review because the repository is no longer actively maintained.

2. A complete form-login script

This runnable template uses placeholder selectors. Inspect the authorized target page and replace every URL, selector, and success condition with values from that site.

import asyncio
import os
from pyppeteer import launch

LOGIN_URL = "https://example.com/login"
USERNAME = os.environ["SITE_USERNAME"]
PASSWORD = os.environ["SITE_PASSWORD"]

async def main():
    browser = await launch(
        headless=True,
        args=["--no-sandbox"],  # Use only when required by your container runtime.
    )
    page = await browser.newPage()
    try:
        await page.setViewport({"width": 1440, "height": 900})
        response = await page.goto(
            LOGIN_URL,
            {"waitUntil": "domcontentloaded", "timeout": 60000},
        )
        if response is None:
            raise RuntimeError("The login navigation returned no response")

        await page.waitForSelector('input[name="username"]', {"visible": True})
        await page.waitForSelector('input[name="password"]', {"visible": True})

        await page.click('input[name="username"]')
        await page.type('input[name="username"]', USERNAME)
        await page.click('input[name="password"]')
        await page.type('input[name="password"]', PASSWORD)

        # For a normal document navigation, start both waits concurrently.
        await asyncio.gather(
            page.waitForNavigation({
                "waitUntil": "networkidle2",
                "timeout": 60000,
            }),
            page.click('button[type="submit"]'),
        )

        # Replace this with a condition that proves this site's login succeeded.
        await page.waitForSelector('[data-test="account-menu"]', {"visible": True})
        print("Authenticated URL:", page.url)
    finally:
        await browser.close()

if __name__ == "__main__":
    asyncio.get_event_loop().run_until_complete(main())

The asyncio.gather pattern matters. Pyppeteer’s click() documentation warns: “If this method triggers a navigation event and there’s a separate waitForNavigation(), you may end up with a race condition that yields unexpected results.” See the Pyppeteer API reference.

3. Choose selectors that match the real page

Selectors such as input[name="username"] and button[type="submit"] are examples, not universal guarantees. Prefer stable attributes supplied by the application:

  • name or an explicit id used by the form.
  • A dedicated data-testid or data-test attribute.
  • A label relationship that remains stable across redesigns.
  • A narrowly scoped selector when a page contains more than one form.

Avoid selectors based only on generated CSS classes or the position of an element. Check that the selector identifies exactly one control before typing. For a field that appears after JavaScript renders the page:

await page.waitForSelector('form#login input[type="email"]', {"visible": True})

In Python, the object passed to Pyppeteer is a JavaScript-style dictionary, so use True in Python code as shown in the complete example. If the site uses an iframe, obtain the frame and query inside it:

frame = next(
    (f for f in page.frames if "login" in (f.url or "")),
    None,
)
if frame is None:
    raise RuntimeError("Login iframe was not found")
await frame.waitForSelector('input[name="username"]', {"visible": True})
await frame.type('input[name="username"]', USERNAME)

4. Submit and verify the authenticated state

Full-page navigation

When the form submits a new document, pair the click and navigation wait:

await asyncio.gather(
    page.waitForNavigation({"waitUntil": "networkidle2"}),
    page.click('button[type="submit"]'),
)
if "/login" in page.url:
    raise RuntimeError("The site kept us on the login page")
await page.waitForSelector('[data-test="account-menu"]', {"visible": True})

A response or a changed URL alone does not prove authentication. Sites can redirect to an error page, return a validation message, or redirect back to login. Verify a user-only element or an authenticated API result.

Single-page applications

An SPA may update its route and DOM without a document navigation. In that case, waiting forever for waitForNavigation() is the wrong synchronization method. Click the button, then wait for the authenticated UI or URL:

await page.click('button[type="submit"]')
await page.waitForSelector('[data-test="account-menu"]', {"visible": True})
if "/login" in page.url:
    raise RuntimeError("SPA login did not leave the login route")

If the application exposes a predictable post-login request, wait for that response or for a page element that appears only after the request succeeds. Use a site-specific condition rather than a fixed sleep whenever possible.

Validation errors

After clicking submit, check for inline errors before declaring failure:

error = await page.querySelector('.form-error, [role="alert"]')
if error:
    message = await page.evaluate('(el) => el.textContent', error)
    raise RuntimeError(f"Login rejected: {message.strip()}")

5. Cookies and reusable sessions

Form login normally creates cookies after the server accepts credentials. You can inspect cookies for diagnosis:

cookies = await page.cookies()
for cookie in cookies:
    print(cookie["name"], cookie.get("domain"), cookie.get("expires"))

To restore a session, set cookies before visiting the protected page. Use values obtained through an authorized login and protect them like passwords:

await page.setCookie(
    {
        "name": "session",
        "value": os.environ["SESSION_COOKIE"],
        "domain": "example.com",
        "path": "/",
        "secure": True,
        "httpOnly": True,
    }
)
await page.goto("https://example.com/account", {"waitUntil": "networkidle2"})

Cookie domains, paths, expiration, Secure, and SameSite rules must match the target site. A cookie copied from one environment may be invalid in another. Do not print cookie values or save them in source control.

6. HTTP authentication is a different login mechanism

page.authenticate() supplies credentials for HTTP authentication, such as a browser dialog generated by a server. It does not fill an ordinary HTML form.

await page.authenticate({
    "username": os.environ["HTTP_AUTH_USER"],
    "password": os.environ["HTTP_AUTH_PASSWORD"],
})
await page.goto("https://protected.example.com", {"waitUntil": "networkidle2"})

Use form interaction for a page with username and password controls, and authenticate() only when the server requests HTTP authentication.

7. MFA, CAPTCHA, and unusual login flows

  • One-time codes: stop after submitting the first factor and obtain the code through an approved test or operational process. Do not hard-code a live code.
  • Passkeys or hardware security keys: these require browser and device capabilities that a simple form script cannot reproduce. Use the site’s supported automation or a pre-authenticated test account.
  • CAPTCHA and bot checks: do not attempt to defeat them. Use an authorized test environment, a provider-supported automation path, or manual completion.
  • Consent screens: handle them as a separate, site-specific step and verify the resulting state.
  • Redirect-based identity providers: allow each expected redirect and verify the final authenticated page, not an intermediate callback.

8. Useful launch and navigation options

Option Purpose Guidance
headless Run with or without a visible browser. Use headless mode in CI; use visible mode while inspecting selectors.
executablePath Select an installed Chromium or Chrome binary. Useful when the runtime cannot download a browser.
args Pass Chromium flags. Use only flags required by your environment. Treat --no-sandbox as a container-specific compromise.
waitUntil Choose navigation completion criteria. domcontentloaded returns earlier; networkidle2 waits for a quieter network.
timeout Prevent an indefinite wait. Set a value appropriate for the target site’s normal latency and retry policy.
setViewport Control responsive layout. Set it before interacting if selectors differ by breakpoint.

Other practical controls include page.setUserAgent(), page.setExtraHTTPHeaders(), request interception, screenshots for diagnostics, and JavaScript evaluation. Apply them only when the target site’s behavior requires them, and keep credentials out of headers, screenshots, and logs.

9. Troubleshooting

Symptom Likely cause Fix
ModuleNotFoundError: pyppeteer The package is not installed in the active interpreter. Activate the intended virtual environment and run python -m pip install pyppeteer.
Chromium executable not found The first-run download was blocked or no browser path is configured. Allow the download or pass executablePath to an installed browser.
TimeoutError in waitForSelector The selector is wrong, hidden, inside an iframe, or rendered later. Inspect the live markup, wait for the correct frame, use visible, and replace brittle class selectors.
Navigation wait hangs The site is an SPA or the click did not trigger navigation. Remove waitForNavigation() and wait for the authenticated UI or route instead.
Navigation timeout after submit The site is slow, keeps connections open, or the credentials were rejected. Check validation errors, choose an appropriate waitUntil, and set a bounded timeout.
Script reports success but account is logged out A redirect or response was mistaken for authentication. Require an authenticated-only element or protected URL and inspect cookies.
Works locally but fails in CI Missing browser dependencies, different viewport, sandbox restrictions, or blocked network. Install runtime dependencies, set the viewport explicitly, configure the executable, and capture diagnostic logs.
Credentials appear in logs Debug output or exception handling printed input values. Use environment variables, redact logs, and never save passwords or session cookies in artifacts.

10. Reliability, performance, and cost considerations

  • Reliability: synchronize on observable conditions, use bounded timeouts, and retry only safe, idempotent steps. Retrying a login can trigger account lockouts or duplicate side effects.
  • Performance: reuse a browser process when performing multiple authorized tasks, create a fresh page per flow, and avoid unnecessary fixed delays. Wait for the smallest condition that proves the next action is safe.
  • Isolation: keep sessions separate when accounts must not share cookies. Close pages and browsers in a finally block.
  • Security: use test accounts where possible, restrict credential access, redact traces, and protect cookies as bearer credentials.
  • Maintenance: selectors and success conditions belong to the target site. Add a small diagnostic mode that records URL, status, and non-sensitive error text when a flow changes.

11. Pyppeteer maintenance status and Playwright

The Pyppeteer repository README states that “this repo is unmaintained and has been outside of minor changes for a long time” and suggests considering Playwright Python. That is a maintenance caveat, not proof that every existing script must be migrated.

Playwright’s authentication guide documents filling forms, waiting for a final URL or authenticated UI, and reusing saved authentication state. Its documented state model includes cookies, local storage, IndexedDB, and passkeys. The reviewed Pyppeteer documentation supports cookie operations, but it does not establish identical all-storage state reuse, so plan a migration around your application’s actual session behavior.

12. Or skip the browser setup

If your goal is to capture the page after it is publicly accessible, ScreenshotNeo provides a website screenshot API and MCP server. It does not replace an interactive login flow, but it can handle the capture step after you have an authorized URL or the required cookies and headers.

One GET request returns PNG, JPEG, WebP, or PDF. Custom cookies, headers, user agents, and Authorization are supported. The API documentation is at screenshotneo.com/docs.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/account -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/account"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/account' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server includes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account to try the 1,000 monthly screenshots without a card.

FAQ

Can Pyppeteer log in to every website?

No. Selectors, redirects, MFA, CAPTCHA, iframes, and session storage differ by site. The script must be adapted and used only with authorization.

Should I use a fixed sleep after clicking Sign in?

Usually no. Wait for navigation, a URL, a response, or an authenticated-only element that proves the next state.

Why does page.authenticate() not fill my login form?

It handles HTTP authentication challenges. An HTML form requires page interaction with the form controls.

Can I reuse a logged-in Pyppeteer session?

You can inspect and restore cookies, but verify domain, path, expiration, and security attributes. Do not expose session cookies.

Is Playwright a drop-in replacement?

No. It is a separate API and migration requires updating launch, locator, waiting, and authentication-state code. Its documentation provides a maintained authentication workflow that may suit new projects.