How to Log In to a Webpage with Pyppeteer
Use Pyppeteer to fill a real login form, handle navigation safely, verify authentication, and troubleshoot redirects, MFA, cookies, and SPA flows.
Direct answer: open the login page, wait for the actual form controls, fill the username and password fields, submit the form, and verify authentication with a site-specific signal such as the final URL or an authenticated-only element. If the submit click causes navigation, await the click and navigation together with asyncio.gather; waiting for them separately can race.
Pyppeteer is an unofficial Python port of Puppeteer. Its project repository currently says it is unmaintained and recommends considering Playwright Python for new work. The procedure below remains useful for maintaining an existing Pyppeteer script, but confirm behavior against the version installed in your environment and the site you are authorized to access.
1. Install Pyppeteer and prepare Chromium
Create an isolated environment and install the package:
python -m venv .venv
source .venv/bin/activate # Windows: .venv\\Scripts\\activate
python -m pip install --upgrade pip
python -m pip install pyppeteer
On first use, Pyppeteer downloads Chromium when it cannot find a suitable executable. In CI or a locked-down server, provide a browser executable explicitly with executablePath or install Chromium through your system image. Keep the browser version and the Pyppeteer package under review because the repository is no longer actively maintained.
2. A complete form-login script
This runnable template uses placeholder selectors. Inspect the authorized target page and replace every URL, selector, and success condition with values from that site.
import asyncio
import os
from pyppeteer import launch
LOGIN_URL = "https://example.com/login"
USERNAME = os.environ["SITE_USERNAME"]
PASSWORD = os.environ["SITE_PASSWORD"]
async def main():
browser = await launch(
headless=True,
args=["--no-sandbox"], # Use only when required by your container runtime.
)
page = await browser.newPage()
try:
await page.setViewport({"width": 1440, "height": 900})
response = await page.goto(
LOGIN_URL,
{"waitUntil": "domcontentloaded", "timeout": 60000},
)
if response is None:
raise RuntimeError("The login navigation returned no response")
await page.waitForSelector('input[name="username"]', {"visible": True})
await page.waitForSelector('input[name="password"]', {"visible": True})
await page.click('input[name="username"]')
await page.type('input[name="username"]', USERNAME)
await page.click('input[name="password"]')
await page.type('input[name="password"]', PASSWORD)
# For a normal document navigation, start both waits concurrently.
await asyncio.gather(
page.waitForNavigation({
"waitUntil": "networkidle2",
"timeout": 60000,
}),
page.click('button[type="submit"]'),
)
# Replace this with a condition that proves this site's login succeeded.
await page.waitForSelector('[data-test="account-menu"]', {"visible": True})
print("Authenticated URL:", page.url)
finally:
await browser.close()
if __name__ == "__main__":
asyncio.get_event_loop().run_until_complete(main())
The asyncio.gather pattern matters. Pyppeteer’s click() documentation warns: “If this method triggers a navigation event and there’s a separate waitForNavigation(), you may end up with a race condition that yields unexpected results.” See the Pyppeteer API reference.
3. Choose selectors that match the real page
Selectors such as input[name="username"] and button[type="submit"] are examples, not universal guarantees. Prefer stable attributes supplied by the application:
nameor an explicitidused by the form.- A dedicated
data-testidordata-testattribute. - A label relationship that remains stable across redesigns.
- A narrowly scoped selector when a page contains more than one form.
Avoid selectors based only on generated CSS classes or the position of an element. Check that the selector identifies exactly one control before typing. For a field that appears after JavaScript renders the page:
await page.waitForSelector('form#login input[type="email"]', {"visible": True})
In Python, the object passed to Pyppeteer is a JavaScript-style dictionary, so use True in Python code as shown in the complete example. If the site uses an iframe, obtain the frame and query inside it:
frame = next(
(f for f in page.frames if "login" in (f.url or "")),
None,
)
if frame is None:
raise RuntimeError("Login iframe was not found")
await frame.waitForSelector('input[name="username"]', {"visible": True})
await frame.type('input[name="username"]', USERNAME)
4. Submit and verify the authenticated state
Full-page navigation
When the form submits a new document, pair the click and navigation wait:
await asyncio.gather(
page.waitForNavigation({"waitUntil": "networkidle2"}),
page.click('button[type="submit"]'),
)
if "/login" in page.url:
raise RuntimeError("The site kept us on the login page")
await page.waitForSelector('[data-test="account-menu"]', {"visible": True})
A response or a changed URL alone does not prove authentication. Sites can redirect to an error page, return a validation message, or redirect back to login. Verify a user-only element or an authenticated API result.
Single-page applications
An SPA may update its route and DOM without a document navigation. In that case, waiting forever for waitForNavigation() is the wrong synchronization method. Click the button, then wait for the authenticated UI or URL:
await page.click('button[type="submit"]')
await page.waitForSelector('[data-test="account-menu"]', {"visible": True})
if "/login" in page.url:
raise RuntimeError("SPA login did not leave the login route")
If the application exposes a predictable post-login request, wait for that response or for a page element that appears only after the request succeeds. Use a site-specific condition rather than a fixed sleep whenever possible.
Validation errors
After clicking submit, check for inline errors before declaring failure:
error = await page.querySelector('.form-error, [role="alert"]')
if error:
message = await page.evaluate('(el) => el.textContent', error)
raise RuntimeError(f"Login rejected: {message.strip()}")
5. Cookies and reusable sessions
Form login normally creates cookies after the server accepts credentials. You can inspect cookies for diagnosis:
cookies = await page.cookies()
for cookie in cookies:
print(cookie["name"], cookie.get("domain"), cookie.get("expires"))
To restore a session, set cookies before visiting the protected page. Use values obtained through an authorized login and protect them like passwords:
await page.setCookie(
{
"name": "session",
"value": os.environ["SESSION_COOKIE"],
"domain": "example.com",
"path": "/",
"secure": True,
"httpOnly": True,
}
)
await page.goto("https://example.com/account", {"waitUntil": "networkidle2"})
Cookie domains, paths, expiration, Secure, and SameSite rules must match the target site. A cookie copied from one environment may be invalid in another. Do not print cookie values or save them in source control.
6. HTTP authentication is a different login mechanism
page.authenticate() supplies credentials for HTTP authentication, such as a browser dialog generated by a server. It does not fill an ordinary HTML form.
await page.authenticate({
"username": os.environ["HTTP_AUTH_USER"],
"password": os.environ["HTTP_AUTH_PASSWORD"],
})
await page.goto("https://protected.example.com", {"waitUntil": "networkidle2"})
Use form interaction for a page with username and password controls, and authenticate() only when the server requests HTTP authentication.
7. MFA, CAPTCHA, and unusual login flows
- One-time codes: stop after submitting the first factor and obtain the code through an approved test or operational process. Do not hard-code a live code.
- Passkeys or hardware security keys: these require browser and device capabilities that a simple form script cannot reproduce. Use the site’s supported automation or a pre-authenticated test account.
- CAPTCHA and bot checks: do not attempt to defeat them. Use an authorized test environment, a provider-supported automation path, or manual completion.
- Consent screens: handle them as a separate, site-specific step and verify the resulting state.
- Redirect-based identity providers: allow each expected redirect and verify the final authenticated page, not an intermediate callback.
8. Useful launch and navigation options
| Option | Purpose | Guidance |
|---|---|---|
headless |
Run with or without a visible browser. | Use headless mode in CI; use visible mode while inspecting selectors. |
executablePath |
Select an installed Chromium or Chrome binary. | Useful when the runtime cannot download a browser. |
args |
Pass Chromium flags. | Use only flags required by your environment. Treat --no-sandbox as a container-specific compromise. |
waitUntil |
Choose navigation completion criteria. | domcontentloaded returns earlier; networkidle2 waits for a quieter network. |
timeout |
Prevent an indefinite wait. | Set a value appropriate for the target site’s normal latency and retry policy. |
setViewport |
Control responsive layout. | Set it before interacting if selectors differ by breakpoint. |
Other practical controls include page.setUserAgent(), page.setExtraHTTPHeaders(), request interception, screenshots for diagnostics, and JavaScript evaluation. Apply them only when the target site’s behavior requires them, and keep credentials out of headers, screenshots, and logs.
9. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
ModuleNotFoundError: pyppeteer |
The package is not installed in the active interpreter. | Activate the intended virtual environment and run python -m pip install pyppeteer. |
| Chromium executable not found | The first-run download was blocked or no browser path is configured. | Allow the download or pass executablePath to an installed browser. |
TimeoutError in waitForSelector |
The selector is wrong, hidden, inside an iframe, or rendered later. | Inspect the live markup, wait for the correct frame, use visible, and replace brittle class selectors. |
| Navigation wait hangs | The site is an SPA or the click did not trigger navigation. | Remove waitForNavigation() and wait for the authenticated UI or route instead. |
| Navigation timeout after submit | The site is slow, keeps connections open, or the credentials were rejected. | Check validation errors, choose an appropriate waitUntil, and set a bounded timeout. |
| Script reports success but account is logged out | A redirect or response was mistaken for authentication. | Require an authenticated-only element or protected URL and inspect cookies. |
| Works locally but fails in CI | Missing browser dependencies, different viewport, sandbox restrictions, or blocked network. | Install runtime dependencies, set the viewport explicitly, configure the executable, and capture diagnostic logs. |
| Credentials appear in logs | Debug output or exception handling printed input values. | Use environment variables, redact logs, and never save passwords or session cookies in artifacts. |
10. Reliability, performance, and cost considerations
- Reliability: synchronize on observable conditions, use bounded timeouts, and retry only safe, idempotent steps. Retrying a login can trigger account lockouts or duplicate side effects.
- Performance: reuse a browser process when performing multiple authorized tasks, create a fresh page per flow, and avoid unnecessary fixed delays. Wait for the smallest condition that proves the next action is safe.
- Isolation: keep sessions separate when accounts must not share cookies. Close pages and browsers in a
finallyblock. - Security: use test accounts where possible, restrict credential access, redact traces, and protect cookies as bearer credentials.
- Maintenance: selectors and success conditions belong to the target site. Add a small diagnostic mode that records URL, status, and non-sensitive error text when a flow changes.
11. Pyppeteer maintenance status and Playwright
The Pyppeteer repository README states that “this repo is unmaintained and has been outside of minor changes for a long time” and suggests considering Playwright Python. That is a maintenance caveat, not proof that every existing script must be migrated.
Playwright’s authentication guide documents filling forms, waiting for a final URL or authenticated UI, and reusing saved authentication state. Its documented state model includes cookies, local storage, IndexedDB, and passkeys. The reviewed Pyppeteer documentation supports cookie operations, but it does not establish identical all-storage state reuse, so plan a migration around your application’s actual session behavior.
12. Or skip the browser setup
If your goal is to capture the page after it is publicly accessible, ScreenshotNeo provides a website screenshot API and MCP server. It does not replace an interactive login flow, but it can handle the capture step after you have an authorized URL or the required cookies and headers.
One GET request returns PNG, JPEG, WebP, or PDF. Custom cookies, headers, user agents, and Authorization are supported. The API documentation is at screenshotneo.com/docs.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/account -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/account"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/account' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server includes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.
Create a free ScreenshotNeo account to try the 1,000 monthly screenshots without a card.
FAQ
Can Pyppeteer log in to every website?
No. Selectors, redirects, MFA, CAPTCHA, iframes, and session storage differ by site. The script must be adapted and used only with authorization.
Should I use a fixed sleep after clicking Sign in?
Usually no. Wait for navigation, a URL, a response, or an authenticated-only element that proves the next state.
Why does page.authenticate() not fill my login form?
It handles HTTP authentication challenges. An HTML form requires page interaction with the form controls.
Can I reuse a logged-in Pyppeteer session?
You can inspect and restore cookies, but verify domain, path, expiration, and security attributes. Do not expose session cookies.
Is Playwright a drop-in replacement?
No. It is a separate API and migration requires updating launch, locator, waiting, and authentication-state code. Its documentation provides a maintained authentication workflow that may suit new projects.


