How to Capture a Web Page Behind a Login Using Selenium for a Report
Use Selenium to sign in through the site’s normal flow, verify the report is ready, and save a screenshot or PDF with the context you need.
To capture a web page behind a login with Selenium, use an account authorized to view the page, complete the site’s normal sign-in flow, wait for a page-specific sign that the report is ready, then save a screenshot or print the page to PDF. Verify the saved artifact before attaching it to a report. Selenium automates a real browser through WebDriver; it does not provide a universal login recipe or a way to bypass access controls.
This guide uses Python and Chrome. Selectors and login steps are examples: replace them with the controls and ready-state marker for the site you are permitted to access. Site-approved single sign-on (SSO) or multi-factor authentication (MFA) may require an approved interactive step or a dedicated reporting integration.
1. Confirm access and prepare a safe environment
Before automating a login, make sure the account is explicitly authorized to access the page and use its contents in the report. Prefer a low-privilege account created for this task. Do not store a password, session cookie, or token in source code, screenshots, logs, or the report.
ChromeDriver’s security guidance says it should not run with a privileged account and recommends a protected environment, such as a container or virtual machine. Keep browser-control ports restricted, especially when using Selenium Server or another remote-control service. See the [ChromeDriver security guidance](https://developer.chrome.com/docs/chromedriver/security-considerations).
Install Selenium with pip and install a compatible Chrome browser and ChromeDriver for your environment. Selenium’s current driver-management behavior may handle driver setup automatically where supported; follow the official [Selenium installation documentation](https://www.selenium.dev/documentation/webdriver/getting_started/install_library/) and your environment’s browser setup instructions.
python -m pip install selenium
The example below reads credentials from environment variables. Set them through your secret manager or shell, not in the script. The URLs and selectors are placeholders; use the real sign-in URL, form fields, and report page for your authorized site.
2. Sign in, wait for the report, and save a screenshot
Navigation reaching the document’s ready state does not guarantee that a JavaScript application has finished loading its report. Use an explicit wait for a condition that is meaningful on the authenticated page, such as the report title or a data table appearing. Selenium advises against mixing implicit and explicit waits because timeout behavior can become unpredictable.
import os
from datetime import datetime, timezone
from pathlib import Path
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
LOGIN_URL = "https://example.com/login"
REPORT_URL = "https://example.com/reports/monthly"
OUTPUT = Path("artifacts/monthly-report.png")
username = os.environ["REPORT_USERNAME"]
password = os.environ["REPORT_PASSWORD"]
OUTPUT.parent.mkdir(parents=True, exist_ok=True)
driver = webdriver.Chrome()
try:
driver.set_window_size(1440, 1000)
wait = WebDriverWait(driver, 30)
driver.get(LOGIN_URL)
wait.until(EC.visibility_of_element_located((By.NAME, "username"))).send_keys(username)
driver.find_element(By.NAME, "password").send_keys(password)
driver.find_element(By.CSS_SELECTOR, "button[type='submit']").click()
# Replace with a site-approved MFA/SSO step if required. Do not attempt to
# defeat a challenge or automate a step the site does not permit.
wait.until(EC.url_contains("/dashboard"))
driver.get(REPORT_URL)
wait.until(EC.visibility_of_element_located((By.CSS_SELECTOR, "h1.report-title")))
wait.until(EC.visibility_of_element_located((By.CSS_SELECTOR, "table.report-data")))
# Optional: fail early if the page still shows a loading indicator.
wait.until(EC.invisibility_of_element_located((By.CSS_SELECTOR, ".report-loading")))
if "login" in driver.current_url.lower():
raise RuntimeError("The browser returned to a login URL; report was not captured")
driver.save_screenshot(str(OUTPUT))
captured_at = datetime.now(timezone.utc).isoformat()
print(f"Saved {OUTPUT}; URL={driver.current_url}; captured_at={captured_at}")
finally:
driver.quit()
Install or update Selenium and use the browser setup recommended by its [official documentation](https://www.selenium.dev/documentation/webdriver/getting_started/install_library/). The sample assumes field names, a dashboard redirect, and report selectors that may not match your site. Inspect the authorized page to identify stable selectors; avoid selectors tied to generated CSS classes when the site offers semantic attributes or stable IDs.
Why wait for more than navigation?
Selenium’s [waiting strategies documentation](https://www.selenium.dev/documentation/webdriver/waits/) explains that navigation waits for a page-load strategy’s readyState—by default, complete—before returning control. Client-side scripts can still fetch and render report data afterward. An explicit wait polls for a specified condition until it succeeds or times out.
Choose a ready marker that belongs to the actual report: its heading, a table with expected columns, an account marker, or the disappearance of a loading indicator. When the report can legitimately contain no rows, wait for a report container and an explicit empty-state marker as alternatives. A visible heading alone may not prove that the data is complete.
Check the resulting file before using it. Confirm it shows the intended report, date range, filters, and account context, rather than a sign-in screen, error page, or loading shell. A rendered capture shows what the browser displayed; it does not prove that underlying data is correct.
3. Choose a screenshot or PDF
| Output | Useful when | Check before attaching |
|---|---|---|
| Screenshot (PNG) | You need visual evidence of the browser-rendered view or a specific screen state. | Whether it captures just the visible viewport or the full page. Full-page support varies by browser and API. |
| You need a paginated, printable attachment or a document that readers can page through. | Page breaks, scale, paper size, margins, background rendering, and whether print layout differs from the browser view. |
Save the current browser view as an image
The Python WebDriver API provides save_screenshot (also known as get_screenshot_as_file) to save the current window as a PNG file. The example above uses it after the report conditions pass. If you need a full-page image, confirm the selected browser and Selenium interface support that capture mode; a viewport screenshot is not automatically a full-page capture.
Print the current page to PDF
Selenium’s print-page capability returns PDF data for the current page. The following example requires Chromium’s headless mode, as noted in Selenium’s [windows and tabs documentation](https://www.selenium.dev/documentation/webdriver/interactions/windows/). Run the same authorized login and report-ready waits first, then replace the screenshot call with this code:
import base64
from pathlib import Path
# Run this after login and the report-ready waits, using a headless Chrome session.
result = driver.print_page()
pdf_bytes = base64.b64decode(result)
Path("artifacts/monthly-report.pdf").write_bytes(pdf_bytes)
Selenium’s [print page documentation](https://www.selenium.dev/documentation/webdriver/interactions/print_page/) describes print options such as orientation, page ranges, paper size, margins, scale, backgrounds, and shrink-to-fit. Check the installed Selenium binding’s API for how to set these options. Preview the PDF: browser print CSS can omit or rearrange content compared with the visible page.
4. Make the capture useful in a report
Record enough context for a reader to understand what the artifact represents. A small sidecar text file or report caption can include:
- The exact report URL and capture time in UTC.
- The report’s date range, filters, and relevant view or account context, when sharing that context is permitted.
- The output type, browser and Selenium versions, and whether the browser ran headlessly.
- Any caveat about pagination, dynamic data, or a view that changed during capture.
Do not include passwords, cookies, authentication tokens, or unnecessary personal information. A screenshot or PDF is not automatically a complete archival record and does not establish the correctness of the data. Follow your organization’s reporting and retention rules; Selenium’s capture documentation does not define a legal chain-of-custody standard.
5. Reliability and performance considerations
- Use condition-based waits. Waiting for the actual report marker avoids both capturing too early and adding an arbitrary long sleep to every run. Set a finite timeout and treat a timeout as a failed capture.
- Keep the browser session scoped. Create the browser for the job and close it in a
finallyblock so failed logins and timeouts do not leave a process running. - Keep output deterministic. Set a consistent viewport, use a fixed report period where possible, and record filters and capture time. Live reports can change between runs.
- Limit unnecessary work. Navigate directly to the report after authentication, wait for only the elements required to establish readiness, and avoid repeated refreshes that may trigger rate limits or account protections.
- Run with least privilege. Use an isolated environment and account with access only to the needed site and data. Keep ChromeDriver inaccessible from untrusted networks.
- Account for site policy. Session expiry, MFA, SSO, bot defenses, and account permissions can interrupt automation. Use the site’s supported flow or an approved reporting interface; do not try to evade challenges.
Local Selenium has no per-screenshot service charge, but you still pay in browser infrastructure, runtime, maintenance, and operational handling of credentials and outputs. There is no universal speed figure: page size, scripts, network conditions, authentication steps, and report rendering all affect completion time.
6. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
TimeoutException waiting for a field or report element |
Selector mismatch, slow rendering, wrong page, or the user lacks access. | Inspect the authorized page and current URL; confirm the selector and account permissions. Increase the timeout only when the page legitimately needs longer, and keep the condition specific. |
| Screenshot shows the login page | Credentials were rejected, session expired, redirect did not complete, or the report URL requires additional authorization. | Check the sign-in result and final URL before capture. Follow approved SSO/MFA steps and verify that the account can open the report manually. |
| Screenshot shows a blank table or loading shell | Document navigation finished before client-side report data arrived. | Wait for a data-ready condition, such as populated rows or a deliberate empty-state marker; also wait for the loading indicator to disappear. |
| Element is present but not interactable | The control is hidden, covered by an overlay, outside the viewport, or still changing. | Wait for visibility or clickability, handle site-approved overlays, and scroll the element into view if needed. Do not click through security challenges. |
| Chrome or ChromeDriver fails to start | Browser and driver are incompatible, browser dependencies are missing, or the process lacks a suitable environment. | Use compatible current browser components and consult Selenium’s setup guidance. Check container dependencies and run as a non-privileged user. |
| PDF printing is unsupported or fails | The selected browser or session does not support the print command, or Chromium is not running headlessly. | Use a supported Chromium headless session for print-page output, or save a screenshot when a PDF is not required. Confirm print options against the installed binding. |
| Waits take much longer than expected | Implicit and explicit waits are mixed, or a broad condition never becomes true. | Use explicit waits for defined conditions, avoid setting an implicit wait alongside them, and inspect the page state when a timeout occurs. |
| Capture contains the wrong report period or filters | Defaults changed, filters did not apply, or the report updated during capture. | Set filters through the supported interface, wait for the report to update, and record the exact period and filter context with the artifact. |
7. Or skip the browser setup
If the page is publicly reachable, [ScreenshotNeo](https://screenshotneo.com) can return an image or PDF with one GET request. A screenshot API cannot sign into a private account for you; use it only for pages it can access without circumventing a login. For an authorized private page, keep the Selenium flow above or use an approved way to make the content public to the capture service.
For a public page, this cURL example saves a WebP image. See the [ScreenshotNeo API documentation](https://screenshotneo.com/docs/) for parameters and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
- Cookie banners are accepted and removed, along with known newsletter popups and chat widgets, before the shot.
- Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers indicate the page verdict and billing status.
- An MCP server lets AI agents use tools to take screenshots, get page information, and capture PDFs.
- 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for ScreenshotNeo’s free plan to try 1,000 screenshots a month with no card.
FAQ
Can Selenium capture a page that requires MFA?
Only through a flow the site permits. Use its approved MFA or SSO process, or ask for an authorized reporting integration. Selenium does not make it appropriate to bypass a challenge.
Does a screenshot prove the report data is accurate?
No. It records a rendered browser view. Include the report period and filters, and validate data through the system or process responsible for it.
Should I use a screenshot or a PDF?
Use a screenshot for a visual state of the browser and a PDF for paginated, print-friendly sharing. Inspect either artifact because print layout and full-page support vary.
Can ScreenshotNeo capture my logged-in private page?
The one-call example is for a page the capture service can reach. It does not log in to a private site or bypass its access controls.


