ScreenshotNeo

BlogHow-to

How to Capture a Paywall Page for an Authorized Subscriber with Selenium

Use Selenium to capture a page already visible in an authorized subscriber session. Choose a viewport, element, full-page image, or PDF capture.

By the ScreenshotNeo team4 October 20269 min read

Selenium can save a screenshot of the page currently displayed in an authorized subscriber’s browser session. In Python, driver.save_screenshot("page.png") saves the current window capture as a PNG. First complete the publisher’s permitted sign-in flow, wait until the content you are allowed to view is actually rendered, and then capture it.

This guide covers saving a viewport image, capturing one element, browser-specific full-document screenshots, and printing to PDF. A subscription grants access to view content under the publisher’s terms; it does not automatically grant permission to retain or redistribute a copy. Check the account terms and applicable permissions for your intended use. See the Selenium WebDriver documentation for the documented APIs.

1. Choose the output you need

Output Use it for Important limitation
Viewport PNG The portion currently visible in the browser window It does not automatically include content below the viewport.
Element PNG A specific article or content region The page needs a stable locator, and the element must be rendered.
Full-document image A long page as one image The documented Python method here is from Selenium’s Firefox API; check support for your browser and binding.
PDF A paginated document representation Selenium documents printing with Chromium in headless mode; inspect pagination and layout.

Start with a viewport image if that is all you need. Use element capture to focus on a stable content container. For a long article, choose a full-document method supported by your browser or print to PDF; these outputs have different rendering and pagination behavior.

2. Capture the visible browser window with Python

Install Selenium in the environment that will run the script. Selenium Manager can handle driver setup for supported browser installations; if your environment manages browser drivers separately, follow that environment’s setup. This outline intentionally does not automate a publisher’s login: authenticate through the subscriber’s permitted workflow and wait for the expected content before taking the screenshot.

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

url = "https://example.com/article"
content_selector = "article"  # Replace with a selector the publisher's page actually exposes.

 driver = webdriver.Chrome()
try:
    driver.get(url)
    # Complete the site's permitted sign-in flow in this authorized browser session.
    # If sign-in is manual, finish it before waiting for the article content.
    WebDriverWait(driver, 30).until(
        EC.visibility_of_element_located((By.CSS_SELECTOR, content_selector))
    )
    driver.save_screenshot("page.png")
finally:
    driver.quit()

Remove the leading space before driver = webdriver.Chrome() if copying the snippet: it should align with url and content_selector. The selector is an example, not a claim about any publisher’s page. Replace it with a stable element that appears when the content you intend to capture is ready. The official Python API describes save_screenshot() as saving the current window capture to a PNG file.

If you do not need an explicit content wait, the minimal capture sequence is:

from selenium import webdriver

driver = webdriver.Chrome()
try:
    driver.get("https://example.com/article")
    # Sign in through the permitted workflow and wait for the desired state.
    driver.save_screenshot("page.png")
finally:
    driver.quit()

This captures the browser state at that moment. It does not verify that the page is an article rather than a sign-in prompt, that all images have loaded, or that a copy may be retained.

3. Capture a specific element

When the article content has a stable CSS selector, save just that element. The locator below is illustrative; inspect the page structure and use a selector that is appropriate for the site and your permitted use.

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

url = "https://example.com/article"
content_selector = "article"  # Replace for the target page.

driver = webdriver.Chrome()
try:
    driver.get(url)
    # Authenticate through the permitted workflow first.
    article = WebDriverWait(driver, 30).until(
        EC.visibility_of_element_located((By.CSS_SELECTOR, content_selector))
    )
    article.screenshot("article.png")
finally:
    driver.quit()

Element screenshots can avoid unrelated page chrome, but a locator that matches a hidden, empty, or wrong element will produce an unusable capture or an error. Confirm the selected element is the intended content and has nonzero size.

4. Full-document image or PDF

Full-document screenshot in Firefox

Selenium’s Firefox Python API documents get_full_page_screenshot_as_file() for a full-document screenshot. This is browser- and binding-specific; check the documentation for your installed Selenium version and browser.

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

url = "https://example.com/article"
content_selector = "article"  # Replace for the target page.

driver = webdriver.Firefox()
try:
    driver.get(url)
    # Authenticate through the permitted workflow, then wait for the content.
    WebDriverWait(driver, 30).until(
        EC.visibility_of_element_located((By.CSS_SELECTOR, content_selector))
    )
    driver.get_full_page_screenshot_as_file("full-page.png")
finally:
    driver.quit()

This method is not interchangeable with a current-window screenshot: it is documented in Firefox’s Python API. If you use another browser or language binding, check its supported full-page capture behavior rather than assuming the same method exists.

Selenium documents print_page() for PDF output and specifies Chromium in headless mode for this capability. The returned value is encoded PDF data, so decode and write it as bytes:

import base64
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
url = "https://example.com/article"
content_selector = "article"  # Replace for the target page.

driver = webdriver.Chrome(options=options)
try:
    driver.get(url)
    # Authenticate through the permitted workflow and wait for the content.
    WebDriverWait(driver, 30).until(
        EC.visibility_of_element_located((By.CSS_SELECTOR, content_selector))
    )
    pdf_data = driver.print_page()
    with open("page.pdf", "wb") as output:
        output.write(base64.b64decode(pdf_data))
finally:
    driver.quit()

Printing creates a paginated representation, not a screenshot. Check page breaks, headers, image loading, and whether the site’s print styles omit or rearrange content. Confirm the installed Chromium and Selenium versions support the headless configuration used in your environment.

5. Wait for the right page state

A successful navigation does not prove that subscriber content has rendered. Prefer an explicit wait tied to a page element or state that indicates the content you need is ready. The example uses Selenium’s visibility condition; other sites may require a different condition, such as presence, a changed page state, or completion of an interaction.

  1. Open the page in the browser session that is authorized to view it.
  2. Complete the publisher’s permitted authentication flow. Do not bypass a subscription wall, fabricate credentials, or use another person’s session.
  3. Wait for an element or state that identifies the intended content on that page.
  4. Capture only after the needed content is visible and rendered. If images or other content load later, wait on a relevant signal exposed by the page.
  5. Keep the resulting file in a protected location if it contains account or personal information.

There is no universal selector or readiness condition for publisher sites. Avoid choosing an arbitrary fixed sleep as the main readiness check: it can waste time on fast pages and still be too short on slow ones.

6. cURL, Python, and Node.js with ScreenshotNeo

For a page your use is authorized to capture, ScreenshotNeo provides a screenshot API and MCP server. This is an API alternative to setting up and maintaining a browser automation session. It cannot authenticate to a subscriber account on your behalf in the examples below; use it only for content the request is permitted and able to access.

See the ScreenshotNeo API documentation for request options. The one-call examples below capture a URL as WebP:

cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com/article \
  -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com/article"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com/article'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

These examples use the API key in a request parameter, so keep it out of public source code and logs. ScreenshotNeo supports other output and capture options, including PNG, JPEG, PDF, full-page capture, selectors, viewport and device presets, custom waits, headers, cookies, and JavaScript. Use account credentials and cookies only in a way permitted by the publisher and protect them as secrets. The API supports an MCP server for AI clients with take_screenshot, get_page_info, and capture_pdf tools.

7. Troubleshooting

Symptom Likely cause What to do
The image shows a login or subscription prompt The browser is not in the intended authorized account, or the content has not reached the expected state. Confirm the account and permitted sign-in flow, then wait for the intended content element before capturing. Selenium records the browser’s current state; it does not grant access.
Text or images are missing The page or its assets were still loading when the screenshot ran. Wait for a relevant content element or site-provided readiness state. For images, wait for the needed image state where the page exposes it.
The bottom of the article is cut off A current-window screenshot covers the viewport, not necessarily the whole document. Use a full-document method supported by your browser, capture the relevant element, or print to PDF with supported headless Chromium.
Element screenshot fails or is blank The selector does not match, matches the wrong element, or points to an element that is hidden or has no rendered size. Inspect the locator, wait for visibility, and verify the element is the intended content.
PDF printing is unavailable The documented Selenium print capability requires Chromium in headless mode. Run a supported Chromium browser headlessly and check the Selenium and browser versions in use.
The capture contains account details The page includes personal or session-specific information. Store the file securely and limit access according to your account terms and data-handling requirements.

8. Reliability, performance, and cost

Selenium runs a browser, so the capture depends on browser startup, page loading, authentication state, and the page’s rendering behavior. Reuse a session only within the authorized account workflow and protect any associated profile or credentials. For repeatable jobs, use explicit waits and save outputs with a clear naming scheme; ensure failures still close the browser, as the finally blocks above do.

Full-page images can become large for long documents, while PDF output divides content into pages and can reflow it. Capture only the artifact you need, and avoid unnecessary repeated captures. Selenium itself does not establish publisher permission, storage rights, or the cost of your browser hosting; those depend on your environment and the content’s terms. ScreenshotNeo offers 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 screenshots, with yearly billing giving two months free. Its stated billing rules exclude bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits; responses identify page verdict and billed status in headers.

9. FAQ

Does Selenium unlock a paywall?

No. Selenium automates a browser and captures the state it can access. Use a subscriber session authorized for the content; do not use automation to bypass access controls.

Does viewing access mean I can share the screenshot?

Not necessarily. Check the publisher’s account terms and permissions for storing or sharing the copy you make.

Can I use the same full-page method in every browser?

No. The full-document screenshot method described here is documented in Selenium’s Firefox Python API. Verify support for your browser and binding.

Should I save an image or a PDF?

Choose an image when you need a visual capture. Choose PDF when a paginated document representation is more useful and your Chromium headless environment supports Selenium printing.

Or skip the browser setup

ScreenshotNeo can capture a URL with one API request. Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com/article \
  -o shot.webp

Use this only where the target page is accessible and your capture is permitted. Read the API docs and start with 1,000 free screenshots a month, no card required.