ScreenshotNeo

BlogHow-to

Access Secured Pages in Python with aiohttp

Use aiohttp for Basic, Digest, bearer-token and cookie-based pages, with secure sessions, redirects, TLS, retries and troubleshooting.

By the ScreenshotNeo team29 September 20269 min read

Access Secured Pages in Python with aiohttp

To access a secured page with Python and aiohttp, identify the server’s authentication scheme, create one ClientSession, send credentials in the form the server expects, and inspect the final response rather than assuming a redirect means success. The session keeps a connection pool and cookie state, so it is the right unit for a login flow and for related requests.

This guide covers HTTP Basic, Digest, bearer or custom authorization headers, and cookie-backed login. It also explains redirects, TLS verification, response handling, retries, performance, cost, and common failures. The examples target current aiohttp documentation behavior; check the version installed in your project before shipping authentication code.

1. Install aiohttp and make a secure request

Install aiohttp in your virtual environment:

python -m pip install aiohttp

A minimal authenticated request uses an async context manager. The session closes its sockets and releases resources even when the request raises an exception.

import asyncio
import aiohttp

async def main():
    url = "https://example.com/private"

    async with aiohttp.ClientSession() as session:
        async with session.get(url) as response:
            print("status:", response.status)
            print("final URL:", response.url)
            print("redirects:", [r.url for r in response.history])
            body = await response.text()
            print(body[:500])

asyncio.run(main())

ClientSession is aiohttp’s recommended interface for HTTP requests. It owns a connection pool and a cookie jar by default. Reuse it for all requests in one logical operation instead of creating a new session per URL.

2. Choose the authentication scheme first

Credentials are not interchangeable. Read the target service’s documentation or inspect its challenge response before selecting an implementation.

Choose the authentication scheme the server requires before writing the aiohttp request.
Choose the authentication scheme the server requires before writing the aiohttp request.
Scheme Use it when aiohttp approach
Basic The server explicitly requests HTTP Basic Encode credentials and send an Authorization header
Digest The server returns an HTTP Digest challenge Use DigestAuthMiddleware documented for your installed version
Bearer or custom header An API specifies a token or custom authorization scheme Set the exact Authorization or service header
Cookie-backed login A login endpoint sets a session cookie Reuse one ClientSession so its cookie jar carries state

Never send a password as a query parameter unless the service explicitly requires it. Keep secrets in environment variables or a secret manager, and avoid logging authorization headers, cookies, or complete login payloads.

3. HTTP Basic authentication

In aiohttp 3.14, constructing BasicAuth is deprecated. The stable reference directs current code toward encode_basic_auth() and the request’s headers parameter. This example reads credentials from the environment and sends the resulting header.

import asyncio
import base64
import os
import aiohttp


def basic_header(username: str, password: str) -> str:
    raw = f"{username}:{password}".encode("utf-8")
    return "Basic " + base64.b64encode(raw).decode("ascii")


async def fetch_basic():
    username = os.environ["SITE_USERNAME"]
    password = os.environ["SITE_PASSWORD"]
    headers = {"Authorization": basic_header(username, password)}

    timeout = aiohttp.ClientTimeout(total=30)
    async with aiohttp.ClientSession(timeout=timeout) as session:
        async with session.get(
            "https://example.com/private",
            headers=headers,
            raise_for_status=False,
        ) as response:
            text = await response.text()
            if response.status == 401:
                raise RuntimeError("Credentials were rejected")
            if response.status >= 400:
                raise RuntimeError(f"HTTP {response.status}: {text[:200]}")
            return text


print(asyncio.run(fetch_basic())[:500])

Basic authentication is only appropriate over HTTPS. It encodes credentials for transport; it does not encrypt them by itself. A 401 response commonly means the username, password, realm, or expected scheme is wrong.

4. Digest authentication

Digest authentication is a challenge-response protocol. The server first returns a challenge containing values such as a realm and nonce; the client then calculates a response. Do not manually copy a Basic example and call it Digest.

The aiohttp advanced client guide documents DigestAuthMiddleware. Middleware APIs can vary by aiohttp release, so verify the signature in the documentation for the version pinned by your application. A representative pattern is:

import asyncio
import aiohttp
from aiohttp import DigestAuthMiddleware

async def main():
    middleware = DigestAuthMiddleware("user", "password")
    async with aiohttp.ClientSession(middlewares=(middleware,)) as session:
        async with session.get("https://example.com/digest-private") as response:
            print(response.status)
            print((await response.text())[:500])

asyncio.run(main())

If this raises an import or constructor error, check the installed aiohttp version and use that version’s official advanced-client documentation. A server that advertises only Basic will not become Digest-compatible because the client uses this middleware.

5. Bearer tokens and custom authorization headers

For token APIs, send exactly the scheme documented by the service. Bearer tokens normally use the Authorization header:

import asyncio
import os
import aiohttp

async def main():
    headers = {
        "Authorization": f"Bearer {os.environ['API_TOKEN']}",
        "Accept": "application/json",
    }
    async with aiohttp.ClientSession(headers=headers) as session:
        async with session.get("https://api.example.com/account") as response:
            print(response.status)
            data = await response.json(content_type=None)
            print(data)

asyncio.run(main())

Session-level headers are convenient when every request uses the same token. Use per-request headers when a session talks to multiple services or when scopes differ. For a custom scheme, replace the value with the service’s required format; do not assume Bearer is accepted.

Many secured pages require a login POST followed by one or more GET requests. Keep those calls in the same session so cookies set by the login response remain available.

import asyncio
import os
import aiohttp

async def fetch_after_login():
    timeout = aiohttp.ClientTimeout(total=30)
    async with aiohttp.ClientSession(timeout=timeout) as session:
        login_payload = {
            "username": os.environ["SITE_USERNAME"],
            "password": os.environ["SITE_PASSWORD"],
        }
        async with session.post(
            "https://example.com/login",
            data=login_payload,
            allow_redirects=False,
        ) as login:
            if login.status not in (200, 201, 302, 303):
                raise RuntimeError(f"Login failed with HTTP {login.status}")

        async with session.get("https://example.com/private") as page:
            if page.status in (401, 403):
                raise RuntimeError("The session is not authorized")
            return await page.text()

print(asyncio.run(fetch_after_login())[:500])

Real login forms may require a CSRF token, a specific content type, hidden fields, a preliminary GET, or an extra verification step. Follow the site’s documented flow. A cookie jar cannot bypass multi-factor authentication, bot checks, or an account that lacks permission.

7. Redirects, cookies and authorization safety

aiohttp follows redirects by default. You can disable that behavior with allow_redirects=False when you need to inspect a login redirect or preserve strict navigation rules.

When a redirect changes host or protocol, aiohttp removes the Authorization header. This protects credentials from being forwarded to a different origin, but it also explains why a request can authenticate at the first URL and arrive unauthenticated at the final one. Always inspect:

async with session.get(url, allow_redirects=True) as response:
    print("final:", response.url)
    print("status:", response.status)
    for redirect in response.history:
        print("redirect:", redirect.status, redirect.url)
    print("content type:", response.headers.get("Content-Type"))

A successful HTTP status does not prove that you reached the private page. Some sites return a 200 login form. Check the final URL, redirect history, content type, and an application-specific marker in the body.

8. TLS verification, timeouts and response handling

TLS certificate validation is enabled by default. Keep it enabled. Setting ssl=False disables certificate validation and is not a normal fix for an authentication error. Resolve missing CA certificates, hostname mismatches, or clock problems in the environment instead.

Set explicit timeouts so a stalled server does not consume a worker forever:

timeout = aiohttp.ClientTimeout(
    total=60,
    connect=10,
    sock_connect=10,
    sock_read=45,
)

async with aiohttp.ClientSession(timeout=timeout) as session:
    async with session.get(url) as response:
        response.raise_for_status()
        content = await response.read()

Use raise_for_status() when every non-2xx response is exceptional. Use raise_for_status=False while diagnosing authentication, redirects, or an API that returns useful JSON error details. Choose response.text() for HTML or text, response.json() for JSON, and response.read() for binary content.

9. cURL, Python and Node.js equivalents

cURL is useful for separating a server or credential problem from an aiohttp implementation problem.

curl --verbose --user "$SITE_USERNAME:$SITE_PASSWORD" \
  --location https://example.com/private

For a bearer token:

curl --verbose \
  -H "Authorization: Bearer $API_TOKEN" \
  https://api.example.com/account

The equivalent Python call with the synchronous requests library is:

import os
import requests

r = requests.get(
    "https://example.com/private",
    auth=(os.environ["SITE_USERNAME"], os.environ["SITE_PASSWORD"]),
    timeout=30,
)
r.raise_for_status()
print(r.text[:500])

Node.js has no built-in cookie-login workflow, but Basic and bearer requests are straightforward:

const username = process.env.SITE_USERNAME;
const password = process.env.SITE_PASSWORD;
const basic = Buffer.from(`${username}:${password}`).toString('base64');

const res = await fetch('https://example.com/private', {
  headers: { Authorization: `Basic ${basic}` }
});

if (!res.ok) throw new Error(`HTTP ${res.status}`);
console.log((await res.text()).slice(0, 500));

10. Performance and reliability

  • Reuse one session for related requests. Connection pooling and keep-alive reduce handshake overhead.
  • Bound concurrency with a semaphore or connector limit when fetching many pages. Unbounded tasks can exhaust sockets or trigger rate limits.
  • Retry only transient failures such as connection resets, timeouts, and selected 5xx responses. Do not blindly retry 401, 403, or a failed login.
  • Use exponential backoff with jitter and honor Retry-After when supplied.
  • Stream large downloads instead of loading them all into memory; authenticate the request before opening the response body.
  • Record status, final URL, elapsed time, and a request identifier, but redact credentials and session cookies.

Authentication itself does not determine your infrastructure cost. Your costs come from compute, bandwidth, proxies, target-service limits, and any paid API you call. Cache permitted resources, avoid duplicate logins, and respect the target’s terms and rate limits.

11. Troubleshooting checklist

Symptom Likely cause Fix
401 Unauthorized Wrong credentials or scheme Read the WWW-Authenticate header; match Basic, Digest, or bearer exactly.
403 Forbidden Valid identity lacks permission, or policy blocks the client Check account roles, IP policy, user agent requirements, and service documentation.
Final page is a login form Cookies were not retained, CSRF was missing, or a redirect changed origin Reuse one session; inspect response.history, final URL, cookies, and required hidden fields.
Authorization disappears Redirect changed host or protocol Use the canonical HTTPS URL, inspect redirects, and authenticate the final permitted origin.
SSL certificate error Broken CA store, hostname, certificate chain, or system clock Repair trust configuration. Do not disable verification as a routine workaround.
Timeout Slow origin, blocked network, or a request waiting indefinitely Set phase-specific timeouts, check connectivity, and retry only transient failures.
Digest middleware error Documentation and installed aiohttp versions differ Pin a supported version and consult its matching advanced-client reference.
JSON decode failure The server returned HTML or an error page Check status and Content-Type before parsing; log a redacted response prefix.

12. Or skip the browser setup

If your end goal is a clean image or PDF of a secured page rather than application data, ScreenshotNeo provides a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. You can pass custom headers, cookies, user agents, and Authorization values, then capture a page without maintaining a browser stack.

A capture service can remove common overlays before returning the page image.
A capture service can remove common overlays before returning the page image.

Or skip the browser setup:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for parameters and response headers. Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and each response reports its result through X-Page-Verdict and X-Billed headers. An MCP server gives Claude, Cursor, and other MCP clients take_screenshot, get_page_info, and capture_pdf tools. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

13. Frequently asked questions

Can aiohttp log in to any website?

No. It can implement the protocol the service exposes, but it cannot bypass missing permissions, MFA, CAPTCHAs, bot controls, or terms that prohibit automation.

Should I create a session for every request?

Usually no. Reuse a session for related requests to retain cookies and connection pooling, then close it with an async context manager.

Why does a 200 response still show “sign in”?

Many login pages return HTTP 200. Inspect the final URL, redirect history, cookies, and page markers before treating the request as authenticated.

Is ssl=False safe?

It disables certificate validation. Keep TLS verification enabled and repair the environment’s trust configuration instead.

When should I use ScreenshotNeo?

Use it when the deliverable is a clean screenshot or PDF and you prefer an API or MCP workflow over operating a browser and its login state.