ScreenshotNeo

BlogHow-to

How to Reuse Browser Cookies for Web Scraping

Transfer authorized browser cookies to Playwright, Selenium, Requests, or cURL safely, fix 401/403 errors, and automate reliable scraping.

By the ScreenshotNeo team1 October 20269 min read

Short answer: obtain cookies from a browser session you control, preserve each cookie’s name, value, domain, path, expiry, Secure, HttpOnly, SameSite, and partitioning metadata, then load them into the same site context before requesting the target URL. Use Playwright or Selenium when the site needs browser JavaScript and interactions; use a persistent Python Requests session or cURL when a normal HTTP client is sufficient.

A cookie is state issued by a server in a Set-Cookie response and returned by a user agent in a Cookie request header when its scope and policy allow it. RFC 6265 defines this mechanism: “This document defines the HTTP Cookie and Set-Cookie header fields.” Keep reuse limited to accounts, domains, and data for which you have explicit authorization.

  1. Log in through a browser session that you own or are authorized to automate.
  2. Export cookies without exposing their values in source control, tickets, screenshots, or logs.
  3. Preserve metadata: name, value, domain, path, expires, httpOnly, secure, sameSite, and, where provided, partitionKey.
  4. Install the cookies in the matching browser context or an HTTP cookie jar.
  5. Navigate to the target URL on the same scheme and host scope.
  6. Check the response, redirect chain, and application state. A cookie alone may not replace CSRF tokens, device checks, or other session state.
  7. Rotate or revoke the session when the job ends.
Attribute Effect Common failure
Domain Limits which host names receive the cookie. Host-only cookies apply only to the host that set them. A cookie from app.example.com is sent to neither every subdomain nor an unrelated domain.
Path Limits the URL paths where the cookie is sent. A cookie scoped to /account may not be sent to /api.
Expires/Max-Age Controls expiration; session cookies disappear when the browser session ends. An exported value is already expired or was evicted.
Secure Cookie is sent only over HTTPS. Testing over HTTP silently omits it.
HttpOnly Blocks JavaScript access through document.cookie. A page script cannot read it; use browser automation or an authorized export.
SameSite Controls cross-site sending, including Strict and Lax behavior. A cross-site request lacks the cookie even though the domain appears correct.
Partitioning Some browsers partition state by the top-level site as well as the cookie domain. A cookie works in one embedded context but not another.

Do not build a global Cookie header from a copied string unless you have verified scope yourself. A cookie jar lets the client apply domain, path, expiry, and secure rules for each request.

3. Reuse cookies with Playwright

Playwright is the browser-faithful option. It executes JavaScript, follows browser navigation rules, and lets you capture and restore a complete context. BrowserContext.cookies() returns all cookies or only cookies affecting supplied URLs; addCookies() installs cookie objects into the context. Either a URL or both a domain and path are required when adding a cookie.

import { chromium } from 'playwright';

const browser = await chromium.launch();
const context = await browser.newContext();

// Log in or otherwise establish an authorized session here.
const loginPage = await context.newPage();
await loginPage.goto('https://example.com/login', { waitUntil: 'networkidle' });
// Fill the login form using your approved test credentials and flow.
// await loginPage.fill('#email', process.env.EXAMPLE_EMAIL);
// await loginPage.fill('#password', process.env.EXAMPLE_PASSWORD);
// await loginPage.click('button[type="submit"]');

const cookies = await context.cookies('https://example.com/target');
// Persist cookies securely if another process needs them. Never print values.
const saved = cookies;

const next = await browser.newContext();
await next.addCookies(saved);
const page = await next.newPage();
const response = await page.goto('https://example.com/target', {
  waitUntil: 'domcontentloaded',
  timeout: 30000
});

console.log({ status: response?.status(), url: page.url() });
console.log((await page.title()).slice(0, 120));
await browser.close();

For a reusable authenticated state file, Playwright can save the context state and load it into a later context. Treat that file like a password: restrict permissions, encrypt it at rest, and delete it when no longer needed.

// Save after an authorized login
await context.storageState({ path: 'state.json' });

// Restore for a later run
const restored = await browser.newContext({ storageState: 'state.json' });
const page = await restored.newPage();
await page.goto('https://example.com/target');

When Playwright is the right choice

  • The page requires JavaScript to render data.
  • You must click, scroll, wait for a selector, or satisfy a normal browser navigation flow.
  • The application uses browser storage or redirects in addition to cookies.
  • You need to observe the final page rather than call a stable HTTP endpoint.

4. Reuse cookies with Selenium

Selenium requires the driver to be in the relevant browser context before adding a cookie. Navigate to the target domain first, add the authorized cookie, then load the protected path.

import os
from selenium import webdriver
from selenium.webdriver.chrome.options import Options

options = Options()
options.add_argument('--headless=new')
driver = webdriver.Chrome(options=options)

try:
    driver.get('https://example.com/')
    driver.add_cookie({
        'name': 'session',
        'value': os.environ['EXAMPLE_SESSION'],
        'domain': 'example.com',
        'path': '/',
        'secure': True,
        # Selenium accepts SameSite values such as 'Strict' and 'Lax'.
        'sameSite': 'Lax',
    })
    driver.get('https://example.com/target')
    print(driver.title)
finally:
    driver.quit()

Use driver.get_cookies() to inspect the current browser jar and driver.get_cookie('name') for one cookie. Redact the value field before writing diagnostics.

5. Reuse cookies with Python Requests

Requests’ Session persists cookies across requests made by that session. This is efficient for stable HTTP endpoints that do not need browser JavaScript, challenge handling, or DOM interaction.

import os
import requests

session = requests.Session()
session.cookies.set(
    'session',
    os.environ['EXAMPLE_SESSION'],
    domain='example.com',
    path='/'
)

response = session.get('https://example.com/target', timeout=20)
response.raise_for_status()
print(response.url)
print(response.text[:500])

If you have an exported cookie-jar file, load it with the appropriate Requests cookie-jar format instead of flattening every value into one header. A session also keeps cookies received during redirects and subsequent requests.

6. Reuse cookies with cURL

For a simple HTTP continuation, cURL can read a Netscape-format cookie file. The file must contain only authorized cookies and should have restrictive permissions.

curl --cookie cookies.txt \
  --location \
  --fail-with-body \
  --connect-timeout 10 \
  --max-time 30 \
  'https://example.com/target' \
  -o response.html

To persist cookies received during a run, combine --cookie with --cookie-jar. Avoid putting live values directly in shell history. A hand-written header such as -H 'Cookie: session=…' bypasses useful scope checks and is easy to send to the wrong host.

Symptom Likely cause Fix
401 immediately Expired session, wrong cookie name, or the server expects another session cookie. Log in again, inspect the authorized browser context, and transfer the complete relevant jar.
403 only on a subdomain Domain or path does not cover that host or route. Capture cookies for the exact target URL and preserve domain/path metadata.
Works in Chrome, fails in HTTP code JavaScript, redirects, CSRF tokens, or browser state are also required. Use Playwright or Selenium, or reproduce the documented API flow instead of copying one cookie.
Works on HTTPS, fails locally over HTTP Secure prevents transmission over an insecure channel. Use HTTPS on the authorized target.
document.cookie is missing the value The cookie is HttpOnly. Use browser automation or an authorized cookie export.
Fails only in an iframe or cross-site request SameSite or browser third-party-cookie policy withholds it. Reproduce the same top-level site context or use the service’s supported authentication method.
Intermittent 403 after several requests Short session lifetime, server-side binding to account/device/IP signals, or rate controls. Keep a stable authorized session, use reasonable request rates, refresh through the normal login flow, and honor access controls.
Redirects to login The cookie was attached to the wrong host, was not loaded, or was evicted. Log status and redirect locations without logging values; verify the jar before navigation.

8. A practical debugging checklist

  • Confirm the target URL uses the same scheme, host, and path assumptions as the browser request.
  • Check expiry and whether the browser evicted the cookie.
  • Compare cookie names and domains, never values in shared logs.
  • Use browser developer tools or automation inspection to confirm whether the browser actually sends the cookie.
  • Check for a CSRF token in a form, header, or page state that must accompany the cookie.
  • Follow redirects and inspect the final status code.
  • Test one authorized request before adding concurrency.
  • Remove secrets from trace output, HAR files, screenshots, crash reports, and CI artifacts.

9. Security and authorization

Cookies can act as bearer credentials. Use only sessions and targets for which you have explicit authorization, and follow the site’s terms, access controls, robots guidance where applicable, and applicable law. Encrypt stored cookies, restrict file permissions, minimize retention, rotate or revoke sessions after use, and redact values from logs and screenshots. Never ask someone to paste a live authentication cookie into chat or a public issue.

10. Performance, reliability, and cost

  • Browser overhead: Playwright and Selenium pay for browser startup, JavaScript execution, rendering, and navigation. Reuse one browser process and create isolated contexts when safe.
  • HTTP efficiency: Requests and cURL are lighter when the endpoint is stable and browser execution is unnecessary. Reuse a session to avoid repeated handshakes and login flows.
  • Concurrency: Start sequentially, then add bounded concurrency that respects the site’s limits and your authorization. More workers do not repair expired or incorrectly scoped cookies.
  • Timeouts and retries: Set connect and total timeouts. Retry transient network failures with backoff, but do not blindly retry authentication failures or 403 responses.
  • State freshness: Session cookies can expire or be revoked. Refresh through the normal authorized login process instead of copying a stale value indefinitely.
  • Cost: Self-hosted browser runs consume compute and maintenance time. An API can move browser setup and rendering out of your application; compare its per-request price with your browser infrastructure and engineering time.

11. Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. You can make one GET request for a PNG, JPEG, WebP, or PDF without maintaining your own browser capture service. See the ScreenshotNeo API documentation for options and authentication.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Before capture, ScreenshotNeo accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and each response reports its result in X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account to try it.

12. Frequently asked questions

Can I reuse cookies from one browser on another machine?

Yes, when you are authorized and you transfer the complete supported cookie or storage state securely. Host, path, expiry, secure, SameSite, and server-side session binding can still make the session invalid on the second machine.

Should I use Playwright, Selenium, or Requests?

Use Playwright or Selenium when you need browser behavior and interaction. Use Requests when a stable HTTP endpoint and cookie jar are enough. cURL is useful for a small, inspectable HTTP continuation.

Cookies marked HttpOnly are intentionally hidden from document.cookie. Read them through an authorized browser automation API or export mechanism instead.

No. A cookie is only one part of session state. The service may evaluate browser behavior, device signals, IP reputation, or a challenge independently.

How long should I retain exported cookies?

Retain them only for the job that needs them, protect them like credentials, and revoke or delete them when the job ends.