ScreenshotNeo

BlogEngineering

How SSL Works in Web Scraping APIs

Learn how TLS handshakes, certificates, proxies, and mTLS affect web scraping APIs—and how to fix SSL errors without disabling verification.

By the ScreenshotNeo team1 October 20267 min read

Short answer: SSL is the older name still used in scraper settings and error messages. Modern HTTPS uses TLS. A scraping client opens a TLS connection, verifies the target certificate and hostname, negotiates encryption keys, and then sends HTTP requests through the encrypted session. If a proxy, CDN, or scraping API sits in the middle, each network leg has its own TLS connection and certificate checks.

TLS protects data in transit and authenticates the endpoint. It does not grant permission to scrape, bypass a bot challenge, or authorize access to a private page.

SSL and TLS: what the terms mean

SSL versions are obsolete. TLS is the protocol used by current HTTPS connections; TLS 1.3 is the current version, while TLS 1.2 remains widely deployed. The three security properties are encryption, integrity, and authentication: traffic is confidential, modifications are detectable, and the client can verify the server identity. See MDN’s TLS overview.

What happens during a scraper’s TLS connection

  1. Client hello: The scraper sends supported TLS versions, cipher suites, extensions, and the requested hostname through SNI.
  2. Server hello and certificate: The server selects parameters and returns an X.509 certificate chain.
  3. Certificate validation: The client checks that the chain leads to a trusted CA, the certificate covers the requested DNS name, it is currently valid, and the server proves possession of the corresponding private key.
  4. Key agreement: The peers exchange key material and derive temporary session keys.
  5. Encrypted HTTP: The scraper sends its request and receives the response inside the authenticated, encrypted TLS session.

The exact handshake messages differ between TLS versions, but the result is the same: HTTP data is protected after the handshake completes.

Certificate verification explained

A certificate is not simply a permission slip. It binds a hostname to a public key and is signed by a certificate authority (CA) that the client trusts. Verification normally checks:

  • the certificate’s validity dates;
  • the requested hostname against the Subject Alternative Name field;
  • the complete chain, including required intermediate certificates;
  • the signature of each certificate up to a trusted root CA;
  • proof that the server controls the private key.

The client also needs two inputs: a trusted CA store and the DNS hostname it is connecting to. OpenSSL’s s_client documentation is useful for inspecting these details.

Why scraper SSL errors happen

Error pattern Likely cause Correct fix
certificate verify failed Unknown CA, incomplete chain, or an expired certificate Repair the target’s chain or configure the correct CA bundle.
hostname mismatch The certificate does not include the hostname in the URL Use the certificate’s real hostname or replace the certificate.
unable to get local issuer certificate A client image lacks a root or intermediate CA Update the OS CA package or point the client at a maintained bundle.
certificate has expired Expired leaf or intermediate certificate Renew the certificate and deploy the full chain.
wrong version number HTTPS was attempted against a plain HTTP port, or a proxy protocol is wrong Check the scheme, port, and proxy configuration.
handshake failure No compatible TLS version or cipher, or a server policy rejected the client Upgrade the runtime and inspect supported TLS policies.
self signed certificate The server uses a private CA that the client does not trust Install the private CA explicitly when you control the environment.

Do not solve these errors by disabling verification. In Python Requests, verify=False accepts expired or mismatched certificates and permits man-in-the-middle attacks. Requests verifies HTTPS certificates by default; use its CA bundle configuration instead.

Inspect a target with OpenSSL

openssl s_client -connect example.com:443 -servername example.com -showcerts -verify_return_error

Look for the negotiated protocol, certificate subject and SAN names, issuer chain, expiration dates, and the final verification result. Always include -servername because virtual hosts often return different certificates based on SNI.

Runnable client examples

cURL

curl --verbose --fail --location https://example.com/ -o response.html

Use a specific CA bundle when required:

curl --cacert /path/to/company-ca.pem --fail https://internal.example/

Use --insecure only for a controlled diagnostic, never in production scraping code.

Python Requests

import requests

url = "https://example.com/"
r = requests.get(url, timeout=(10, 60))
r.raise_for_status()
print(r.status_code, r.url)
open("response.html", "wb").write(r.content)

For a private CA, pass its bundle:

r = requests.get(
    "https://internal.example/",
    verify="/path/to/company-ca.pem",
    timeout=(10, 60),
)
r.raise_for_status()

Requests also honors the REQUESTS_CA_BUNDLE environment variable. Keep hostname verification enabled.

Node.js

const res = await fetch('https://example.com/');
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const body = await res.text();
console.log(body.length);

For a private CA, configure an HTTPS agent with the CA certificate in the client library you use. Do not set rejectUnauthorized: false in production.

TLS when a proxy or scraping API is involved

Draw every network leg separately:

  1. Your application to the scraping API: your client verifies the API hostname and certificate.
  2. Scraping API to the target: the API verifies the target hostname and certificate independently.
  3. CDN to origin: a CDN may terminate TLS at its edge and establish another TLS connection to the origin.

A successful first leg does not prove that the target’s certificate is valid. Conversely, a target certificate problem may be visible only to the gateway. Log which hop failed and preserve the underlying TLS error.

When to use mutual TLS (mTLS)

Standard TLS authenticates the server to the client. mTLS adds client authentication: the caller presents a client certificate and the server validates it. Use mTLS when an API or private origin must restrict access to approved scraper services. Cloudflare describes this as bidirectional certificate-based trust in its mTLS documentation.

mTLS requires certificate issuance, private-key protection, rotation, and revocation procedures. It is separate from API keys, HTTP basic authentication, cookies, and authorization headers. A server can require both mTLS and an application credential.

Security checklist for scraping clients

  • Keep certificate verification enabled.
  • Use TLS 1.2 or TLS 1.3 through a maintained runtime.
  • Keep the operating system CA store current.
  • Set connection and read timeouts.
  • Limit redirects or validate redirected hosts when sensitive credentials are present.
  • Never log private keys, cookies, authorization headers, or full signed URLs.
  • Rotate private CAs and mTLS certificates before expiration.
  • Verify that scraping is permitted by the target’s terms, robots policy, and applicable law.

Performance, reliability, and cost

TLS adds a handshake before the first request. Reusing HTTP keep-alive connections avoids repeated handshakes; HTTP/2 can multiplex requests over one connection when supported. DNS lookup, TCP setup, TLS negotiation, server processing, and response transfer should be measured separately.

For reliable workers, use bounded connect and read timeouts, retry only transient network failures, and apply exponential backoff with a limit. Do not blindly retry certificate validation failures: they usually require configuration or server repair. Cache the CA bundle locally, but refresh it through your operating-system update process. Monitor certificate expiry and alert before renewal windows.

In a hosted scraping API, cost may be based on requests, successful captures, bandwidth, or compute time. Read the provider’s billing definition and record status, cache, and failure outcomes so retries do not create unexpected charges.

Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. It handles the browser and HTTPS connection for a capture request. Before the shot, it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing result.

See the ScreenshotNeo API documentation for options such as custom headers, cookies, authorization, user agents, waiting rules, request blocking, caching, full-page capture, element selection, PDFs, async jobs, bulk capture, and signed links.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

An MCP server also lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.

Troubleshooting checklist

  1. Confirm the URL scheme and hostname exactly match the certificate.
  2. Run OpenSSL with SNI and inspect the complete chain.
  3. Check the scraper machine’s clock and timezone.
  4. Update the CA store inside the container or virtual machine.
  5. Test without the proxy, then test each proxy leg separately.
  6. Compare the runtime’s TLS versions and cipher policy with the server’s requirements.
  7. If the target uses a private CA, install that CA explicitly rather than disabling verification.
  8. For mTLS, verify the client certificate, private-key permissions, chain, and server trust store.
  9. Capture the original exception, hostname, proxy endpoint, and negotiated protocol in diagnostics.

FAQ

Is SSL still used by web scraping APIs?

The name remains in configuration and error messages, but secure HTTPS connections use TLS.

Does TLS hide my scraper’s identity?

TLS encrypts the connection and authenticates endpoints. The target can still observe the API gateway’s IP address, HTTP headers, cookies, and request behavior.

Can TLS bypass a CAPTCHA?

No. TLS protects transport; bot detection and authorization are application-layer controls.

Should I pin a certificate?

Only when you control the operational lifecycle and can rotate pins safely. A normal maintained CA store is more resilient for public websites.

Why does a browser work while my scraper fails?

The browser may have a newer CA store, different TLS support, installed enterprise roots, or a different proxy path. Compare those environments before changing verification settings.