ScreenshotNeo

BlogHow-to

How to Convert a Password-Protected Web Page to PDF with PDFCrowd

Convert a protected page with PDFCrowd by supplying HTTP Basic credentials or session cookies. Learn the workflow, code, alternatives, and troubleshooting steps.

By the ScreenshotNeo team4 October 20269 min read

Yes. PDFCrowd can convert password-protected pages. Use HTTP Basic authentication for pages protected with a username and password, or pass the session cookie for a page that requires a logged-in browser session. For the HTTP API, send a form-encoded POST request to the versioned endpoint and save the successful response body as PDF bytes. PDFCrowd’s FAQ documents these approaches and two alternatives.

Keep the credentials for your source website separate from the PDFCrowd username and API key. The latter authenticate your conversion request to PDFCrowd; source-site credentials let the converter load the protected page. See the HTTP API guide and parameter reference.

1. Identify how the page is protected

First determine what a browser must provide to access the page:

  • HTTP Basic authentication: The browser requests a username and password, often through a browser-native authentication dialog. Configure PDFCrowd’s HTTP authentication for the source page.
  • Cookie or session authentication: You first sign in to the site, and the browser then sends a session cookie with page requests. Configure the cookie with PDFCrowd’s cookie support.
  • Other authentication flows: A login form, MFA challenge, SSO redirect, device verification, or bot check may require additional steps. The cited documentation does not promise that every such flow will work. If the converter cannot obtain the page, use an alternative workflow below.

Do not put your PDFCrowd API key in source-page authentication fields, or source-site passwords in PDFCrowd’s API authentication. These credentials serve different systems.

2. Convert a page protected by HTTP Basic authentication

PDFCrowd’s FAQ names setHttpAuth() for Basic authentication. In its HTTP API, the request itself uses HTTP Basic authentication with your PDFCrowd username and API key. The source page’s authentication is configured separately as conversion input. Check the parameter reference for the exact accepted source-authentication fields for your integration.

The following command shows the HTTP API request shape, including the PDFCrowd API credentials and page URL. Add the source-page HTTP-auth fields supported by your account/API integration; do not substitute the PDFCrowd API credentials for those source credentials.

curl --fail-with-body --silent --show-error \
  --user 'PDFCROWD_USERNAME:PDFCROWD_API_KEY' \
  --output protected.pdf \
  --form-string 'url=https://example.com/protected/' \
  --form-string 'content_viewport_width=balanced' \
  'https://api.pdfcrowd.com/convert/24.04/'

For an SDK integration, follow the documented API FAQ approach: configure the source page’s HTTP authentication with setHttpAuth(), then submit the URL for conversion. Keep the SDK client’s PDFCrowd username and API key in its own client configuration.

For cookie/session-based access, PDFCrowd documents setCookies(). Obtain the cookie through your application’s normal, authorized login flow, then pass the relevant session cookie to the converter. A cookie is a bearer credential: anyone who obtains a valid session cookie may be able to act as that logged-in session. Avoid putting cookies in source control, logs, URLs, or client-side code.

The HTTP API request still posts a URL as a form field and authenticates to PDFCrowd with HTTP Basic authentication. Configure the source session cookie using the documented cookie option for your API client; consult the parameter reference for its supported representation.

curl --fail-with-body --silent --show-error \
  --user 'PDFCROWD_USERNAME:PDFCROWD_API_KEY' \
  --output account.pdf \
  --form-string 'url=https://example.com/account/report' \
  --form-string 'content_viewport_width=balanced' \
  'https://api.pdfcrowd.com/convert/24.04/'

This command is the basic URL conversion request. Add the source session cookie through PDFCrowd’s documented setCookies() configuration in the client integration; the command above does not itself authenticate to the protected source page.

4. HTTP API examples in Python and Node.js

These runnable examples make a standard PDFCrowd URL conversion request, save the returned bytes, and keep the PDFCrowd credentials separate from source-page authentication. To convert protected content, configure its HTTP authentication or cookie with the corresponding PDFCrowd client option (setHttpAuth() or setCookies()) as described in the FAQ and API reference. Do not treat successful authentication to the conversion endpoint as proof that the source page was authenticated.

Python

import requests
from requests.auth import HTTPBasicAuth

endpoint = "https://api.pdfcrowd.com/convert/24.04/"
response = requests.post(
    endpoint,
    auth=HTTPBasicAuth("PDFCROWD_USERNAME", "PDFCROWD_API_KEY"),
    data={
        "url": "https://example.com/protected/",
        "content_viewport_width": "balanced",
    },
    timeout=120,
)
response.raise_for_status()

content_type = response.headers.get("Content-Type", "")
if "pdf" not in content_type.lower():
    raise RuntimeError(
        f"Expected a PDF response, got {content_type!r}: "
        f"{response.text[:1000]}"
    )

with open("protected.pdf", "wb") as pdf_file:
    pdf_file.write(response.content)

Node.js

import { writeFile } from "node:fs/promises";

const endpoint = "https://api.pdfcrowd.com/convert/24.04/";
const credentials = Buffer.from(
  "PDFCROWD_USERNAME:PDFCROWD_API_KEY"
).toString("base64");
const form = new FormData();
form.set("url", "https://example.com/protected/");
form.set("content_viewport_width", "balanced");

const response = await fetch(endpoint, {
  method: "POST",
  headers: { Authorization: `Basic ${credentials}` },
  body: form,
  signal: AbortSignal.timeout(120_000),
});

if (!response.ok) {
  const detail = await response.text();
  throw new Error(`PDFCrowd returned HTTP ${response.status}: ${detail}`);
}

const contentType = response.headers.get("content-type") ?? "";
if (!contentType.toLowerCase().includes("pdf")) {
  throw new Error(`Expected PDF bytes, got Content-Type: ${contentType}`);
}

await writeFile("protected.pdf", Buffer.from(await response.arrayBuffer()));

In production, read the PDFCrowd credentials and source-site credentials from a secret store or environment variables. The examples use placeholders and do not include a source cookie or password. Do not add credentials to a URL query string.

5. Choose the right input and formatting options

Authentication answers whether the converter can access the page. Conversion settings control the resulting PDF’s layout. The HTTP API accepts form fields, not a JSON request body, and a successful response contains PDF bytes. Keep the documented versioned endpoint explicit: https://api.pdfcrowd.com/convert/24.04/.

Need Approach
Convert a live page Send its address in the url form field.
Control page layout Use settings such as page_size, margins, headers, and footers.
Fix print styling Use custom_css; use custom_javascript for post-load DOM changes.
Wait for page content Use javascript_delay or wait_for_element where appropriate.
Convert a specific part Use element_to_convert.
Convert HTML you already retrieved Use text for HTML content or file for an uploaded document.
Include local HTML assets Package the HTML and assets in an archive while preserving relative paths. For accessible remote assets, use absolute URLs or a suitable <base> element.

These settings do not log in to the source site. For the complete accepted values, defaults, and constraints, use the parameter reference. For diagnosis, the guide documents ?errfmt=json, response headers, debug_log=true, and preserving the error response body.

6. Alternatives when direct authenticated loading is unsuitable

PDFCrowd documents two alternatives:

  1. Temporarily expose the content at a hard-to-guess URL. The content owner can make it available briefly and revoke or invalidate the URL after it has been visited or after a time limit. Treat the URL as a secret while it is active.
  2. Render the protected HTML yourself and submit it. Retrieve the page through your application’s authorized session, then send the HTML using PDFCrowd’s convertString() or convertFile() client methods, or the HTTP API’s HTML input. Include required assets in an archive or refer to assets at accessible absolute URLs; a <base> element can resolve relative URLs.

Choose an alternative only if you are authorized to access and transfer the page content. Rendered HTML can contain sensitive data, and external assets must be reachable by the converter for the document to look as intended. See the PDFCrowd FAQ for the documented alternatives.

7. Troubleshooting

Symptom Likely cause What to check
Conversion returns a login page or access-denied page The source authentication was missing, expired, or configured in the wrong form. Identify Basic auth versus session cookies. Configure source credentials separately from PDFCrowd API credentials; refresh the session cookie if needed.
PDFCrowd rejects the API request The PDFCrowd username/API key is wrong or the request was not encoded as form fields. Check the Basic auth pair and send form data. The HTTP API guide says JSON request bodies are not supported.
Output is HTML or an error message rather than a PDF The HTTP request failed, or the client saved an error response as if it were PDF bytes. Check the status code and Content-Type before writing the file. Preserve the response body; try ?errfmt=json for structured API errors.
Images or styles are missing The converter cannot access relative or local asset paths. Use accessible absolute URLs or a <base> element, or package the HTML and local assets in an archive.
Content is incomplete or dynamic sections are absent The page content was not ready when captured, or authentication expired before resources loaded. Verify the session remains valid, then consider a documented wait setting such as javascript_delay or wait_for_element.
Client times out The page or its resources take longer to load than the client’s timeout. Set a suitable client timeout, inspect the API error, and reduce unnecessary page resources where you control the page. Do not retry indefinitely.

For a useful diagnostic request, the API guide shows saving response headers and body, printing the HTTP status, enabling debug_log=true, and using --fail-with-body with cURL. Treat debug output and error bodies as potentially sensitive if they include request details.

8. Security, reliability, and cost considerations

  • Protect credentials: Keep API keys, passwords, and session cookies on the server. Redact authorization headers and cookie values from logs. Use a narrowly scoped, short-lived session where your site supports it.
  • Check the output: A successful HTTP response should be a PDF response. Validate status and content type before persisting bytes; optionally verify the file opens as a PDF in your downstream workflow.
  • Handle retries carefully: Retry transient transport failures with a bounded policy. Do not retry authentication or permission errors as if they were temporary network problems. Check whether the page’s session remains valid before retrying.
  • Account for page variability: A page may depend on scripts, fonts, images, or session state. These influence render time and appearance. Test the necessary output settings against the actual page.
  • Plan cost from your PDFCrowd account terms: The research sources do not establish a price or conversion quota, so check PDFCrowd’s current plan and billing information before estimating spend. Avoid assuming failed conversions are free or billed without verifying its terms.

Or skip the browser setup

If you need a screenshot rather than a PDF, ScreenshotNeo is a website screenshot API and MCP server. Its one-call API returns PNG, JPEG, or WebP. For this protected-page workflow, use PDFCrowd’s PDF methods above; ScreenshotNeo’s documented endpoint is for screenshots and PDFs, and its supported authentication options are described in its API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

ScreenshotNeo accepts cookie and consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, with the page verdict and billing status shown in response headers. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. See the docs for the options and supported outputs.

Sign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.

FAQ

Can PDFCrowd convert a page protected by a password?

Yes. PDFCrowd documents HTTP Basic authentication and cookie/session authentication. The correct method depends on how the source website protects the page.

Are my PDFCrowd API key and source website password interchangeable?

No. The PDFCrowd username and API key authenticate the conversion request. Source credentials or cookies authorize access to the page being converted.

Can I send a JSON body to the HTTP API?

The HTTP API guide specifies form fields and says JSON request bodies are not supported.

What if I can access the page but the converter cannot?

PDFCrowd documents temporarily exposing content at a hard-to-guess URL or submitting rendered HTML with its string/file conversion methods. Include or make available the HTML’s assets as needed.