ScreenshotNeo

BlogHow-to

How to Convert a PDF URL to a PDF File

A PDF URL usually needs downloading, not conversion. Save it in a browser or use curl, Wget, Python, or Node.js with redirect and file checks.

By the ScreenshotNeo team1 October 20266 min read

Direct answer: a URL that points directly to a PDF normally needs to be downloaded, not converted. In Chrome, right-click the PDF link and choose Save Link As. From a terminal, use curl -L 'PDF_URL' -o 'document.pdf' or wget -O document.pdf 'PDF_URL'. A .pdf suffix alone does not prove that the server returned a PDF, so verify the downloaded file before using it.

1. Check whether the URL is really a PDF

A URL identifies a resource; it does not guarantee the response format. The address may redirect, require authentication, return an HTML page containing a PDF link, or return an error document. The server response headers and the file contents are more useful than the filename extension.

  • Open the URL in a browser. If a PDF viewer appears, use its download button.
  • If the link is on another page, copy the direct download link rather than the page URL.
  • For scripts, follow redirects and fail on HTTP errors so an error page is not saved as .pdf.

See Chrome Help’s download instructions and the curl documentation on saving downloaded files and what downloading means.

2. Download one PDF in Chrome

  1. Find the PDF link on the source page.
  2. Right-click the link and choose Save Link As.
  3. Select a folder and confirm the filename ends in .pdf.

If clicking the link opens Chrome’s PDF viewer, use the viewer’s download/save control. You can also return to the source page and use Save Link As on the link itself. A browser session may already contain cookies or a sign-in that a command-line request does not.

3. Download with cURL

Choose the local filename

curl -L 'https://example.com/path/document.pdf' -o 'document.pdf'

-L follows redirects. -o writes to the exact local path you choose.

Use the filename from the URL

curl -L -O 'https://example.com/path/document.pdf'

Uppercase -O derives the output name from the URL. To honor a server-provided Content-Disposition filename, combine -J and -O:

curl -L -OJ 'https://example.com/path/document.pdf'

Review the destination before running these commands. Remote-name mode can overwrite an existing file with the same name.

Make scripts fail on HTTP errors

curl -L --fail --show-error --silent \
  'https://example.com/path/document.pdf' \
  -o 'document.pdf'

--fail prevents a response body from being written for an unsuccessful HTTP status, which helps avoid saving an HTML error page as a PDF. Add -I when you only want to inspect headers:

curl -I -L 'https://example.com/path/document.pdf'

4. Download with Wget

wget -O document.pdf 'https://example.com/path/document.pdf'

GNU Wget’s -O (also written --output-document) writes to the chosen path. The destination is truncated immediately, so do not use a filename containing an important existing file unless overwriting is intended. Wget’s documented options are in the GNU Wget manual.

5. Download a PDF with Python

This example follows redirects, checks the HTTP status, and writes the response in binary mode.

import requests

url = "https://example.com/path/document.pdf"
response = requests.get(url, allow_redirects=True, timeout=90)
response.raise_for_status()

with open("document.pdf", "wb") as output:
    output.write(response.content)

print(f"Downloaded {len(response.content)} bytes")

For a large file, stream it instead of keeping the complete response in memory:

import requests

url = "https://example.com/path/document.pdf"
with requests.get(url, stream=True, timeout=90) as response:
    response.raise_for_status()
    with open("document.pdf", "wb") as output:
        for chunk in response.iter_content(chunk_size=1024 * 1024):
            if chunk:
                output.write(chunk)

6. Download a PDF with Node.js

Node.js 18 and later include fetch. This version follows redirects, rejects HTTP errors, and writes the binary response.

const fs = require('node:fs/promises');

const url = 'https://example.com/path/document.pdf';
const res = await fetch(url, { redirect: 'follow' });

if (!res.ok) {
  throw new Error(`HTTP ${res.status} ${res.statusText}`);
}

const bytes = Buffer.from(await res.arrayBuffer());
await fs.writeFile('document.pdf', bytes);
console.log(`Downloaded ${bytes.length} bytes`);

If the site requires a permitted authenticated workflow, provide the required headers or cookies according to that site’s documentation. Do not assume that changing the URL suffix grants access.

7. Verify that the downloaded file is a PDF

Use more than the filename extension when a download will enter an automated workflow.

file document.pdf
head -c 5 document.pdf

A normal PDF begins with the %PDF- signature. If file reports HTML or the first bytes look like an error page, inspect the response and use the direct download URL. You can also compare the response’s Content-Type header:

curl -L -D headers.txt -o document.pdf 'PDF_URL'
cat headers.txt

Servers are not required to send a perfect MIME type, so treat headers as a useful signal rather than the sole proof. The file signature and whether a PDF reader opens the file are stronger checks.

8. Troubleshooting common failures

Symptom Likely cause Fix
The saved file is an HTML page The URL returned an error, login page, or wrapper page. Use curl -L --fail, inspect headers, and locate the site’s direct PDF download link.
Only a small file downloads A redirect or access check returned a short response. Follow redirects with -L; open the URL in a browser to determine whether a permitted sign-in is required.
The PDF viewer says the file is damaged The response is not PDF bytes, or the transfer was interrupted. Check the first bytes for %PDF-, retry, and compare the downloaded size with the source’s information when available.
The command downloads nothing The server returned an unsuccessful HTTP status or refused the request. Run with --show-error, inspect the status, and use the site’s normal download workflow.
The output overwrote another file -O, Wget’s -O, or a reused destination selected an existing name. Choose a unique path with curl -o or Wget -O; check the destination before running.
The URL opens in a PDF viewer but curl is denied Your browser has cookies, credentials, or other session state. Download through the authenticated site workflow, or supply credentials only as that site permits.
The URL is a page containing a PDF link The address is not the document resource. Copy the PDF link from the page and download that URL.

9. Performance, reliability, and cost notes

  • For a one-off file, Chrome is usually the least setup.
  • For repeatable jobs, curl or Wget gives a deterministic destination and can follow redirects.
  • Use streaming writes in Python or Node.js for large documents so the entire file is not held in memory.
  • Use timeouts in application code and check status codes before writing output.
  • Retries should be limited and deliberate; retry transient network failures, but do not repeatedly retry authentication or permission errors.
  • Downloading a PDF consumes the bandwidth and storage of your client or server. The tools themselves do not require a paid conversion service.

10. Or skip the browser setup

If your goal is to capture a clean visual or PDF representation of a URL rather than preserve the original PDF bytes, ScreenshotNeo provides a single HTTP request. Its API can return PNG, JPEG, WebP, or PDF output; see the ScreenshotNeo documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/path/document.pdf -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/path/document.pdf"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/path/document.pdf' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the page verdict and billing result in headers. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account to start with 1,000 shots a month and no card.

11. FAQ

Is a PDF URL itself a PDF file?

No. It is an address. The server response must contain PDF data for the downloaded result to be a PDF.

Should I use curl -o or -O?

Use -o when you want an explicit filename. Use -O when the URL’s filename is suitable.

Why is -L commonly included?

Many short or canonical URLs redirect to the final document. -L tells curl to follow those redirects.

Can I rename an HTML response to .pdf?

Renaming changes only the filename. It does not convert the content; verify the PDF signature first.

When should I use a browser instead of a script?

Use a browser when the download depends on a session, consent step, or interactive access control. Use a script when you need repeatable downloads and a controlled output path.