ScreenshotNeo

BlogHow-to

How to Get the File Type of a URL in Python

Use Python’s mimetypes module for an offline extension guess, or inspect HTTP Content-Type for the live resource. Learn how to handle redirects, compressed files, and unknown types.

By the ScreenshotNeo team29 September 202610 min read

How to Get the File Type of a URL in Python

To guess a file type from a URL without making a network request, use Python’s standard-library mimetypes.guess_type(). It returns a MIME type and a separate content-encoding hint; either can be None when the suffix is missing or unknown. For a live resource, prefer the HTTP response’s Content-Type header, fall back to the final URL’s path suffix, and call the result a declaration or guess—not proof of the bytes.

For example, mimetypes.guess_type("https://example.com/archive.tar.gz?download=1") commonly returns ("application/x-tar", "gzip"). The URL suffix check is fast and offline. Checking a live resource costs a request, but works for extensionless URLs and accounts for redirects.

1. Guess the type from the URL suffix

Python documents mimetypes as a way to convert a filename or URL into a MIME type associated with its extension. No request is made, and the input need not be a real, reachable URL.

from mimetypes import guess_type

url = "https://example.com/archive.tar.gz?download=1"
mime_type, encoding = guess_type(url)

print(mime_type)  # commonly application/x-tar
print(encoding)   # commonly gzip

The first result describes the media type associated with the filename suffix. The second describes an encoding inferred from the suffix, such as gzip for .gz. Preserve both values if the distinction matters to your application.

For inputs that may contain query parameters or fragments, split the URL and pass only its path. This makes it explicit that ?download=report.pdf is not itself a filename suffix.

from mimetypes import guess_type
from urllib.parse import urlsplit

url = "https://example.com/download?file=report.pdf#details"
path = urlsplit(url).path
mime_type, encoding = guess_type(path)

print(path)      # /download
print(mime_type) # None

The urllib.parse module provides the standard URL component-splitting interface. Parsing the path avoids accidentally treating query or fragment text as evidence about the resource.

Strict and non-strict mappings

guess_type() uses strict=True by default. In strict mode the mapping table is limited to official IANA media types. Set strict=False to include common non-standard mappings too.

import mimetypes

filename = "example"
print(mimetypes.guess_type(filename))                 # (None, None)
print(mimetypes.guess_type("image.svg", strict=True))
print(mimetypes.guess_type("image.svg", strict=False))

The mappings available can depend on the Python runtime and its environment. If your application needs a stable allowlist, define and validate that allowlist explicitly instead of assuming every machine has identical MIME mappings.

2. Check the live HTTP response

A suffix is only a filename-based hint. When you need to know what a server says it is returning, inspect Content-Type. The following example tries HEAD first, follows redirects, and uses the final response URL for a suffix fallback.

import mimetypes
from urllib.parse import urlsplit

import requests


def file_type_from_url(url: str) -> str | None:
    response = requests.head(
        url,
        allow_redirects=True,
        timeout=10,
    )

    content_type = response.headers.get("Content-Type", "")
    declared = content_type.split(";", 1)[0].strip().lower()
    if declared and declared != "application/octet-stream":
        return declared

    final_path = urlsplit(response.url).path
    path_type, _encoding = mimetypes.guess_type(final_path)
    return path_type


print(file_type_from_url("https://example.com/report.pdf"))

Install the dependency with python -m pip install requests. Requests exposes response headers through the returned Response object. This function returns a normalized declared type when available, otherwise the suffix guess, or None if neither supplies a type.

Why filter the header value?

Headers often include parameters, such as text/html; charset=utf-8. Splitting at the semicolon yields the media type. The example also treats application/octet-stream as generic, allowing a more informative suffix guess to be used. If your application needs to distinguish “server declared generic binary” from “header absent,” retain the original header separately rather than discarding that information.

HEAD support and streamed GET fallback

HEAD is intended to retrieve response metadata without the response body, so it can be an efficient first request. Some servers reject it or behave differently for HEAD than for GET. There is no universal guarantee that it works for every URL. A robust client can fall back to GET, stream the response, and inspect headers before consuming the body.

import mimetypes
from urllib.parse import urlsplit

import requests


def file_type_from_url(url: str, timeout: float = 10) -> str | None:
    with requests.Session() as session:
        try:
            response = session.head(
                url, allow_redirects=True, timeout=timeout
            )
            if response.status_code in (405, 501):
                response.close()
                response = session.get(
                    url,
                    allow_redirects=True,
                    stream=True,
                    timeout=timeout,
                )
        except requests.RequestException:
            # A production caller may choose a GET fallback for other
            # HEAD failures too; handle network errors at the call site.
            raise

        try:
            header = response.headers.get("Content-Type", "")
            declared = header.split(";", 1)[0].strip().lower()
            if declared and declared != "application/octet-stream":
                return declared

            path_type, _encoding = mimetypes.guess_type(
                urlsplit(response.url).path
            )
            return path_type
        finally:
            response.close()

The streamed GET requests headers and leaves the body unread. Closing the response releases its resources. A real download can still consume bandwidth if the server sends data quickly or the client reads it; use streaming and close promptly when you only need headers. Choose a timeout appropriate to your application and pass it consistently to every request.

3. Choose the signal that fits the job

Signal Network request Useful when Main limitation
URL suffix via mimetypes No You need a cheap offline guess Extension can be absent or misleading
HTTP Content-Type Yes You need the server’s declared type, including extensionless routes Header can be absent, generic, stale, or wrong
Inspect downloaded bytes Usually Format correctness or security policy matters Requires format-specific parsing or signature detection

Use the header first for a fetched resource, but retain the suffix as a fallback. Treat both as claims or hints. If your code will parse, execute, store, or serve a file based on its type, inspect the bytes with a parser or signature detector suitable for the formats you allow. There is no universal magic-byte check that verifies every possible format.

A URL suffix is an offline hint; the response header is the server’s declaration for the live resource.
A URL suffix is an offline hint; the response header is the server’s declaration for the live resource.

4. Handle common URL and response edge cases

  • Query strings and fragments: Split the URL with urlsplit() and inspect .path, not the full string. Query parameters can contain filenames that do not describe the resource.
  • Redirects: A redirect can lead to a different path or resource. Enable redirect following when appropriate, then use response.url for the suffix fallback.
  • Compressed suffixes: Keep the encoding result. For a name ending in .tar.gz, the type can describe the tar archive while the encoding indicates gzip.
  • Unknown suffix: Return None or an explicit “unknown” state. Do not invent a type from a filename that has no recognized extension.
  • Generic header: Decide deliberately whether application/octet-stream is useful to your caller. It is a valid declaration, but it does not identify a more specific format.
  • Data URLs: CPython’s mimetypes implementation handles data: URLs from their declared media type. Do not assume that behavior validates embedded content.
  • Case and parameters: Normalize header media types if comparing them. Remove parameters before comparison, but preserve the original value if you need diagnostics.
  • Request failures: DNS failures, TLS errors, timeouts, and HTTP errors are different from an unknown suffix. Report or handle them separately so callers can distinguish “could not fetch” from “fetched, but type unknown.”
After redirects, use the final response URL for suffix fallback and inspect bytes when validation matters.
After redirects, use the final response URL for suffix fallback and inspect bytes when validation matters.

5. cURL, Python, and Node.js examples

The same distinction applies outside Python: inspect the response header for a live resource, then use its path suffix if the server’s declaration is unavailable or generic.

cURL: read response headers

curl -sS -I -L --max-time 10 "https://example.com/report.pdf"

Look for the final response’s Content-Type header. Some servers do not support HEAD; try a GET that asks for headers while discarding the body:

curl -sS -L --max-time 10 -D - -o /dev/null \
  "https://example.com/report.pdf"

Python: inspect a response directly

import requests

url = "https://example.com/report.pdf"
with requests.get(url, stream=True, timeout=10) as response:
    print(response.status_code)
    print(response.headers.get("Content-Type"))
    print(response.url)  # final URL after redirects

Check the status code according to your use case. A server can return an error page with a type such as text/html; the header describes that response, not necessarily the file you originally expected.

Node.js: fetch headers

const response = await fetch("https://example.com/report.pdf", {
  method: "HEAD",
  redirect: "follow",
  signal: AbortSignal.timeout(10_000),
});

console.log(response.status);
console.log(response.headers.get("content-type"));
console.log(response.url); // final URL after redirects

If HEAD is not supported, use GET and avoid reading the body. Fetch follows redirects by default in standard server-side use, but the final response URL remains the value to check. Network errors and non-success HTTP statuses should be handled separately from a missing content type.

6. Troubleshooting

Symptom Likely cause Fix
guess_type() returns (None, None) No recognized suffix, or the URL path is a route such as /download Check the live Content-Type; keep unknown as a valid outcome
The result seems to include query text The full URL was treated as a filename Parse with urlsplit(url).path before guessing
HEAD returns 405 or 501 The server does not implement HEAD for that endpoint Use a streamed GET and inspect headers without reading its body
The header says application/octet-stream The server sent a generic binary declaration Use a suffix fallback, or inspect bytes if a specific format is required
The header says HTML for a supposed PDF The URL may redirect to an error, login page, or HTML response Check status, final URL, and response body policy before treating it as a PDF
Request hangs or fails intermittently No timeout, slow server, or transient network failure Set connect/read timeouts, handle request exceptions, and retry only when suitable
Type differs across machines Different Python or system MIME mappings Use an explicit mapping or allowlist for application-critical behavior

7. Performance, reliability, and cost

A suffix guess is constant, local work: it needs no DNS lookup, TLS connection, or server response. It is the best choice when a quick hint is enough. A HEAD request generally avoids downloading the response body, but still incurs network latency and can fail or be unsupported. A streamed GET gives compatibility with servers that reject HEAD, yet needs careful response cleanup and can still consume resources. Reading and validating bytes costs more bandwidth and CPU, but is the appropriate path when correctness or security depends on the actual format.

For reliability, define the result states your application needs: declared type, suffix guess, validated type, and unknown are useful distinct states. Keep request errors separate from unknown MIME types. Apply bounded timeouts, respect redirect behavior, and avoid trusting remote metadata to choose unsafe processing paths. When a format is accepted, validate it with a parser or detector designed for that format.

For batch jobs, reuse an HTTP session where appropriate and avoid fetching bodies solely to classify them. Cache metadata only when the resource identity and freshness requirements allow it; a URL can serve different content over time. The cost of metadata lookup is a network request and its operational load, while a suffix check is free of network cost. The exact cost depends on your hosting and traffic environment.

8. Or skip the browser setup

This task is about identifying a URL’s response type; ScreenshotNeo is useful when the next step is capturing how that page looks. It is a website screenshot API and MCP server from ScreenshotNeo. The API returns an image or PDF from a URL, and its options include full-page capture, custom waits, and output format selection. See the API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Create a free account and get 1,000 screenshots a month with no card.

9. Frequently asked questions

Can I get a URL’s file type without downloading it?

Yes. Use mimetypes.guess_type() for a suffix-based guess, or make a HEAD request to inspect the server’s declared type. Neither validates the resource’s bytes.

What is the difference between MIME type and encoding?

The MIME type describes the media format, while the separate encoding result from guess_type() can describe compression inferred from a suffix. For .tar.gz, keep both instead of treating gzip as the archive’s media type.

Should I trust the file extension or Content-Type?

For a live response, prefer the header as the server’s declaration and use the final URL’s suffix as a fallback. Trust neither as proof when your application needs to enforce a format; validate the actual bytes.

What should I return when the type is unknown?

Return None or a clearly named unknown state. Keeping uncertainty explicit prevents downstream code from mistaking an unsupported URL for a verified format.

Primary references