How to Decode URLs in Python
Use Python’s urllib.parse functions to decode percent-encoded URL components, form values, query strings, or bytes—and choose the right function for each.
Use urllib.parse.unquote() to decode a percent-encoded URL component as text. Use unquote_plus() for form-style values where + means a space. To extract fields from a complete query string, use parse_qs() or parse_qsl(); use unquote_to_bytes() when you need bytes.
Python’s standard library provides these functions in urllib.parse. Decoding changes representation; it does not validate a URL or make untrusted input safe. See the official urllib.parse documentation.
1. Pick the right decoding function
| Your input or goal | Use | What it does |
|---|---|---|
| A percent-encoded component, such as a path segment, as text | unquote() |
Replaces percent escapes such as %20. A plus sign stays a plus. |
| A form-encoded value | unquote_plus() |
Decodes percent escapes and changes + to a space. |
| A query string whose named fields you need | parse_qs() |
Returns a mapping from field names to lists of values. |
| A query string where order or repeated pairs matter | parse_qsl() |
Returns a list of name/value pairs in input order. |
| Percent-encoded content that must remain bytes | unquote_to_bytes() |
Returns bytes, not decoded text. |
Use a component decoder on a component, not indiscriminately on an entire URL. If you need query parameters, parse the query string so separators and repeated values retain their intended structure.
2. Decode a component with unquote()
unquote() replaces %xx escapes and decodes the resulting bytes to text. Its defaults are UTF-8 and errors="replace".
from urllib.parse import unquote
encoded_path = "/products/El%20Ni%C3%B1o"
decoded_path = unquote(encoded_path)
print(decoded_path)
# /products/El Niño
# A plus sign is ordinary data to unquote().
print(unquote("Ada+Lovelace"))
# Ada+Lovelace
For a str input, non-ASCII percent-encoded bytes are interpreted using the selected text encoding. The documentation’s default is UTF-8. If you have a legacy source encoded differently, pass its known encoding explicitly; do not guess from the decoded output.
from urllib.parse import unquote
value = unquote("caf%E9", encoding="latin-1")
print(value)
# café
The errors argument controls how invalid byte sequences are handled during text conversion. The default, replace, inserts the replacement character. You can select a stricter behavior such as strict to raise an error when the byte sequence cannot be decoded with the chosen encoding.
from urllib.parse import unquote
try:
print(unquote("%FF", encoding="utf-8", errors="strict"))
except UnicodeDecodeError as exc:
print(f"Invalid UTF-8 escape sequence: {exc}")
3. Decode form values with unquote_plus()
HTML form-style encoding uses + to represent a space. unquote_plus() applies that rule as well as decoding percent escapes. It accepts a string.
from urllib.parse import unquote_plus
print(unquote_plus("name=Ada+Lovelace"))
# name=Ada Lovelace
print(unquote_plus("C%2B%2B+guide"))
# C++ guide
Choose this only when the input uses form-style encoding. If a plus sign is literal component data, unquote_plus() changes its meaning; use unquote() instead.
4. Parse a complete query string
When you need parameter names and values, use parse_qs() or parse_qsl() rather than splitting on & and decoding by hand. Both parse the form-style query representation, including plus-as-space behavior.
Use parse_qs() for a mapping
from urllib.parse import parse_qs
query = "name=Ada+Lovelace&tag=python&tag=urls"
params = parse_qs(query)
print(params)
# {'name': ['Ada Lovelace'], 'tag': ['python', 'urls']}
print(params["tag"])
# ['python', 'urls']
Values are lists because a query can contain the same key more than once. If your application expects one value, decide explicitly whether to reject duplicates, take the first, or handle all of them.
Use parse_qsl() when pairs and order matter
from urllib.parse import parse_qsl
query = "tag=python&tag=urls&sort=recent"
pairs = parse_qsl(query)
print(pairs)
# [('tag', 'python'), ('tag', 'urls'), ('sort', 'recent')]
For an entire URL, first split it into components, then parse its query. This avoids treating the path or fragment as query data.
from urllib.parse import urlsplit, parse_qs
url = "https://example.com/search?q=red+fox&tag=python#results"
parts = urlsplit(url)
query_params = parse_qs(parts.query)
print(parts.path)
# /search
print(query_params)
# {'q': ['red fox'], 'tag': ['python']}
print(parts.fragment)
# results
Python’s URL parsing functions split and interpret components; they do not validate that a URL is well-formed or safe for your application. Validate hosts, schemes, allowed characters, and any other application-specific rules before trusting input.
5. Get bytes with unquote_to_bytes()
Use unquote_to_bytes() if the decoded result is binary data or if another layer must decide how to interpret the bytes. It returns bytes, avoiding an automatic text-decoding choice.
from urllib.parse import unquote_to_bytes
raw = unquote_to_bytes("payload=%E2%9C%93")
print(raw)
# b'payload=\xe2\x9c\x93'
text = raw.decode("utf-8")
print(text)
# payload=✓
When its input is a string, unescaped non-ASCII characters are encoded as UTF-8 in the returned bytes. If you need to preserve the original wire bytes exactly, retain the original byte input rather than first converting it to text.
6. Runnable command-line example
This small script accepts a component or query string and applies the selected operation. Save it as decode_url.py and run it with Python 3:
from urllib.parse import unquote, unquote_plus, parse_qs, parse_qsl
component = "/El%20Ni%C3%B1o/C%2B%2B"
form_value = "name=Ada+Lovelace"
query = "tag=python&tag=urls&q=red+fox"
print("component:", unquote(component))
print("form value:", unquote_plus(form_value))
print("query mapping:", parse_qs(query))
print("ordered pairs:", parse_qsl(query))
python3 decode_url.py
Expected output:
component: /El Niño/C++
form value: name=Ada Lovelace
query mapping: {'tag': ['python', 'urls'], 'q': ['red fox']}
ordered pairs: [('tag', 'python'), ('tag', 'urls'), ('q', 'red fox')]
7. Common mistakes and troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| A plus sign became a space | unquote_plus() or query parsing was used on data where plus was literal. |
Use unquote() for ordinary components, or ensure literal plus is encoded as %2B in form data. |
| A space remained a plus sign | unquote() was used for a form-encoded value. |
Use unquote_plus() or parse the query with parse_qs()/parse_qsl(). |
| Non-ASCII text looks wrong | The bytes were decoded using the wrong character encoding. | Confirm the encoding used by the producer; pass it explicitly or decode bytes with the known encoding. |
| Invalid characters appear as � | The default errors="replace" replaced a byte sequence invalid for the selected encoding. |
Use errors="strict" when invalid data should fail, then handle UnicodeDecodeError. |
| A parameter appears more than once | Query strings may contain repeated keys; parse_qs() preserves values in a list. |
Handle the list deliberately or use parse_qsl() to process each pair. |
| Decoded output still contains percent escapes | The source may be encoded more than once. | Only decode again if the format contract explicitly says it is double-encoded. Repeated decoding can turn intentionally escaped data into active delimiters or other changed content. |
TypeError when passing bytes to unquote_plus() |
unquote_plus() expects a string. |
Keep form data as text or use a byte-aware approach appropriate to the input; unquote_to_bytes() returns bytes. |
Version note: the current Python 3.14 documentation says unquote() accepted only strings before Python 3.9; bytes support was added in 3.9. Check the documentation for the Python version you deploy if you rely on version-specific behavior.
8. Security, performance, reliability, and cost
Security
Decoding is not validation. A decoded path can contain separators or characters that were hidden in encoded form, so apply authorization and path rules to the representation your application actually uses. Avoid decoding repeatedly unless the protocol requires it, and do not treat successful parsing as proof that a URL is safe.
Performance
For ordinary application inputs, choose the function that matches the data shape and decode once at the boundary where the encoding is defined. Avoid hand-written replacement loops: they are easy to get wrong around UTF-8, repeated parameters, plus signs, and malformed escapes. If processing very large or adversarial inputs, enforce input-size limits in your application.
Reliability
Be explicit about the text encoding and error policy when malformed or legacy data is possible. Test cases should include spaces, literal plus signs, non-ASCII characters, repeated query keys, empty values, and malformed percent escapes. Keep raw input when you need auditability or need to re-interpret it later.
Cost
urllib.parse is part of Python’s standard library, so this decoding step does not require a third-party package or a ScreenshotNeo API call. If the actual task is capturing a web page as an image or PDF, that is a separate operation.
9. Or skip the browser setup
URL decoding is a local Python task. If what you need next is a clean capture of a page, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for options.
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Sign up free for 1,000 screenshots a month, with no card required.
10. FAQ
Does unquote() decode a complete URL?
It replaces percent escapes in the supplied string, but it does not parse the URL into scheme, host, path, query, and fragment. Use urlsplit() for components and query parsing functions for query fields.
How do I keep duplicate query parameters?
Use parse_qsl() to receive an ordered list of pairs, or keep the lists returned by parse_qs().
Does URL decoding make an input safe?
No. Validate the decoded components against the rules of your application before using them.
Should I decode an input twice?
Only when the format explicitly requires multiple encoding layers. Otherwise a second pass can change data that was intentionally escaped.


