ScreenshotNeo

BlogEngineering

UTF-8 Decoder: Encode and Decode UTF-8 Text

Learn how to encode Unicode text as UTF-8, decode bytes safely, handle invalid data and BOMs, and fix garbled characters.

By the ScreenshotNeo team1 October 20266 min read

UTF-8 is a byte encoding for Unicode text. Encoding converts Unicode scalar values into one to four bytes. Decoding converts valid UTF-8 bytes back into text. If the bytes were produced with another encoding, truncated, or malformed, a decoder may emit replacement characters such as � or fail in fatal mode.

This guide shows how to encode and decode UTF-8 in JavaScript, Python, Node.js and command-line workflows, explains BOM handling and invalid input, and provides a troubleshooting checklist.

How UTF-8 works

UTF-8 preserves ASCII: characters U+0000 through U+007F use their original single-byte values. Other Unicode scalar values use two, three or four bytes:

Code-point range Bytes Byte pattern
U+0000–U+007F 1 0xxxxxxx
U+0080–U+07FF 2 110xxxxx 10xxxxxx
U+0800–U+FFFF 3 1110xxxx 10xxxxxx 10xxxxxx
U+10000–U+10FFFF 4 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx

Continuation bytes must match 10xxxxxx, and UTF-8 must not encode the UTF-16 surrogate range directly. RFC 3629 defines these validity rules and warns that permissive decoding of malformed sequences can create security problems. Read RFC 3629.

Decode UTF-8 in JavaScript

In browsers and modern JavaScript runtimes, TextDecoder accepts bytes such as Uint8Array and returns a string.

const bytes = new Uint8Array([0x48, 0xC3, 0xA9, 0x6C, 0x6C, 0x6F]);
const text = new TextDecoder("utf-8").decode(bytes);
console.log(text); // Héllo

Reject malformed bytes

By default, TextDecoder uses replacement behavior. Pass { fatal: true } when corrupted input must be reported instead of silently changed.

const decoder = new TextDecoder("utf-8", { fatal: true });
try {
  const text = decoder.decode(new Uint8Array([0xC3, 0x28]));
  console.log(text);
} catch (error) {
  console.error("Invalid UTF-8:", error.message);
}

Decode a response body

const response = await fetch("https://example.com/data.txt");
const bytes = new Uint8Array(await response.arrayBuffer());
const text = new TextDecoder("utf-8", { fatal: true }).decode(bytes);
console.log(text);

Stream large responses

A multibyte character can be split between chunks. Keep the decoder state by passing stream: true for intermediate chunks, then flush it with a final decode.

const decoder = new TextDecoder("utf-8", { fatal: true });
let output = "";
for await (const chunk of response.body) {
  output += decoder.decode(chunk, { stream: true });
}
output += decoder.decode(); // flush pending bytes
console.log(output);

The WHATWG Encoding Standard specifies TextDecoder, replacement and fatal modes, and streaming behavior. See the standard.

Encode UTF-8 in JavaScript

const text = "こんにちは — café";
const bytes = new TextEncoder().encode(text);
console.log([...bytes]);

// Convert bytes to a hexadecimal diagnostic string
const hex = [...bytes].map(byte => byte.toString(16).padStart(2, "0")).join(" ");
console.log(hex);

TextEncoder always emits UTF-8. It does not accept an alternative character-set label.

Encode and decode UTF-8 in Python

Encode a string

text = "こんにちは — café"
data = text.encode("utf-8")
print(data)

with open("message.txt", "wb") as file:
    file.write(data)

Decode bytes

data = bytes([0x48, 0xC3, 0xA9, 0x6C, 0x6C, 0x6F])
text = data.decode("utf-8")
print(text)  # Héllo

Choose an error policy

raw = b"valid: \xc3\xa9 invalid: \xc3("

# Fail immediately (recommended when data must be correct)
try:
    print(raw.decode("utf-8", errors="strict"))
except UnicodeDecodeError as error:
    print("Invalid UTF-8:", error)

# Replace invalid sequences with U+FFFD
print(raw.decode("utf-8", errors="replace"))

# Drop invalid bytes (use only when loss is acceptable)
print(raw.decode("utf-8", errors="ignore"))

Read and write files explicitly

with open("input.txt", "r", encoding="utf-8", errors="strict") as file:
    text = file.read()

with open("output.txt", "w", encoding="utf-8", newline="") as file:
    file.write(text)

Node.js UTF-8 APIs

import { readFile, writeFile } from "node:fs/promises";

const bytes = await readFile("input.txt");
const text = bytes.toString("utf8");
console.log(text);

await writeFile("output.txt", text, { encoding: "utf8" });

For strict validation in Node.js, use TextDecoder with fatal: true, as shown in the JavaScript section. Buffer.toString("utf8") is convenient but uses replacement behavior for malformed input.

Decode UTF-8 from the command line

Fetch bytes with cURL, then decode them with a local tool whose input encoding is known:

curl --fail --location https://example.com/data.txt --output data.bin
iconv -f UTF-8 -t UTF-8 data.bin > data.txt

To inspect bytes before decoding:

xxd -g 1 data.bin | head

Do not assume arbitrary bytes are UTF-8. Confirm the producer’s declared encoding, protocol specification or file metadata first.

UTF-8 BOM: EF BB BF

A UTF-8 byte-order mark is the three-byte sequence EF BB BF. UTF-8 has no big-endian or little-endian byte order; the mark is an encoding signature, not an endianness indicator. It can break formats that require the first bytes to be an ASCII token, such as a shebang.

The WHATWG UTF-8 decode algorithm consumes an initial BOM. A decode-without-BOM operation preserves it as U+FEFF. Check which operation your library uses before comparing the first character or parsing a file. Unicode’s BOM FAQ explains the distinction.

const bytes = new Uint8Array([0xEF, 0xBB, 0xBF, 0x41]);
console.log(new TextDecoder("utf-8").decode(bytes)); // A
console.log(new TextDecoder("utf-8", { ignoreBOM: true }).decode(bytes)); // may expose U+FEFF

Invalid UTF-8 and replacement characters

The symbol � is U+FFFD, the replacement character. It usually means a decoder encountered invalid bytes and chose replacement behavior. Common causes include:

  • Bytes were encoded as Windows-1252, Latin-1 or another charset.
  • A multibyte sequence was truncated during a read or network transfer.
  • Text was decoded once and then incorrectly decoded again.
  • A permissive conversion replaced errors before your code received the data.

Use fatal or strict mode at trust boundaries, log the original bytes, and fix the producer or transport. Never “repair” data by guessing an encoding when the protocol defines one.

Encoding and decoding checklist

  1. Keep text as Unicode strings internally.
  2. Encode only when writing bytes to a file, socket, database or HTTP body.
  3. Decode bytes exactly once, using the producer’s declared charset.
  4. Use strict or fatal errors when silent corruption is unacceptable.
  5. Handle streaming boundaries with a stateful decoder.
  6. Check whether your API consumes or preserves a leading BOM.
  7. Reject overlong sequences and surrogate encodings.

Troubleshooting

Symptom Likely cause Fix
� appears in output Malformed bytes or wrong source encoding Verify the producer charset; decode with strict/fatal mode.
Text looks like é UTF-8 bytes decoded as Latin-1/Windows-1252 Keep the original bytes and decode them as UTF-8 once.
First character is invisible or comparisons fail Leading BOM preserved Use BOM-consuming decode or remove U+FEFF deliberately.
Only the end of a streamed character is broken Chunks split a multibyte sequence Use streaming decoder state and flush at end.
Decoder throws immediately Fatal/strict mode found invalid input Inspect the byte offset and correct the upstream encoding or truncation.
Security filter behaves unexpectedly Overlong or surrogate sequence accepted by a permissive parser Use a standards-conforming UTF-8 decoder and reject malformed input.

Performance, reliability and cost

UTF-8 conversion is linear in the number of input bytes. Avoid repeated string-to-byte conversions in tight loops; encode once and reuse the resulting buffer. Stream large files instead of loading them all into memory, while preserving decoder state across chunks. Strict validation adds error checking but prevents corrupted data from entering storage or downstream parsers. For network payloads, verify transfer completion and content metadata before decoding.

Or skip the browser setup

If your goal is to obtain a clean screenshot of a page rather than process text bytes, ScreenshotNeo provides a single HTTP request. Its API accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture. Bot checks, blank pages, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. An MCP server lets Claude, Cursor and other MCP clients call screenshot tools directly. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Free accounts include 1,000 screenshots each month with no card. Paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.

FAQ

Is UTF-8 the same as Unicode?

No. Unicode defines scalar values and characters; UTF-8 is one byte representation of those values.

Can every byte sequence be decoded as UTF-8?

No. Continuation-byte, range and surrogate rules make many sequences invalid.

Should I remove every BOM?

Only when the format or parser requires no leading signature. Otherwise follow that format’s BOM rules.

When should decoding be fatal?

Use fatal or strict behavior for protocols, security-sensitive parsing and data pipelines where replacement would hide corruption.

Why does ASCII work when UTF-8 is broken?

ASCII bytes are identical in UTF-8, so encoding mismatches often remain hidden until non-ASCII characters appear.

For new protocols and formats, use the utf-8 label and define the encoding at the boundary. The WHATWG standard describes UTF-8 as the required encoding for new formats and protocols. Read the Encoding Standard.