How to Convert Plain Text to HTML
Convert plain text to safe HTML, preserve line breaks, create structure, process Markdown, and avoid XSS with runnable Python and JavaScript examples.
Short answer: first decide whether the input should remain literal text or whether its conventions should become HTML. To display ordinary text safely, escape HTML-significant characters such as < and &, then place the result in a text element. Escaping does not infer headings, paragraphs, lists, or links. If the source is Markdown, use a Markdown parser; if it is untrusted, sanitize the parser output before rendering.
1. Choose the conversion you actually need
| Input and goal | Correct approach |
|---|---|
| Show user-entered prose exactly as written | Contextual HTML output encoding, then render as text |
| Show paragraphs and preserve intentional line breaks | Encode the text, then map lines or blank lines to an explicit layout |
| Create headings, lists, links, and emphasis from conventions | Use a parser for the source format, such as Markdown |
| Render untrusted Markdown | Parse, sanitize the resulting HTML, then render it |
HTML parsing gives special meaning to characters such as < and &. OWASP describes output encoding as converting untrusted input into a form displayed as data rather than executed as browser code (OWASP XSS Prevention Cheat Sheet).
2. Display plain text literally
Python
Python’s standard library provides html.escape(). Its default quote=True also encodes quotation marks, which is useful when the result may be reused in an attribute.
import html
plain_text = 'Use <tag> & "quotes"'
safe_text = html.escape(plain_text)
html_fragment = f'<p>{safe_text}</p>'
print(html_fragment)
# <p>Use <tag> & "quotes"</p>
Read the Python html module documentation for the standard-library behavior. Encode at the point where the value enters an HTML text context; do not permanently store an escaped copy as your canonical data.
Browser JavaScript
When inserting a string as text into an existing element, use textContent. It treats the value as text instead of parsing it as markup.
const input = 'Use <tag> & "quotes"';
const output = document.querySelector('#output');
output.textContent = input;
For this use case, MDN’s textContent reference documents the browser API. This does not make arbitrary attribute, URL, CSS, or JavaScript contexts safe; each context has different encoding rules.
Server-rendered HTML
If your template engine auto-escapes variables, pass the original plain string to the template and let the engine encode it. Avoid a raw or “safe HTML” output option unless the value has been deliberately sanitized and is intended to contain markup.
3. Preserve line breaks without creating unsafe markup
Escaping protects characters; it does not decide how newline characters should appear. Choose a presentation rule explicitly.
Option A: CSS preserves whitespace
<pre class="plain-text"></pre>
<style>
.plain-text {
white-space: pre-wrap;
overflow-wrap: anywhere;
}
</style>
<script>
document.querySelector('.plain-text').textContent = sourceText;
</script>
This is usually the simplest choice for logs, messages, configuration, and other content where spaces and newlines matter.
Option B: Convert blank lines to paragraphs
import html
source = "First paragraph.\n\nSecond paragraph with a\nline break."
paragraphs = []
for block in source.split("\n\n"):
text = html.escape(block)
paragraphs.append(f"<p>{text.replace(chr(10), '<br>')} </p>")
html_fragment = "".join(paragraphs)
print(html_fragment)
Remove the extra space before </p> in production if you use this illustrative snippet. For robust processing, normalize \r\n and \r to \n first, then define whether one newline means a line break and two newlines mean a new paragraph.
Option C: Build DOM nodes instead of concatenating HTML
function renderPlainText(text, container) {
container.replaceChildren();
const blocks = text.replaceAll('\r\n', '\n').replaceAll('\r', '\n').split(/\n{2,}/);
for (const block of blocks) {
const p = document.createElement('p');
p.textContent = block;
container.append(p);
}
}
This avoids constructing an HTML string. If you need single newlines inside each paragraph, split the block and append text nodes plus <br> elements.
4. Create semantic HTML structure
Plain text has no reliable metadata telling you which line is a heading, which lines form a list, or where a link target should come from. Define those rules yourself or use a source format with established syntax.
import html
heading = html.escape("Release notes")
items = ["Faster startup", "Improved error messages"]
list_items = "".join(f"<li>{html.escape(item)}</li>" for item in items)
result = f"<h1>{heading}</h1><ul>{list_items}</ul>"
print(result)
Every value in this example is encoded before insertion. A rule such as “the first line is an h1” is an application decision, not something HTML escaping can infer.
5. Convert Markdown to HTML
Use Markdown only when the input is actually Markdown and its syntax should become HTML. Python-Markdown’s markdown() or Markdown.convert() functions perform the conversion.
from markdown import markdown
source = "# Release notes\n\n- Faster startup\n- **Safer** rendering"
html_fragment = markdown(source)
print(html_fragment)
Install the package with python -m pip install markdown. See the Python-Markdown documentation for extensions and configuration.
Untrusted Markdown needs sanitization
Markdown conversion is not sanitization. Python-Markdown explicitly leaves responsibility for sanitizing generated HTML to the caller. For user-supplied Markdown, parse it, sanitize the resulting HTML with a policy appropriate to your application, and only then render it. Do not assume that escaping the original Markdown after parsing will make the generated HTML safe.
6. Convert a text file in a complete Python script
from pathlib import Path
import html
source_path = Path("input.txt")
output_path = Path("output.html")
text = source_path.read_text(encoding="utf-8")
text = text.replace("\r\n", "\n").replace("\r", "\n")
encoded = html.escape(text)
document = f"""<!doctype html>
<html lang="en">
<meta charset="utf-8">
<title>Converted text</title>
<style>body {{ white-space: pre-wrap; overflow-wrap: anywhere; }}</style>
<body>{encoded}</body>
</html>
"""
output_path.write_text(document, encoding="utf-8")
The generated file displays the text literally while preserving whitespace through CSS. If you need paragraphs or headings, replace the layout step with explicit rules rather than expecting html.escape() to create structure.
7. Security checklist
- Encode untrusted values for their exact destination: HTML text, an attribute, a URL, JavaScript, and CSS have different parsing rules.
- Prefer
textContentor a framework’s normal escaped interpolation for plain text. - Do not concatenate raw user input into an HTML string.
- Do not use HTML entity escaping as a universal sanitizer.
- Do not render Markdown output from untrusted users without sanitizing it.
- Escape exactly once for the destination context. Repeated escaping can display entity spellings such as
&. - Keep canonical data unescaped so it can later be rendered in a different context.
8. Common errors and fixes
| Symptom | Cause | Fix |
|---|---|---|
<tag> disappears or becomes an element |
Raw text was inserted as HTML | Use contextual escaping or textContent |
Literal < appears on screen |
The value was escaped twice | Keep the source unescaped and encode once at output |
| Newlines appear as spaces | HTML collapses ordinary whitespace | Use white-space: pre-wrap, <pre>, or explicit paragraphs and breaks |
| Markdown markers remain visible | Plain text was escaped instead of parsed | Run a Markdown parser when Markdown syntax is intended |
| Unexpected HTML or script executes | Parsed or supplied HTML was trusted | Sanitize generated HTML with an allowlist before rendering |
| Quotes break an attribute | Text-node encoding was used in an attribute context | Use the framework or encoder designed for that attribute context |
| Accented characters are garbled | Input or output encoding is inconsistent | Read and write UTF-8 and declare <meta charset="utf-8"> |
9. Performance, reliability, and cost
Escaping is linear in the input length and inexpensive for normal request sizes. A Markdown parser performs more work because it tokenizes and builds structure. For large files, stream or chunk the input when your output format permits it, and avoid repeatedly concatenating large strings in loops. Normalize line endings once, encode once, and write the output once.
For reliable conversion, make the input format explicit, fix the character encoding, define newline behavior, and add tests for ampersands, angle brackets, quotes, empty lines, Unicode, long lines, and already escaped-looking text. If Markdown is accepted, test both legitimate syntax and sanitization cases.
10. Or skip the browser setup
If your converted HTML is published at a URL and you need an image or PDF of it, ScreenshotNeo captures the page through one API request. Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server lets Claude, Cursor, and other MCP clients use take_screenshot, get_page_info, and capture_pdf.
See the ScreenshotNeo API documentation for all options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 screenshots each month with no card. Paid plans start at $5 for 3,000 shots, and every feature is included on every plan. Create a free ScreenshotNeo account.
11. FAQ
Does escaping plain text convert it into HTML?
It makes the text safe to place in an HTML text context. It does not create semantic elements such as headings, lists, or links.
Should I use <pre> for every text conversion?
No. Use it or white-space: pre-wrap when whitespace is meaningful. Use paragraphs when the content should read like prose.
Can I safely put escaped text in an attribute?
Use encoding designed for that attribute context and its surrounding syntax. Text-node escaping is not a universal solution.
Is Markdown safer than HTML?
Markdown can reduce the amount of HTML authors write, but the generated HTML still needs sanitization when the source is untrusted.
How do I keep the original text for later conversions?
Store the original Unicode text or Markdown. Encode or sanitize only when producing a specific output.


