How to Fix Selectable Text in PDFs Generated with jsPDF
Learn why jsPDF PDFs contain unselectable screenshots and how to generate searchable text with direct APIs, HTML rendering, and embedded fonts.
Direct answer: your PDF text is not selectable because the page contains an image, not PDF text operators. A common html2canvas plus jsPDF.addImage() pipeline rasterizes the page into a canvas and places that bitmap in the PDF. Keep written content on a text-producing path: use doc.text() for known data, or use doc.html() and verify the generated file. Keep screenshots, charts, and other visual regions as images only when that trade-off is intentional.
1. Identify the image-only PDF
Open the generated PDF and try all three checks:
- Search for a word that is visibly present.
- Drag across a sentence and copy it.
- Extract text with a PDF viewer or parser.
If none works, inspect your source for html2canvas, canvas, toDataURL(), or addImage(). Those calls usually indicate that the page is being rendered as pixels. The html2pdf.js documentation explicitly describes this mode as an image placed in a PDF, which makes text non-selectable and can increase file size. html2canvas reconstructs a visual representation from the DOM; it is not a semantic PDF text writer.
2. The failing pattern
<script src="https://cdn.jsdelivr.net/npm/html2canvas@latest/dist/html2canvas.min.js"></script>
<script src="https://cdn.jsdelivr.net/npm/jspdf@latest/dist/jspdf.umd.min.js"></script>
<script>
html2canvas(document.querySelector('#content')).then(canvas => {
const { jsPDF } = window.jspdf;
const pdf = new jsPDF();
pdf.addImage(canvas.toDataURL('image/png'), 'PNG', 0, 0, 210, 297);
pdf.save('file.pdf');
});
</script>
This can look perfect while producing a one-page bitmap. Changing PNG to JPEG, increasing canvas scale, or changing the PDF dimensions does not create selectable text; it only changes the image.
3. Fix known content with jsPDF text APIs
When the document is made from data you control, write the data directly to the PDF.
<script src="https://cdn.jsdelivr.net/npm/jspdf@latest/dist/jspdf.umd.min.js"></script>
<script>
const { jsPDF } = window.jspdf;
const doc = new jsPDF({ unit: 'mm', format: 'a4' });
doc.setFont('helvetica', 'normal');
doc.setFontSize(12);
doc.text('Selectable PDF text', 20, 30);
doc.setFontSize(10);
doc.text('Order: 1042', 20, 40);
doc.text('Status: Paid', 20, 47);
doc.save('selectable.pdf');
</script>
doc.text() creates text content that viewers can search and select. Set the font and size before writing each group of text. For wrapping, measure or split lines yourself and advance the Y coordinate between lines.
Simple wrapping and page breaks
const { jsPDF } = window.jspdf;
const doc = new jsPDF({ unit: 'mm', format: 'a4' });
const margin = 20;
const pageWidth = doc.internal.pageSize.getWidth();
const pageHeight = doc.internal.pageSize.getHeight();
const maxWidth = pageWidth - margin * 2;
const lineHeight = 6;
let y = 25;
const lines = doc.splitTextToSize(
'A long paragraph remains real PDF text when it is split into lines and written with doc.text.',
maxWidth
);
for (const line of lines) {
if (y > pageHeight - margin) {
doc.addPage();
y = margin;
}
doc.text(line, margin, y);
y += lineHeight;
}
doc.save('wrapped-selectable.pdf');
4. Render HTML with the jsPDF HTML module
For an existing HTML layout, try the official HTML module instead of converting the entire element to a canvas and calling addImage().
<script src="https://cdn.jsdelivr.net/npm/jspdf@latest/dist/jspdf.umd.min.js"></script>
<script>
const { jsPDF } = window.jspdf;
const doc = new jsPDF({ unit: 'mm', format: 'a4' });
doc.html(document.querySelector('#content'), {
x: 15,
y: 15,
width: 180,
autoPaging: 'text',
callback: pdf => pdf.save('html-selectable.pdf')
});
</script>
autoPaging: 'text' asks the module to avoid cutting text in half at page boundaries. The HTML module also exposes options for html2canvas, jsPDF, and fontFaces. For HTML strings, the jsPDF documentation notes that the pipeline also depends on DOM sanitization.
What this approach can and cannot preserve
| Requirement | Best approach | Trade-off |
|---|---|---|
| Searchable body copy | doc.text() or verified doc.html() |
More layout work |
| Exact complex CSS appearance | Canvas screenshot plus addImage() |
Text is an image |
| Charts and screenshots | Embed those regions as images | Those regions are not selectable |
| Predictable pagination | Direct text with explicit page-break logic | You implement layout rules |
A hybrid PDF is often the practical answer: write headings, paragraphs, and metadata as text, then embed charts or screenshots as images.
5. Embed a TrueType font for Unicode text
jsPDF’s standard fonts cover a limited code page. For accented Latin text, Greek, Cyrillic, Chinese, or other Unicode characters, register a TrueType font that contains every glyph you need.
const { jsPDF } = window.jspdf;
const doc = new jsPDF();
const fontBinary = await fetch('/fonts/Inter-Regular.ttf')
.then(response => {
if (!response.ok) throw new Error(`Font request failed: ${response.status}`);
return response.arrayBuffer();
})
.then(buffer => String.fromCharCode(...new Uint8Array(buffer)));
doc.addFileToVFS('Inter-Regular.ttf', fontBinary);
doc.addFont('Inter-Regular.ttf', 'Inter', 'normal');
doc.setFont('Inter', 'normal');
doc.setFontSize(12);
doc.text('Unicode: café, Ελληνικά, 中文', 20, 40);
doc.save('unicode.pdf');
Register the font before writing text. If characters are blank or replaced, confirm that the selected TTF actually contains those glyphs and that the font file was loaded successfully.
6. Verify the generated file in code
A successful doc.save() call only proves that a file was written. It does not prove that the file contains text. Add a manual verification step to your workflow:
- Generate a PDF containing a unique test word.
- Open it in two different PDF viewers.
- Search, select, and copy the word.
- Run your normal text-extraction tool and confirm the expected characters.
- Repeat with representative Unicode characters and a multi-page document.
If a direct doc.text('test', 20, 20) page works but the HTML version does not, isolate the HTML until you find an unsupported CSS property, canvas element, or cross-origin asset.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Nothing can be selected | The page ends in addImage() |
Write text with doc.text() or test doc.html(). |
| Text is searchable but page breaks are awkward | Automatic pagination cannot infer your layout | Use autoPaging: 'text', reduce content width, or add explicit page-break logic. |
| Accents or non-Latin characters are missing | The standard font lacks glyphs | Embed a suitable TTF with addFileToVFS() and addFont(). |
| HTML styling disappears | The HTML renderer does not implement that CSS property | Simplify the markup, use supported CSS, or reserve that region for an image. |
| Images are blank or tainted | Cross-origin restrictions or inaccessible assets | Serve assets with appropriate CORS headers, use same-origin files, and check the browser console. |
| Output is huge | A high-resolution full-page bitmap is embedded | Keep body text as text, crop visual assets, and choose a suitable image format and scale. |
| PDF looks right but extraction is empty | Visual fidelity was achieved through canvas rendering | Inspect the source for canvas, toDataURL, and addImage; replace the text path. |
8. Performance, reliability, and cost
- Performance: Direct text is usually cheaper to construct and keeps files compact. Full-page canvas rendering consumes browser memory proportional to pixel dimensions, especially at high scale.
- Reliability: Direct text has fewer browser rendering variables. HTML output still depends on the supported CSS subset, loaded fonts, image access, and pagination behavior.
- Fonts: Embedding a font increases the PDF but avoids missing glyphs. Load it once and reuse the same registered font.
- Cost: jsPDF itself is a client-side library. Your costs come from the browser, hosting, storage, and any external assets or services you add.
- Accessibility: Real text gives assistive technologies and extraction tools content to work with; an image-only page does not provide equivalent semantics.
9. Or skip the browser setup
If your goal is a reliable screenshot or PDF of a web page rather than a semantically generated report, ScreenshotNeo provides a single API request. It removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and each response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server lets Claude, Cursor, and other MCP clients use take_screenshot, get_page_info, and capture_pdf.
For a PDF capture, see the ScreenshotNeo documentation. The same endpoint also supports PNG, JPEG, and WebP screenshots.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
require('node:fs').writeFileSync('shot.webp', bytes);
ScreenshotNeo includes full-page capture with lazy images loaded, element capture by CSS selector, custom CSS and JavaScript, waits, headers, cookies, user agents, timezone and geolocation, PDF paper and page-range controls, caching, signed links, asynchronous jobs, bulk capture, and a usage API. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
10. FAQ
Can OCR make an image-only jsPDF searchable?
OCR can add a text layer, but it is a separate recognition step. If you control the source content, generating text directly is simpler and preserves the original characters.
Does changing PNG to JPEG make text selectable?
No. Both formats remain images when passed to addImage().
Can I make every CSS property selectable?
No. Selectability comes from the PDF text layer. Complex visual styling may still need images, while the written content should use text APIs.
Why does doc.html() still produce imperfect output?
It is an HTML rendering pipeline with supported-feature and asset limitations. Test the actual PDF, simplify unsupported markup, and use direct text when exact semantics matter.
Is an image-only PDF always wrong?
No. It can be appropriate for a faithful visual snapshot. It is the wrong format when readers must search, copy, extract, or access the text.


