ScreenshotNeo

BlogHTML to image & PDF

How to Fix Hash Characters Appearing in Converted PDFs

Hashes in a PDF usually indicate missing glyphs, encoding errors, or bad OCR. Find the failure point and repair fonts, Unicode, conversion, or scans.

By the ScreenshotNeo team1 October 20267 min read

Short answer: hash characters usually appear because the PDF renderer cannot map a source character to a glyph in the selected font. They can also result from non-Unicode source text, unsupported characters, encoding mistakes, font substitution, or poor OCR. First determine whether the hashes are already in the source, visible in the PDF, or introduced only when text is copied or extracted. Then repair the source text, choose a font with the required glyphs, embed it when licensing permits, or use OCR only for image scans.

Use this diagnostic order

  1. Open the original document and search for the affected characters.
  2. Open the PDF visually. Note whether hashes are visible on the page.
  3. Select and copy the affected text into a plain-text editor.
  4. Check whether the PDF page contains selectable text or only a scanned image.
  5. Record the source application, converter and version, font names, language or script, and whether the problem affects display, copied text, or both.
What you observe Most likely area to investigate
Hashes are already in the source Source content or input encoding
Source is correct, PDF visibly shows hashes Font glyph coverage, substitution, embedding, or renderer
PDF looks correct, copied text contains hashes Character mapping, encoding, or text extraction
Page is an image scan Scan quality and OCR workflow

1. Check the source text and encoding

Repair the original document whenever possible, then export again. Preserve text as Unicode and avoid legacy or non-Unicode fonts. Amazon Kindle Direct Publishing documents unsupported characters, non-Unicode fonts, and Unicode encoding errors as causes of conversion failures; treat that guidance as specific to its workflow, but the same checks are useful in other converters.

Test with a small file containing the exact failing characters before converting a long document:

English: Hello
Cyrillic: Привет
Greek: Καλημέρα
Chinese: 你好
Japanese: こんにちは
Arabic: مرحبا
Symbols: € ™ ✓ — 漢字

If the sample is wrong before conversion, fix the application’s encoding or replace the source text. If only one script fails, continue with font coverage checks.

2. Choose a font that contains the missing glyphs

A font can support Latin text while lacking Cyrillic, Chinese, Japanese, Arabic, emoji, or specialist symbols. A documented Better PDF Exporter for Jira case explains that when a glyph is missing, its renderer substitutes #. This is a strong clue when only particular scripts are affected, but it is not a universal diagnosis for every PDF converter.

  • Identify the exact characters that become hashes.
  • Open the font’s character map or specimen and confirm those characters exist.
  • Select a font with coverage for the whole language, not just a few sample letters.
  • Use an automatic or fallback-font option when the converter provides one and the document contains many languages.
  • Export the short multilingual test file again.

If the converter supports font fallback, configure a fallback chain so a missing glyph is taken from a permitted secondary font. Keep fallback fonts consistent across build machines to avoid different output.

3. Embed fonts when the license permits

Embedding places font data in the PDF so a reader does not need the same font installed. Adobe’s font documentation explains that embedding can prevent font substitution, but font vendors can restrict embedding and embedding cannot add glyphs that the font does not contain.

  1. Check the font license or embedding permission.
  2. Enable full embedding or the converter’s permitted subset-embedding mode.
  3. Inspect the generated PDF’s font properties to confirm the intended font is embedded.
  4. Open the PDF on a machine without the source font and check the affected script.

Do not treat an “embed fonts” checkbox as a complete fix. A font may lack the glyph, the license may forbid embedding, or the converter may still write an incorrect character map.

4. Distinguish text conversion from OCR

OCR is for recognizing text in page images. It is not a general repair for a PDF that already contains renderable text. Adobe documents an Acrobat OCR error when recognition is attempted on a page that already has renderable text. If your PDF has selectable text, repair the source, font, encoding, or converter instead.

For a genuine scan:

  • Rescan pages straight, clean, and at a sufficient resolution.
  • Remove skew, smudges, and marks before recognition.
  • Set the OCR language to match the document.
  • Review names, numbers, punctuation, and non-Latin scripts manually.
  • Export a new searchable PDF and verify by searching and copying text.

Amazon KDP warns that PDF-to-Word files produced by some OCR tools can contain empty boxes or unrecognizable characters. Follow the destination platform’s conversion guidance rather than repeatedly OCRing a text-based PDF.

5. Re-export through the intended conversion path

After fixing the font or source, regenerate the PDF using the converter’s direct export route. In an Acrobat Word workflow, Adobe recommends converting the document to PDF instead of relying on Print to PDF or Scan to PDF when conversion quality is poor. Other applications have their own preferred export paths.

Keep a reproducible test case with the failing characters. Compare the new file visually, through search, and through copy/paste. A PDF that opens successfully can still contain incorrect text mappings.

Command-line checks for a PDF

These tools help determine whether text is present and which fonts are recorded. They do not repair a broken PDF by themselves.

# List fonts and embedding status (Poppler)
pdffonts input.pdf

# Extract text for a copy/search check (Poppler)
pdftotext -layout input.pdf extracted.txt

# Look for replacement hashes in extracted text
rg -n '#' extracted.txt

If pdffonts shows a substituted font or a font that is not embedded, revisit the export settings. If pdftotext produces hashes while the page looks correct, investigate the PDF’s character map and the extraction tool.

Python inspection example

This script reports extracted text and helps separate an image-only page from a text page. Install the dependency with python -m pip install pypdf.

from pathlib import Path
from pypdf import PdfReader

path = Path("input.pdf")
reader = PdfReader(str(path))

for number, page in enumerate(reader.pages, start=1):
    text = page.extract_text() or ""
    print(f"Page {number}: {len(text)} extracted characters")
    if "#" in text:
        print("  Hash characters found in extracted text")
    if not text.strip():
        print("  No selectable text; investigate scan/OCR workflow")

Troubleshooting common symptoms

Symptom Cause to test Fix
Only Chinese, Japanese, or Korean text becomes # Selected font lacks CJK glyphs Use a CJK-capable font, configure fallback, and export again
Cyrillic or Greek fails while Latin works Partial font coverage or legacy encoding Use Unicode source text and a font covering the script
PDF looks correct but copied text is wrong Bad ToUnicode map or extraction conversion Regenerate with a converter that writes correct mappings; compare another extractor
Output changes on another computer Font substitution because fonts are not embedded Embed a licensed font and verify PDF properties
OCR creates boxes or nonsense Wrong OCR language, low-quality scan, or unsuitable OCR route Improve the scan, set language correctly, and review the recognized text
Hashes are present in the source file Input data already contains replacement characters Repair the source data and its encoding before conversion
Only emoji or specialist symbols fail Font lacks those code points or the renderer cannot draw them Use a font with coverage and test whether the converter supports the symbols

Performance, reliability, and cost considerations

  • Performance: font fallback and embedding increase output work and file size. Test a representative multilingual page rather than judging from a Latin-only sample.
  • Reliability: pin converter versions and fonts in automated builds. Store the source, export settings, and a small regression document containing affected scripts.
  • Verification: check visual rendering plus search and copy/paste. These validate different parts of the PDF pipeline.
  • Cost: repairing the source and re-exporting is usually cheaper than manually correcting a large PDF. OCR adds processing time and review effort, so use it only for image text.
  • Licensing: confirm that fonts may be embedded before distributing the PDF.

Or skip the browser setup

If your workflow starts with a web page that must be captured as an image or PDF, ScreenshotNeo provides a single HTTP request. Its capture pipeline accepts cookie and consent banners before the shot and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response reports the result with X-Page-Verdict and X-Billed headers.

See the ScreenshotNeo API documentation for PDF options, custom CSS and JavaScript, waits, headers, cookies, device presets, and other settings.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also has an MCP server for AI agents, including Claude, Cursor, and other MCP clients, with take_screenshot, get_page_info, and capture_pdf tools. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Does a valid PDF guarantee that its text is correct?

No. A PDF can open and print successfully while displaying substituted glyphs or containing an incorrect text map.

Will embedding any font solve the problem?

No. The font must contain the glyph, and its license must allow embedding. Encoding and renderer bugs can remain.

Should I convert every broken PDF to images and OCR it?

No. OCR is intended for scanned page images. Converting selectable text to images can reduce searchability and introduce recognition errors.

Why do hashes appear only after copy and paste?

The visible glyph may render correctly while the PDF’s character mapping for extraction is wrong. Re-export with correct Unicode mappings and compare extraction tools.

What information should I send to a converter vendor?

Provide the smallest source file that reproduces the issue, affected characters, source application, converter and version, font names and licenses, whether fonts are embedded, and whether hashes are visible or extraction-only.