ScreenshotNeo

BlogHTML to image & PDF

How to Create Accessible PDFs with DocRaptor

Create tagged PDFs with DocRaptor’s PDF/UA-1 profile, then inspect structure, reading order, contrast, and assistive technology behavior before publishing.

By the ScreenshotNeo team4 October 20267 min read

To request a tagged PDF from DocRaptor, set prince_options.profile to PDF/UA-1. Then inspect and test the generated PDF: a successful conversion or profile setting alone does not establish that the document is usable with assistive technology or conforms to every WCAG criterion or legal requirement.

The workflow is: author semantic HTML, request PDF/UA-1 output, inspect the tag tree and reading order, validate and test the result, and repair any defects before publishing.

1. Author accessible HTML before conversion

The PDF inherits structure and content from its HTML source. Make the source express the document’s actual meaning and reading sequence.

  • Use semantic elements and a logical heading hierarchy. DocRaptor says Prince maps common elements such as headings and block quotes to PDF tags.
  • Set the document language on the root element, for example <html lang="en">. Mark passages in another language with an appropriate lang attribute.
  • Give informative images useful alternative text. Ensure decorative images can be ignored by assistive technology.
  • Use table headers and captions for data tables. Do not rely on visual placement alone to communicate relationships.
  • Do not convey meaning through color alone. Make links distinguishable by more than color.
  • Use sufficient text contrast. WCAG 2.2 SC 1.4.3 specifies at least 4.5:1 for ordinary text and 3:1 for large-scale text, subject to exceptions including inactive components, decoration, and logos.
  • Set a meaningful HTML <title>. Consider adding author and subject metadata as appropriate.
  • Keep source order aligned with the intended reading order. CSS positioning can change visual placement without changing the logical sequence.

These practices improve the source, but still inspect the converted PDF: HTML semantics do not guarantee correct tags for every layout or content type.

2. Request PDF/UA-1 from DocRaptor

DocRaptor’s documentation specifies the PDF/UA-1 profile through prince_options[profile]. In its JSON request body, that becomes a nested prince_options object. Check the API reference for pipeline support applicable to your account; the reference lists PDF/UA-1 on pipeline 7 and later.

{
  "type": "pdf",
  "document_content": "<html lang=\"en\"><head><title>Accessible report</title></head><body><h1>Accessible report</h1><p>Report text.</p></body></html>",
  "prince_options": {
    "profile": "PDF/UA-1"
  }
}

Use that structure in the JSON request your DocRaptor integration sends. For production, use the authentication and endpoint pattern in the current DocRaptor API reference; do not expose API credentials in browser code or public source files.

DocRaptor accepts HTML content for PDF generation. JavaScript is disabled by default; enable the javascript option only when the source requires client-side rendering, and account for the extra rendering behavior when debugging output.

Sources: DocRaptor’s accessible and tagged PDFs guide, API reference, and HTML and JavaScript tutorial.

3. Inspect the generated PDF

Treat conversion success as a file-generation result, not an accessibility pass. Review both the PDF’s structure and how people encounter its content.

  1. Inspect the tag tree. Confirm that headings, paragraphs, lists, table headers, captions, image alternatives, abbreviations, quotes, footnotes, and bibliography entries have appropriate structure.
  2. Review bookmarks. DocRaptor creates bookmarks from headings by default, but confirm that they follow the document outline and help readers navigate.
  3. Check reading order. Compare the tag sequence with the visual meaning, especially in columns, sidebars, positioned content, and complex tables. Tagged element order is central to reading order.
  4. Test interaction. Read the document with a screen reader or read-aloud tool. For interactive content, check keyboard tab order as well.
  5. Run a validator or inspection tool. DocRaptor names CommonLook PDF Validator and PAC as checking tools, axesPDF as a checking and remediation tool, and pdfGoHTML for inspecting tags. pdfGoHTML does not itself test conformance against standards.
  6. Repair and repeat. Fix source HTML or apply custom tagging where necessary, regenerate the PDF, and review the changed output.

No single automated check establishes practical usability for every reader. Combine automated inspection with manual review and assistive technology testing.

4. Handle complex tagging and layout

Automatic mapping is useful, but some content needs deliberate review or custom tagging. DocRaptor documents Prince’s -prince-pdf-tag-type CSS property for assigning tag types, including examples involving table-of-contents tags, decorative content, and ARIA roles.

Use a non-structural or artifact treatment only for content that is genuinely decorative. Do not hide meaningful repeated content such as page numbers. Review complex tables, multi-column pages, generated content, and tables of contents in the resulting tag tree because visual correctness does not guarantee logical structure.

5. Common problems and fixes

Symptom Likely cause What to do
The request rejects the profile or does not produce tagged output. The account’s pipeline may not support PDF/UA-1, or the option is in the wrong place. Check the current API reference for pipeline support and send profile nested inside prince_options.
The PDF is untagged or has a poor tag tree despite a successful response. Conversion success does not guarantee meaningful structure; source markup or layout may not map as intended. Inspect the tag tree, improve semantic HTML, and use custom tagging where needed. Regenerate and inspect again.
Screen reader order differs from the visual order. CSS positioning or a complex layout changed visual placement while the source order remained different. Reorder the HTML to match the intended sequence where possible, then test the PDF with assistive technology.
Expected content is missing or stale. The page depended on JavaScript, which DocRaptor disables by default. Prefer complete server-rendered HTML. If JavaScript is required, enable the documented option and verify the rendered PDF content.
Images are announced poorly or skipped. Alternative text is absent or unsuitable, or decorative and informative images were not distinguished. Provide meaningful alternatives for informative images and ensure decorative images are treated as decorative; verify the PDF tags.
Bookmarks are missing or confusing. The heading outline may be incomplete or not reflect the intended navigation. Use a meaningful heading hierarchy and inspect generated bookmarks after conversion.
A validator reports issues even though the file opens. Opening the PDF proves only that a viewer can load it, not that its structure and content meet accessibility needs. Use validator findings to guide repair, then manually review reading order, contrast, and assistive technology behavior.

6. Reliability, performance, and cost considerations

The supplied documentation does not provide a conversion benchmark or an accessibility pass-rate figure, so plan around your own document complexity and review workflow rather than assuming a fixed processing time or success rate.

  • Make output review reproducible. Keep representative source documents and inspect regenerated PDFs when templates, CSS, fonts, or content structure change.
  • Reduce rendering uncertainty. Supply complete HTML where practical. If JavaScript is necessary, it is an explicit configuration choice because it is disabled by default.
  • Budget for remediation. PDF generation is only one step; manual investigation and fixes may be needed for complex layouts or content that does not receive appropriate tags automatically.
  • Keep claims scoped. PDF/UA-1 is a tagging profile request. It is not by itself proof of full WCAG conformance or satisfaction of a legal obligation.
  • Check current service terms and pricing directly. The sources in this guide do not establish DocRaptor pricing or service-level figures.

Or skip the browser setup

For website screenshots used in reports, documentation, or review workflows, ScreenshotNeo is a website screenshot API and MCP server. It captures a URL as PNG, JPEG, WebP, or PDF. It does not create a tagged PDF or replace the accessibility workflow above.

One GET request returns a capture. See the ScreenshotNeo API documentation for the request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; each cleanup step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents.
  • The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots.

Sign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.

FAQ

Does setting PDF/UA-1 guarantee an accessible PDF?

No. It requests tagged output. Source semantics, visual presentation, reading order, interaction, and user context still need review.

Does PDF/UA-1 mean the document meets every WCAG requirement?

No. Do not treat a PDF profile setting as proof of full WCAG conformance or legal compliance.

Should I enable JavaScript for every conversion?

No. It is disabled by default. Enable it only when the HTML depends on client-side rendering, then check that the resulting PDF contains the intended content.

Are WCAG techniques mandatory implementation recipes?

No. W3C describes techniques as informative examples of ways to meet WCAG; the outcome and applicable success criteria matter.

References: W3C WCAG 2.2 techniques and Understanding SC 1.4.3 contrast minimum.