ScreenshotNeo

BlogHow-to

PDFShift API Returns 422 Invalid HTML: How to Troubleshoot

A 422 response alone does not identify the cause. Capture the full error body, verify the request and source, then isolate raw HTML and URL failures.

By the ScreenshotNeo team4 October 20269 min read

If PDFShift returns 422 invalid HTML, start with the complete response body and the exact request you sent. The available PDFShift documentation shows how to submit raw HTML or a URL as the source, and how to inspect unsuccessful responses. It does not define this exact error phrase or establish one specific cause. Treat malformed markup, missing or mis-encoded input, and source retrieval as hypotheses to investigate—not confirmed explanations.

This guide walks through a repeatable diagnosis: preserve the response, verify the request envelope, determine whether you sent HTML or a URL, then reduce the input until you can identify what changes the result.

1. Capture the complete error response

Do not discard the response body when a request fails. Log the HTTP status and response text, then compare them with the request method, endpoint, and a safe description of the source. PDFShift’s examples check for unsuccessful responses and expose the response content; its Python examples use raise_for_status(). PDFShift’s Python raw HTML guide and httplib2 example show these patterns.

Keep API keys and private document contents out of application logs. If the response includes sensitive input, redact it before sharing it with a teammate or support. Preserve an unmodified copy in an appropriately protected debugging environment if you need it for comparison.

HTTP status: 422
Response body: [copy the complete response body here]
Request: POST https://api.pdfshift.io/v3/convert/pdf
Source mode: raw HTML | URL

2. Verify the request envelope

PDFShift’s documented v3 conversion endpoint is https://api.pdfshift.io/v3/convert/pdf. The request examples use POST, a JSON body, an API key in the X-API-Key header, and a source value. Confirm these parts match your code or workflow before changing the HTML.

  • Use the documented conversion endpoint and POST method.
  • Send a JSON object with a source field.
  • Send the API key in the configured header; do not print it to logs.
  • Check that your HTTP client is actually sending JSON, rather than form data or a JSON string wrapped inside another JSON string.
  • Record the request mode and a source length or URL host, rather than logging a private full document.

These checks validate the documented request shape; they do not prove that any one mismatch causes this particular 422.

3. Identify which source path failed

The source field accepts either raw HTML or a URL. Test the two paths separately because they involve different inputs and failure points. PDFShift documents raw HTML input; its guides also demonstrate URL sources and the raise_for_status option for failing when a remote source does not return a successful status.

When source is raw HTML

  • Inspect the final string immediately before JSON serialization. Confirm it contains the complete document you expect, rather than an empty value, a template placeholder, or a truncated fragment.
  • Let your JSON library serialize the string. Do not manually escape quotes, backslashes, or newlines and then serialize it again.
  • Check how your template handles untrusted or dynamic content. A malformed template result is a possibility to test, not a known explanation for PDFShift’s wording.
  • Try a minimal document, then add your generated markup, styles, scripts, fonts, and images back in stages.

When source is a URL

  • Check that the URL is correct and the page is reachable without a browser session, VPN, or login that PDFShift cannot access.
  • Inspect redirects, access controls, and whether the route responds successfully to a server-side fetch.
  • Separate failure to retrieve the page from problems with resources the page loads, such as stylesheets, scripts, fonts, and images.
  • Try sending the page’s raw HTML as the source. If that changes the result, investigate remote retrieval and dependencies separately.

A URL that works in your browser is not proof that a conversion service can retrieve it: the browser may carry cookies or follow a path that an unauthenticated service request cannot.

4. Reduce the document to a minimal reproduction

Use a controlled sequence so each request answers one question. This is a diagnostic method, not a PDFShift-published remedy for the exact 422.

  1. Save the failing request’s status, full response body, source mode, and a safe copy of the relevant input.
  2. Send a minimal raw HTML document with a title and one paragraph.
  3. If it succeeds, restore the generated body while keeping external assets out.
  4. Add styles, scripts, fonts, and images one group at a time.
  5. If the failure only occurs with a URL, test a publicly reachable simple page, then compare it with your target URL.
  6. Keep each test’s response beside the exact input variant so the change is clear.
<!doctype html>
<html>
  <head><meta charset="utf-8"><title>PDFShift diagnostic</title></head>
  <body><h1>Minimal document</h1><p>Conversion input check.</p></body>
</html>

5. Runnable request examples

The following examples submit raw HTML as JSON and preserve the error body when the request fails. Replace the placeholder API key. Use the URL-source variant shown after them if the problem only occurs with a remote page.

cURL

HTML='<!doctype html><html><head><meta charset="utf-8"><title>Diagnostic</title></head><body><h1>Hello</h1></body></html>'
curl --silent --show-error --include \
  --request POST 'https://api.pdfshift.io/v3/convert/pdf' \
  --header 'X-API-Key: YOUR_API_KEY' \
  --header 'Content-Type: application/json' \
  --data "$(python3 -c 'import json,sys; print(json.dumps({"source":sys.argv[1]}))' "$HTML")" \
  --output result.pdf

For reliable diagnostics, separate response headers and body rather than relying on --include when saving a PDF. This shell example displays headers and writes the response body to a temporary file; inspect the HTTP status and body before treating the file as a PDF:

curl --silent --show-error \
  --request POST 'https://api.pdfshift.io/v3/convert/pdf' \
  --header 'X-API-Key: YOUR_API_KEY' \
  --header 'Content-Type: application/json' \
  --data '{"source":"<!doctype html><html><body><h1>Diagnostic</h1></body></html>"}' \
  --dump-header response-headers.txt \
  --output response-body.bin
cat response-headers.txt

For production scripts, check the status before naming or storing the body as a PDF. An error response is not a PDF file.

Python

import os
import requests

api_key = os.environ["PDFSHIFT_API_KEY"]
html = """<!doctype html>
<html><head><meta charset=\"utf-8\"><title>Diagnostic</title></head>
<body><h1>Hello</h1></body></html>"""

response = requests.post(
    "https://api.pdfshift.io/v3/convert/pdf",
    headers={"X-API-Key": api_key},
    json={"source": html},
    timeout=90,
)

if not response.ok:
    # Keep this output in a protected diagnostic environment.
    print(f"HTTP {response.status_code}: {response.text}")
    response.raise_for_status()

with open("result.pdf", "wb") as output:
    output.write(response.content)

Install the dependency with python -m pip install requests and set PDFSHIFT_API_KEY in the environment. The timeout is a client-side limit for this example, not a statement about PDFShift’s service limits.

Node.js

const apiKey = process.env.PDFSHIFT_API_KEY;
if (!apiKey) throw new Error("Set PDFSHIFT_API_KEY first");

const html = `<!doctype html>
<html><head><meta charset="utf-8"><title>Diagnostic</title></head>
<body><h1>Hello</h1></body></html>`;

const response = await fetch("https://api.pdfshift.io/v3/convert/pdf", {
  method: "POST",
  headers: {
    "X-API-Key": apiKey,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({ source: html }),
  signal: AbortSignal.timeout(90_000),
});

if (!response.ok) {
  const errorBody = await response.text();
  console.error(`HTTP ${response.status}: ${errorBody}`);
  throw new Error("PDFShift conversion failed; inspect the response above");
}

const pdf = Buffer.from(await response.arrayBuffer());
await import("node:fs/promises").then(({ writeFile }) =>
  writeFile("result.pdf", pdf)
);

This uses the built-in fetch available in current Node.js releases. If your runtime lacks it, use an HTTP client that sends JSON and exposes the full error response.

URL-source variant

In each example, replace the HTML value with a URL string, for example "https://example.com/", while leaving the JSON source field intact. Use a page that the conversion service can retrieve. PDFShift documents raise_for_status for making a failed remote fetch fail the conversion; a browser-visible page alone does not establish that the service can fetch it.

6. Common errors and fixes

Symptom What to check Next action
Only “422” is logged The response body was dropped by the client or error handler. Read and preserve the complete response body before raising or returning from the request.
Raw HTML request fails, URL request differs How the string is generated and serialized; compare the final source value. Use the HTTP library’s JSON serializer and test a minimal HTML string.
URL request fails but browser opens the page Authentication, cookies, redirects, access controls, and server-side reachability. Test a public route or submit raw HTML to isolate retrieval from conversion.
The saved “PDF” is unreadable The response may be an error payload saved with a PDF filename. Check status and content before writing a successful result as a PDF.
Minimal HTML succeeds but the full document fails Template output and external resources added to the document. Add components back in stages; inspect the response for every variant.
Request behaves differently across environments Different API key configuration, source generation, JSON encoding, or network access. Compare method, endpoint, headers, source mode, and a redacted source fingerprint across environments.

These are investigation paths, not claims about the undocumented meaning of PDFShift’s exact error phrase. If the response body does not clarify the failure, send PDFShift support the status, complete response body, a minimal reproducible request, and whether the source is HTML or a URL. Redact credentials and confidential document data.

7. Performance, reliability, and cost considerations

When diagnosing a conversion, reduce variables before optimizing. PDFShift recommends avoiding unnecessary network requests, sending raw HTML instead of a URL when appropriate, inlining CSS and JavaScript where possible, removing unneeded scripts, using base64 image data where suitable, and optimizing image dimensions. Its conversion-time guide gives those recommendations. They can reduce external loading dependencies and conversion time, but are not documented as guaranteed fixes for this 422.

  • Performance: A URL source requires a remote fetch, and external assets can add more requests. Inlining suitable dependencies can make a test more controlled.
  • Reliability: Keep error responses and input variants; distinguish request failures from successful PDF bytes. Use an explicit client timeout appropriate to your application and handle non-success responses.
  • Cost: No pricing or billing conclusion follows from a 422 response in the cited troubleshooting material. Check your account’s current usage and plan details rather than assuming whether failed attempts are billable.

Or skip the browser setup

If your immediate goal is a visual capture of a web page rather than a PDF conversion, ScreenshotNeo is a website screenshot API with a one-request workflow. It is not a fix for PDFShift’s 422 and does not replace a PDF conversion when you need a PDF document. Its API can return screenshots or PDFs, and its clean-capture options address common page overlays.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Does 422 prove that my HTML is malformed?

No. The documentation available for this guide does not define the exact “422 invalid HTML” message. Use the response body and a minimal reproduction to determine what the failing request reveals.

Should I switch from a URL source to raw HTML?

It is a useful diagnostic when you can obtain the page markup. It removes the initial remote-page fetch from the test, though external assets in the markup can still require network requests.

What should I send support if the error remains unclear?

Send the status, full response body, a minimal reproducible request, and whether source contains raw HTML or a URL. Remove API keys and sensitive content first.

Can ScreenshotNeo diagnose a PDFShift conversion error?

No. ScreenshotNeo captures web pages; it can help when the separate task is taking a visual screenshot of a page. Use PDFShift’s response details and support for its conversion error.