ScreenshotNeo

BlogHTML to image & PDF

How to Fix Damaged PDFs Returned as Puppeteer Blob Responses

Trace PDF corruption from Puppeteer to the browser Blob, preserve bytes, validate headers, and fix the boundary where data changes.

By the ScreenshotNeo team1 October 20268 min read

Direct answer: a Blob rarely damages a valid PDF by itself. Find the first boundary where the bytes change: Puppeteer output, your HTTP response, or the browser download. Keep the PDF as Uint8Array/Buffer bytes, never decode it as UTF-8 or stringify it, return a 2xx response with Content-Type: application/pdf, and make Content-Encoding match the actual representation. Check the response status before creating a Blob so an HTML or JSON error is not saved as .pdf.

1. Identify which Puppeteer path you use

There are two different problems that are often described as a “corrupt Blob.”

  • PDF generation: await page.pdf() creates PDF bytes and resolves to a Uint8Array. Puppeteer documents saving that value directly to a file: Page.pdf().
  • PDF response capture: you navigate to a URL and read an existing PDF with HTTPResponse.buffer() or content(). Puppeteer warns that the browser may re-encode the buffer using headers or heuristics, which can produce incorrect bytes: HTTPResponse.buffer().

Test generation and capture separately. If page.pdf() is already invalid before your server sends it, the HTTP layer is not the cause. If only captured responses fail, investigate the response-buffer caveat and the origin’s headers first.

2. Generate and return a PDF without text conversion

Minimal Puppeteer generation

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch({ headless: true });
try {
  const page = await browser.newPage();
  await page.goto('https://example.com', { waitUntil: 'networkidle0' });

  const pdfBytes = await page.pdf({
    format: 'A4',
    printBackground: true
  });

  // pdfBytes is a Uint8Array. Keep it binary.
  await Bun.write('document.pdf', pdfBytes);
} finally {
  await browser.close();
}

With Node’s filesystem API, use Buffer.from(pdfBytes) and fs.writeFile. Do not call pdfBytes.toString(), pass it through JSON, or interpolate it into a string.

Express endpoint

import express from 'express';
import puppeteer from 'puppeteer';

const app = express();
app.get('/document.pdf', async (req, res) => {
  let browser;
  try {
    browser = await puppeteer.launch({ headless: true });
    const page = await browser.newPage();
    await page.goto('https://example.com', { waitUntil: 'networkidle0' });
    const pdfBytes = await page.pdf({ format: 'A4', printBackground: true });

    res.status(200);
    res.set({
      'Content-Type': 'application/pdf',
      'Content-Disposition': 'attachment; filename="document.pdf"',
      'Content-Length': String(pdfBytes.byteLength),
      'Cache-Control': 'no-store'
    });
    res.end(Buffer.from(pdfBytes));
  } catch (error) {
    if (!res.headersSent) res.status(500).json({ error: 'PDF generation failed' });
  } finally {
    await browser?.close();
  }
});

app.listen(3000);

Content-Disposition: attachment controls download presentation; it does not repair bytes. Use an inline disposition or omit it when you want a viewer to open the PDF.

3. Capture an existing PDF response safely

const response = await page.goto('https://example.com/report.pdf', {
  waitUntil: 'load'
});

if (!response) throw new Error('No navigation response');
const headers = response.headers();
console.log({ status: response.status(), contentType: headers['content-type'] });

if (response.status() < 200 || response.status() >= 300) {
  throw new Error(`Origin returned HTTP ${response.status()}`);
}

const bytes = await response.buffer();
await fs.promises.writeFile('captured.pdf', bytes);

Record the URL, status, Content-Type, Content-Encoding, and byte length before writing. If the origin is returning a PDF through a redirect, authentication page, proxy, or bot check, the final body may be HTML or JSON even when the URL ends in .pdf.

4. Diagnose every byte boundary

  1. Hash or count the Puppeteer value immediately after page.pdf() or response.buffer().
  2. Hash or count the exact bytes passed to your framework’s response method.
  3. Inspect the browser’s downloaded Blob or saved file with the same measurements.
  4. The first boundary whose length or digest differs contains the defect.

This digest workflow is a practical diagnostic inference. Do it in a local or protected environment and never log PDF contents or personal data.

Quick binary checks

import crypto from 'node:crypto';

function fingerprint(value) {
  const bytes = Buffer.from(value);
  return {
    bytes: bytes.length,
    sha256: crypto.createHash('sha256').update(bytes).digest('hex')
  };
}

console.log(fingerprint(pdfBytes));

A PDF normally begins with a PDF header, but a header check alone cannot prove the entire file is valid. Open the final artifact in a PDF reader or run a PDF parser/validator after comparing boundaries.

5. Check the browser Fetch and Blob code

const response = await fetch('/document.pdf');

if (!response.ok) {
  const diagnostic = await response.text();
  throw new Error(`HTTP ${response.status}: ${diagnostic.slice(0, 500)}`);
}

const contentType = response.headers.get('content-type') || '';
if (!contentType.toLowerCase().includes('application/pdf')) {
  throw new Error(`Expected PDF, received ${contentType}`);
}

const blob = await response.blob();
const url = URL.createObjectURL(blob);
const link = document.createElement('a');
link.href = url;
link.download = 'document.pdf';
link.click();
// Revoke the object URL after the consumer has started using it.
setTimeout(() => URL.revokeObjectURL(url), 0);

Response.blob() consumes the body and creates a Blob containing those response bytes; its type comes from Content-Type. An opaque response produces an empty Blob with an empty type. A Blob constructor cannot repair bytes that were already changed upstream.

For byte-oriented diagnostics, use arrayBuffer():

const response = await fetch('/document.pdf');
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const arrayBuffer = await response.arrayBuffer();
const bytes = new Uint8Array(arrayBuffer);
console.log('downloaded bytes:', bytes.byteLength);

Response.arrayBuffer() can fail when body decoding fails, including a wrong Content-Encoding. Do not call response.text() on a response known to be a PDF merely to inspect it; that interprets binary data as text.

6. Content-Type versus Content-Encoding

Header Purpose Correct PDF example
Content-Type Describes the media type application/pdf
Content-Disposition Suggests inline display or download attachment; filename="document.pdf"
Content-Encoding Describes how the representation is compressed gzip only when the body is actually gzip encoded

These headers affect interpretation and decoding; they cannot make already-corrupted bytes valid. Check compression at every proxy, CDN, framework, and server boundary. Avoid manually setting Content-Encoding unless you are actually applying that encoding.

7. Common causes and fixes

Symptom Likely cause Fix
File contains readable HTML Login page, bot check, proxy error, or application error Check status and content type before reading the body as a PDF; authenticate or fix the upstream route.
Size changes after server handling UTF-8 conversion, JSON serialization, or framework string handling Pass a Buffer/Uint8Array directly and end the response with binary data.
Only captured origin PDFs fail Puppeteer response-buffer re-encoding based on headers or heuristics Compare the origin response and Puppeteer buffer; verify origin headers and test a direct HTTP client.
Browser reports decoding failure Incorrect Content-Encoding or double compression Remove the header unless compression is applied, or configure the proxy and app consistently.
Blob has zero bytes Opaque Fetch response or an empty upstream body Fix CORS/request mode and inspect response.ok, status, and length before creating the Blob.
Download opens but pages are blank Generation completed before fonts, images, or lazy content loaded Wait for a selector, a delay, or network idle before calling page.pdf(); verify the page itself first.
Works locally, fails in production Different Puppeteer/browser version, proxy, compression, or runtime response adapter Record versions and headers, then compare hashes at each boundary. Match behavior to your installed Puppeteer version; documentation versions can differ.

8. Full-buffer versus streaming

page.pdf() returns a complete Uint8Array. page.createPDFStream() returns a ReadableStream<Uint8Array> (createPDFStream()). A full buffer is simplest for setting Content-Length and hashing. A stream can reduce application buffering for large documents, but every consumer must still preserve byte chunks and avoid text transforms. Choose the interface your framework supports reliably, then verify the final file.

9. Reliability, performance, and cost notes

  • Close the browser in a finally block so failed jobs do not leave processes behind.
  • Set navigation and application timeouts appropriate to the target page, and distinguish a timeout from a malformed PDF.
  • Reuse a browser process when safe, but isolate pages and clear per-request state such as cookies and headers.
  • Hashing adds a linear read of the bytes; use it for diagnostics, not as a replacement for a PDF validator.
  • Streaming can reduce peak memory; full buffering simplifies retries, size checks, and response headers.
  • Do not retry blindly on a deterministic HTML error. Record status and content type, then retry only transient navigation or upstream failures.

10. A compact debugging checklist

  • Confirm whether the source is page.pdf() or an existing PDF response.
  • Check status, final URL, Content-Type, and Content-Encoding.
  • Measure and hash bytes immediately after Puppeteer.
  • Search for toString(), UTF-8 decoding, JSON serialization, and accidental base64 handling.
  • Send a Buffer/Uint8Array or byte stream directly from the server.
  • Check response.ok before blob() or arrayBuffer().
  • Compare the browser Blob/file with the server bytes.
  • Open or validate the final file independently.

11. Or skip the browser setup

If you need a hosted screenshot or PDF capture endpoint, ScreenshotNeo returns a PDF or image from one GET request. Its capture flow accepts cookie and consent banners, removes more than 60 known consent platforms plus newsletter popups and chat widgets, and lets you turn each step off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and whether it was billed. It also offers an MCP server for AI agents with take_screenshot, get_page_info, and capture_pdf.

See the ScreenshotNeo API documentation for all options.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes full-page capture, PDF paper and margin controls, custom CSS and JavaScript, waits, request blocking, headers, cookies, user agents, geolocation, caching, signed links, asynchronous jobs, bulk capture, and a usage API. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

12. FAQ

Can changing the Blob MIME type fix a damaged PDF?

No. It changes metadata only. Find and fix the boundary that changed the bytes.

Should I use blob() or arrayBuffer()?

Use blob() when you need a browser download or object URL. Use arrayBuffer() when you need byte counts, hashes, or another binary API.

Does Content-Disposition affect PDF integrity?

No. It controls presentation or download behavior; the body bytes remain the same.

Why can a 200 response still be invalid?

HTTP success describes transport, not media. A route can return a login page, JSON error, or proxy message with status 200.

Is a valid PDF header enough?

No. Validate the complete file with a reader or parser after comparing bytes across boundaries.