ScreenshotNeo

BlogHTML to image & PDF

Add a Text Watermark to PDFs in Python with aiohttp

Download PDFs asynchronously with aiohttp, add a visible or background text watermark, and handle large files, rotations, errors, and uploads safely.

By the ScreenshotNeo team1 October 20268 min read

To add a text watermark to a PDF with Python and aiohttp, use aiohttp only for HTTP transfer and a PDF library for editing. Download the response, verify that it succeeded and is a PDF, insert text with PyMuPDF (direct placement), then save to a new file. For a background stamp with pypdf, first create a one-page PDF containing the watermark text and merge it with over=False.

Complete example: download with aiohttp and watermark with PyMuPDF

This example streams the remote PDF to a temporary file, inserts a centered diagonal watermark on every page, and writes a separate output file. It avoids loading a large response into memory.

import asyncio
import math
import tempfile
from pathlib import Path

import aiohttp
import fitz  # PyMuPDF

SOURCE_URL = "https://example.com/document.pdf"
OUTPUT_PATH = Path("watermarked.pdf")
WATERMARK = "CONFIDENTIAL"


async def download_pdf(url: str, destination: Path) -> None:
    timeout = aiohttp.ClientTimeout(total=90, connect=15, sock_read=60)
    async with aiohttp.ClientSession(timeout=timeout) as session:
        async with session.get(url, allow_redirects=True) as response:
            response.raise_for_status()
            content_type = response.headers.get("Content-Type", "").lower()
            if "pdf" not in content_type and not url.lower().split("?", 1)[0].endswith(".pdf"):
                raise ValueError(f"Expected a PDF, got Content-Type {content_type!r}")

            with destination.open("wb") as output:
                async for chunk in response.content.iter_chunked(1024 * 1024):
                    output.write(chunk)


def add_text_watermark(source: Path, destination: Path, text: str) -> None:
    document = fitz.open(source)
    try:
        for page in document:
            rectangle = page.rect
            point = fitz.Point(rectangle.width * 0.18, rectangle.height * 0.58)
            page.insert_text(
                point,
                text,
                fontsize=max(18, min(rectangle.width, rectangle.height) * 0.06),
                fontname="helv",
                color=(0.55, 0.55, 0.55),
                fill_opacity=0.28,
                overlay=True,
                rotate=45,
            )
        document.save(destination)
    finally:
        document.close()


async def main() -> None:
    with tempfile.TemporaryDirectory() as temporary_directory:
        source = Path(temporary_directory) / "source.pdf"
        await download_pdf(SOURCE_URL, source)
        add_text_watermark(source, OUTPUT_PATH, WATERMARK)
        print(f"Saved {OUTPUT_PATH}")


if __name__ == "__main__":
    asyncio.run(main())

Install the dependencies with python -m pip install aiohttp PyMuPDF. The rotate, coordinates, opacity, and font size are document-specific choices. Test them on portrait, landscape, rotated, and mixed-size pages.

Why aiohttp and the PDF library have separate jobs

aiohttp sends the request and exposes the response body. It does not edit PDF page content. The aiohttp quickstart warns that convenience methods such as read(), text(), and json() read the complete body into memory; use response.content.iter_chunked() for large files. Reuse one ClientSession for related requests so connections can be pooled and kept alive. See the aiohttp Client Quickstart and Client Reference.

Small PDFs: process bytes in memory

For a known, small file, await response.read() is convenient. Keep an application-level size limit when the URL or file size is untrusted.

import asyncio
import io
import aiohttp
import fitz

async def watermark_small_pdf(url: str, output_path: str) -> None:
    async with aiohttp.ClientSession(
        timeout=aiohttp.ClientTimeout(total=90)
    ) as session:
        async with session.get(url) as response:
            response.raise_for_status()
            body = await response.read()

    document = fitz.open(stream=body, filetype="pdf")
    try:
        for page in document:
            page.insert_text(
                (72, 72),
                "INTERNAL",
                fontsize=24,
                color=(0.7, 0.7, 0.7),
                fill_opacity=0.3,
                overlay=True,
            )
        document.save(output_path)
    finally:
        document.close()

asyncio.run(watermark_small_pdf("https://example.com/file.pdf", "marked.pdf"))

Background watermark with pypdf

pypdf’s documented approach merges a one-page stamp PDF onto each target page. The stamp must already contain the text; pypdf’s example does not render text itself. Set over=False to place the watermark beneath existing page content, or over=True for a foreground stamp. The pypdf watermark documentation describes this draw order.

from pathlib import Path
from pypdf import PdfReader, PdfWriter

source = PdfReader("source.pdf")
stamp = PdfReader("watermark-stamp.pdf").pages[0]
writer = PdfWriter()

for page in source.pages:
    page.merge_page(stamp, over=False)
    writer.add_page(page)

with Path("watermarked.pdf").open("wb") as output:
    writer.write(output)

Create watermark-stamp.pdf with a PDF-capable renderer such as ReportLab, or use a prebuilt one-page template. If the stamp has a different page size, scale or translate it before merging and inspect the result on every page size in your input set.

Choosing placement, layer order, and appearance

Decision Recommendation
Above or below content Use overlay=True in PyMuPDF when the mark must always be visible. Use pypdf over=False for a background watermark.
Position Start with page-relative coordinates, then test margins, rotations, and mixed dimensions.
Contrast Use a muted color and opacity that remains legible without hiding important text.
Rotation A diagonal mark is common, but page rotation can change its apparent position. Test actual files.
Font Use a font supported by the library and embed it when your compliance or rendering requirements demand it.

PyMuPDF uses page coordinates and direct text insertion APIs. Its basics guide is the reference for page geometry and text operations. A PDF can contain rotations, unusual media boxes, clipped content, encrypted pages, or malformed objects, so no single coordinate works universally.

Uploading the watermarked PDF with aiohttp

Use FormData for a multipart endpoint. Set the filename and content type explicitly.

import aiohttp

async def upload_pdf(endpoint: str, path: str) -> str:
    form = aiohttp.FormData()
    form.add_field(
        "file",
        open(path, "rb"),
        filename="watermarked.pdf",
        content_type="application/pdf",
    )
    async with aiohttp.ClientSession(
        timeout=aiohttp.ClientTimeout(total=90)
    ) as session:
        async with session.post(endpoint, data=form) as response:
            response.raise_for_status()
            return await response.text()

Close file handles in production with a context manager or an explicit try/finally. If you send an async generator or another non-rewindable body, a redirect may not be able to replay it; configure the final endpoint or handle redirects deliberately.

cURL, Python, and Node.js transfer examples

These snippets show the HTTP download step independently from PDF editing.

curl -L --fail --output source.pdf https://example.com/document.pdf
import requests

response = requests.get("https://example.com/document.pdf", timeout=90)
response.raise_for_status()
with open("source.pdf", "wb") as output:
    output.write(response.content)
const response = await fetch('https://example.com/document.pdf');
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const bytes = Buffer.from(await response.arrayBuffer());
await require('node:fs').promises.writeFile('source.pdf', bytes);

For large files, stream the response to disk in your chosen runtime instead of collecting the entire body.

Production checklist

  • Set connect, read, and total timeouts.
  • Check the HTTP status before parsing bytes.
  • Validate content type and the PDF signature rather than trusting a filename.
  • Apply a maximum download size and clean up temporary files.
  • Write to a new output path and replace the original only after a successful save.
  • Reuse a ClientSession for batches.
  • Test portrait, landscape, rotated, mixed-size, encrypted, and dense PDFs.
  • Log the source URL, response status, processing error, and output path without logging secrets.

Troubleshooting

Symptom Cause Fix
ClientResponseError The server returned 4xx or 5xx. Call raise_for_status(), inspect the status and response headers, and correct authentication, URL, or permissions.
PDF parser error The response is HTML, JSON, a bot page, or a truncated file. Check status, content type, size, and the first bytes before opening it.
Memory spikes read() loaded a large body. Stream with iter_chunked() to a temporary file.
Watermark is invisible It is behind opaque content, clipped, too transparent, or outside the page. Use foreground insertion, adjust coordinates and opacity, and inspect page bounds.
Watermark is rotated or misplaced Page rotation or a nonstandard media box changed the coordinate system. Normalize or account for rotation and test representative pages.
Output overwrites the source Input and output paths are identical. Always save to a distinct path, then replace atomically after validation.
Upload fails after redirect The request body cannot be replayed. Use the final upload URL or a rewindable file body.
Encrypted or restricted PDF fails The document requires a password or disallows edits. Obtain authorization and password, then handle the library’s encryption and permission errors explicitly.

Performance, reliability, and cost notes

Network time, PDF size, number of pages, and font or image complexity determine total processing time. Streaming reduces peak transfer memory but does not make PDF editing constant-memory; the editor still parses the document. Keep the source and output on fast local storage, reuse sessions for batches, and retry only failures that are safe to retry. Use bounded concurrency so many large PDFs do not exhaust memory or file descriptors. The cited documentation does not provide a universal throughput or memory benchmark, so measure with your own representative files.

aiohttp and the open-source PDF libraries do not charge per request, but your hosting, storage, bandwidth, and any upstream PDF service may have costs. Enforce limits before downloading untrusted URLs.

Or skip the browser setup

If your PDF workflow starts with a webpage, ScreenshotNeo can capture the page through one API request before you process the resulting asset. Its consent handling removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing result. An MCP server lets Claude, Cursor, and other MCP clients use take_screenshot, get_page_info, and capture_pdf.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for capture and PDF options. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Can aiohttp add a watermark by itself?

No. It transfers bytes; PyMuPDF, pypdf, or another PDF library edits page content.

Should I use pypdf or PyMuPDF?

Use PyMuPDF when you want direct text placement. Use pypdf when you already have a stamp PDF and need documented merge order.

Should the watermark be above or below the page?

Choose foreground placement for visibility and background placement when preserving the original page’s visual hierarchy matters.

Is reading the whole response always wrong?

No. It is practical for small, trusted PDFs. Stream large or untrusted responses and impose a size limit.

How do I watermark only selected pages?

Iterate with a page index and insert or merge the watermark only when that index is in your target set.