Generating Documents with an API
Learn how to generate PDFs, DOCX files, and collaborative documents from JSON with templates, APIs, validation, retries, and delivery workflows.
Direct answer: generating a document with an API means sending structured data and document instructions to a service that merges a template, creates or updates a document resource, or converts an existing file into the format you need. Choose template merging for repeatable branded PDFs and DOCX files, a document-resource API for collaborative editing, and conversion APIs when your source is already HTML, Word, spreadsheets, slides, images, or another supported file.
A reliable implementation validates the input before submission, authenticates with a server-side credential, submits an idempotent request, checks the returned artifact, stores it with access controls, and monitors failures and vendor limits.
Choose the document model first
| Requirement | Best fit | Why |
|---|---|---|
| Branded invoice, contract, proposal, certificate, or statement | Template merge | The template controls typography, layout, headers, and pagination while JSON supplies recipient-specific values. |
| Users need to keep editing and commenting in a shared workspace | Collaborative document API | A document resource can be created, retrieved, and updated over time. |
| Your system already produces HTML, DOCX, XLSX, PPTX, images, or text | Conversion API | The API focuses on producing a fixed-layout PDF from an existing source. |
| Free-form generated content or many file formats | AI-assisted file generation | The model can create files, but your application still needs validation, review, and retention controls. |
Adobe describes its Document Generation API as merging JSON data into Word-based templates to produce PDF and Word documents from an application (Adobe Document Generation API). Adobe lists invoices, contracts, sales proposals, and work orders as examples of stable template content combined with dynamic data. The Google Docs API exposes documents.create, documents.get, and documents.batchUpdate; batch updates apply a set of edit requests atomically. Google lists bulk documentation, formatting, invoices, and contracts among its use cases.
Design a document-generation workflow
- Define the canonical schema. Decide required fields, types, currency rules, date and timezone rules, locale, and maximum lengths.
- Select the output contract. PDF is fixed-layout, DOCX remains editable in Office, HTML is useful for web delivery, and Google Docs is suited to shared cloud editing.
- Create the template or document structure. Use stable placeholders such as
customer.nameandinvoice.total. Keep formatting in the template rather than embedding presentation rules in business data. - Validate and normalize input. Reject missing fields, invalid totals, unsupported locales, and values that could break the template.
- Authenticate server-side. Keep API keys and OAuth credentials out of browsers, mobile binaries, and generated files.
- Submit with correlation data. Use an idempotency key when the provider supports it; otherwise store your own request ID and result mapping.
- Validate the result. Check status, content type, file size, page count, required text, and any provider verdict before delivery.
- Render representative samples. Inspect long names, large tables, page breaks, fonts, right-to-left text, and localized numbers.
- Store and deliver. Apply retention limits, encryption, access controls, and a delivery path such as download, object storage, email, or a shared workspace.
- Monitor operations. Record request IDs, latency, retries, template versions, quota responses, and failure categories.
Template merging from JSON
Template merging is the most direct pattern for repeatable branded documents. Keep a versioned DOCX template with fields, validate a JSON payload, send both to the provider, and save the returned DOCX or PDF. The same approach works for invoices, contracts, proposals, statements, certificates, and work orders.
Example input schema
{
"template_version": "invoice-v3",
"invoice_number": "INV-1042",
"issue_date": "2026-10-01",
"due_date": "2026-10-31",
"currency": "USD",
"customer": {
"name": "Ada Lovelace",
"email": "ada@example.com",
"address": "12 Analytical Engine Way"
},
"items": [
{"description": "API integration", "quantity": 2, "unit_price": "125.00"}
],
"notes": "Thank you for your business."
}
Validate before calling the provider (Python)
from decimal import Decimal
from datetime import date
def validate_invoice(data):
required = ["template_version", "invoice_number", "issue_date", "due_date", "currency", "customer", "items"]
missing = [key for key in required if not data.get(key)]
if missing:
raise ValueError(f"Missing fields: {', '.join(missing)}")
date.fromisoformat(data["issue_date"])
date.fromisoformat(data["due_date"])
if data["currency"] not in {"USD", "EUR", "GBP"}:
raise ValueError("Unsupported currency")
if not data["items"]:
raise ValueError("At least one line item is required")
total = Decimal("0")
for item in data["items"]:
quantity = Decimal(str(item["quantity"]))
unit_price = Decimal(str(item["unit_price"]))
if quantity <= 0 or unit_price < 0:
raise ValueError("Invalid line item")
total += quantity * unit_price
return total.quantize(Decimal("0.01"))
Provider request shape
Provider field names differ, but the request normally contains a template identifier, the validated JSON data, the requested output format, and a correlation or idempotency key. Keep this adapter behind your own interface so changing vendors does not change the rest of your application.
{
"template": "invoice-v3",
"output": ["docx", "pdf"],
"data": { "...": "validated invoice JSON" },
"correlation_id": "order-8f2c"
}
Creating and updating a collaborative Google document
Use Google Docs when the result should remain a document resource that people can edit and comment on. The API returns a document ID from creation; later calls retrieve the resource or apply edits with batchUpdate. OAuth setup and scopes are required, so the examples assume an access token in GOOGLE_ACCESS_TOKEN.
cURL: create a document
curl -X POST "https://docs.googleapis.com/v1/documents" \
-H "Authorization: Bearer $GOOGLE_ACCESS_TOKEN" \
-H "Content-Type: application/json" \
-d '{"title":"Invoice INV-1042"}'
Python: create, then insert text
import os
import requests
TOKEN = os.environ["GOOGLE_ACCESS_TOKEN"]
headers = {
"Authorization": f"Bearer {TOKEN}",
"Content-Type": "application/json",
}
created = requests.post(
"https://docs.googleapis.com/v1/documents",
headers=headers,
json={"title": "Invoice INV-1042"},
timeout=30,
)
created.raise_for_status()
document_id = created.json()["documentId"]
update = requests.post(
f"https://docs.googleapis.com/v1/documents/{document_id}:batchUpdate",
headers=headers,
json={"requests": [{"insertText": {"location": {"index": 1}, "text": "Invoice INV-1042\\n"}}]},
timeout=30,
)
update.raise_for_status()
print(document_id)
Node.js: create and update
const token = process.env.GOOGLE_ACCESS_TOKEN;
const headers = {
Authorization: `Bearer ${token}`,
'Content-Type': 'application/json'
};
const created = await fetch('https://docs.googleapis.com/v1/documents', {
method: 'POST', headers,
body: JSON.stringify({ title: 'Invoice INV-1042' })
});
if (!created.ok) throw new Error(`Create failed: ${created.status}`);
const { documentId } = await created.json();
const update = await fetch(`https://docs.googleapis.com/v1/documents/${documentId}:batchUpdate`, {
method: 'POST', headers,
body: JSON.stringify({
requests: [{ insertText: { location: { index: 1 }, text: 'Invoice INV-1042\\n' } }]
})
});
if (!update.ok) throw new Error(`Update failed: ${update.status}`);
console.log(documentId);
Conversion-focused PDF generation
Conversion is appropriate when your upstream system already emits a supported source. Adobe PDF Services documents conversion from HTML, Word, PowerPoint, Excel, text, images, ZIP files, and URLs (PDF Services API). Treat conversion as a separate stage from data assembly: first produce and validate the source, then convert it, then inspect the PDF.
AI-assisted file generation
OpenAI Code Interpreter can return files in formats including DOCX, HTML, PDF, PPTX, XLSX, JSON, Markdown, and text through file annotations (Code Interpreter documentation). ChatGPT Work can create or edit documents from instructions, source material, or reusable templates depending on plan, workspace, file type, and surface availability. For production workflows, add deterministic schema validation, content checks, rendering review, and retention rules around any AI-generated artifact.
Or skip the browser setup
If your workflow needs a screenshot of a generated document preview or any public page, ScreenshotNeo provides a single GET request for PNG, JPEG, WebP, or PDF output. Cookie banners, newsletter popups, and chat widgets are removed before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Responses identify the result with X-Page-Verdict and X-Billed headers. Its MCP server gives Claude, Cursor, and other MCP clients take_screenshot, get_page_info, and capture_pdf tools.
See the ScreenshotNeo API documentation for all options. The same request can be made from cURL, Python, or Node.js:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
There are 1,000 free screenshots per month with no card. Paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Production safeguards
Validation and correctness
- Use a versioned schema and template together.
- Calculate totals on the server and compare them with supplied totals.
- Normalize dates, decimal precision, currencies, units, and timezones.
- Escape user content and prevent template expressions from being interpreted as code.
- Check generated files for expected MIME type, nonzero size, page count, required identifiers, and readable text.
Retries and idempotency
Retry transient network failures, rate limits, and server errors with exponential backoff and jitter. Do not blindly retry validation errors or authentication failures. If the provider supports idempotency keys, reuse the same key for the same logical document. Otherwise persist a correlation ID and avoid creating a duplicate when a response is lost after submission.
Performance
Measure latency for your actual template size, number of pages, fonts, images, and concurrency. Cache immutable templates and static assets, batch independent work where the provider supports it, and move long jobs to a queue. Keep request and response bodies out of verbose logs when they contain personal or financial data.
Reliability and retention
Keep the source JSON, template version, provider request ID, artifact checksum, and delivery status long enough to reproduce an incident. Define retention and deletion rules for both source data and generated files. Use signed, expiring download links for private artifacts and separate operational logs from document contents.
Cost
Cost depends on the provider, output format, page count, conversion volume, storage, and retries. Measure cost per completed document rather than cost per request, because failed calls and duplicate retries can distort totals. Set quotas and alerts before enabling bulk generation.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| 400 or validation error | Missing field, wrong type, invalid date, or malformed template data | Log the schema error, validate locally, and reject the request before submission. |
| 401 or 403 | Expired token, wrong scope, revoked key, or credential sent from an unsafe client | Refresh or rotate credentials, verify scopes, and keep secrets server-side. |
| Document has blank placeholders | Field name does not match the template or nested data path | Compare the template field map with the canonical schema and add a fixture test for every required field. |
| Unexpected page breaks | Long text, table growth, missing fonts, or different rendering engine | Render boundary cases, constrain text where appropriate, embed or configure fonts, and inspect the output PDF. |
| Missing images or fonts | Private asset URL, blocked network request, unsupported format, or unavailable font | Use accessible assets, verify content types, package required fonts where supported, and check conversion logs. |
| Duplicate documents | Client timeout caused a retry after the first request succeeded | Use idempotency keys or persist a correlation ID and reconcile provider results before retrying. |
| 429 or quota errors | Rate limit or daily quota exceeded | Apply exponential backoff, queue work, reduce concurrency, and request a quota change when justified. |
| Unreadable localized output | Locale, timezone, decimal, or right-to-left handling was implicit | Pass locale and timezone explicitly and test representative languages, currencies, and date boundaries. |
Short FAQ
Should I generate PDF or DOCX?
Use PDF when the visual layout must remain fixed for delivery or archiving. Use DOCX when recipients need to edit the file.
Is a template required?
No. Resource APIs and AI-assisted tools can create documents programmatically, but templates are usually easier to control for branded, repeatable business documents.
Can I use Google Docs as a PDF generator?
You can create and update a Google document, then add an export step. Choose a conversion-focused PDF API when PDF production is the primary requirement.
How should I test document generation?
Use fixture data for short and long names, empty and large tables, boundary dates, multiple currencies, different locales, and page-break cases. Inspect both machine-readable fields and rendered pages.
Where should generated files live?
Use controlled object storage or the provider's document workspace with encryption, access policies, retention rules, and expiring delivery links.


