PDFShift Data Privacy and Security for Confidential Web Pages
PDFShift says its basic conversion flow does not retain requests or generated documents. Here is what that claim covers, where optional storage applies, and what to verify before sending confidential HTML.
Short answer: PDFShift’s FAQ says it does not store conversion requests or generated documents in the basic flow. It also says you can submit HTML directly, so the source page does not have to be public. Those are vendor statements, not independent verification of the service’s controls. Optional output storage changes the data flow: when you use the filename option, PDFShift says the generated PDF is stored in its Amazon S3 in Paris for up to two days; it also describes delivery to your own S3.
Before sending confidential pages, identify exactly what you will submit, whether the request uses an optional storage feature, which contractual terms apply, and how the service’s subprocessors and deletion commitments fit your requirements. This guide summarizes PDFShift’s published materials; it is not a legal or security assessment of your implementation.
1. Does PDFShift store the HTML or PDFs I convert?
PDFShift’s FAQ says: “We do not store requests or generated documents.” Its described basic flow is to receive the source, convert it, return the generated PDF as binary data, and not retain the request or document. The FAQ says this means PDFShift cannot view the converted content. Treat this as the vendor’s description of the basic flow, not as an independently tested guarantee for every endpoint or configuration.
Keep these data categories separate when you assess “storage”:
| Data or flow | What the published materials say | What to verify |
|---|---|---|
| Submitted HTML and request | The FAQ says requests are not stored in the basic conversion flow. | Confirm the exact endpoint and options your integration uses. |
| Returned PDF in the basic flow | The FAQ says generated documents are not stored; the help-center index describes binary output returned in the response. | Confirm how your own application handles response bodies, logs, temporary files, and backups. |
PDFShift S3 output using filename |
The FAQ says the generated PDF is stored in PDFShift’s Amazon S3. The DPA annex says this output is in Paris, France, retained up to two days, then automatically deleted. The annex says submitted HTML, header, and footer content are not stored in S3 as part of that feature. | Check whether this option is enabled and whether its output location and retention meet your requirements. |
| Customer-owned S3 delivery | The FAQ says direct delivery to a customer’s own S3 is available for increased privacy. | Review your bucket policy, encryption, access logs, lifecycle rules, region, and retention separately. |
| Account and service metadata | The privacy policy describes personal data PDFShift processes in its role as a service provider and data controller, including account/contact, transaction, IP/log, and usage information. | Do not interpret “documents are not stored” as “no personal data is processed.” |
A page does not need to be reachable from the public internet if you send its HTML directly. That can be useful for authenticated or internal pages, but it means your application is responsible for safely collecting and transmitting the correct rendered HTML, assets, and any necessary styles. Do not put credentials or unnecessary personal data into HTML, headers, URLs, or logs.
2. How to submit HTML without publishing the page
PDFShift’s FAQ links to a guide for sending HTML directly. The examples below show the general shape of a direct HTML POST using a bearer-style API key and writing the binary response to a local PDF. Check PDFShift’s current API documentation for the exact endpoint, authentication format, and request fields for your account before using them; the research reviewed for this article does not establish those implementation details.
cURL pattern
curl --fail-with-body \
-X POST "YOUR_PDFSHIFT_HTML_CONVERSION_ENDPOINT" \
-H "Authorization: Bearer YOUR_PDFSHIFT_API_KEY" \
-H "Content-Type: application/json" \
--data '{"source":"<html><body>Confidential report</body></html>"}' \
--output report.pdf
Replace the placeholder endpoint and request format with PDFShift’s current documented values. Do not pass a filename option unless you intend to use the corresponding output-storage behavior and have reviewed its retention terms.
Python pattern
import os
import requests
endpoint = os.environ["PDFSHIFT_HTML_CONVERSION_ENDPOINT"]
api_key = os.environ["PDFSHIFT_API_KEY"]
html = "<html><body>Confidential report</body></html>"
response = requests.post(
endpoint,
headers={
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json",
},
json={"source": html},
timeout=(5, 60),
)
response.raise_for_status()
with open("report.pdf", "wb") as output:
output.write(response.content)
Install the dependency with python -m pip install requests. Set the endpoint and key in your environment; keep secrets out of source control and logs.
Node.js pattern
const endpoint = process.env.PDFSHIFT_HTML_CONVERSION_ENDPOINT;
const apiKey = process.env.PDFSHIFT_API_KEY;
if (!endpoint || !apiKey) {
throw new Error("Set the PDFShift endpoint and API key environment variables");
}
const response = await fetch(endpoint, {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
source: "<html><body>Confidential report</body></html>",
}),
signal: AbortSignal.timeout(60_000),
});
if (!response.ok) {
throw new Error(`Conversion failed with HTTP ${response.status}`);
}
const pdf = Buffer.from(await response.arrayBuffer());
await import("node:fs/promises").then(({ writeFile }) =>
writeFile("report.pdf", pdf, { mode: 0o600 })
);
These are integration patterns, not verified copy-and-paste PDFShift endpoint specifications. Before deployment, adapt the URL, authentication, JSON schema, and response handling to the live PDFShift API docs. In production, write files to an appropriately protected location or stream them to your application’s intended destination.
3. What does PDFShift’s DPA say?
The current published PDFShift Data Processing Agreement describes the customer as Controller and PDFShift as Processor for covered Company Personal Data. It says PDFShift processes that data only on documented customer instructions, except where law requires otherwise; treats it as confidential; and limits personnel access to people who need it for service delivery or their duties. The customer remains responsible for the lawfulness, accuracy, quality, and acquisition of its data and for its own compliance.
The DPA describes technical and organizational security measures appropriate to risk, including physical, logical, access, transfer, instruction, entry, availability, and separation controls. It says PDFShift will provide an overview on request and will not materially decrease overall service security during a subscription term. PDFShift’s terms separately promise commercially reasonable efforts to maintain administrative, technical, and physical measures that protect confidentiality, availability, and integrity.
These are contractual and published vendor commitments. The sources reviewed do not establish an independent audit, certification, or evidence that controls have been assessed for your particular use case. If you need assurance evidence, ask PDFShift for the applicable material and review it with your security and privacy teams.
4. Where are PDFShift files stored, and which subprocessors are named?
The current DPA annex identifies three subprocessors in the reviewed materials:
| Subprocessor | Published purpose and location | Privacy detail to consider |
|---|---|---|
| Amazon Web Services (AWS) | Storage of generated output when the customer supplies filename; S3 infrastructure in Paris, France, for up to two days, followed by automatic deletion. |
The annex says submitted HTML, header, and footer are not stored in S3 as part of this functionality. The FAQ says the generated PDF is subject to Amazon’s privacy policy. |
| OVH | PDFShift-managed servers in Gravelines and Strasbourg, France. | Confirm which service components and data categories use this infrastructure for your configuration. |
| Sentry | Error monitoring and incident diagnostics. | The DPA says HTML source, header/footer, and generated documents are not intentionally transmitted to Sentry. Incidental personal data, such as IP addresses, request metadata, or technical identifiers, may appear in technical error information. |
The DPA says transfers outside the European Economic Area require appropriate safeguards. PDFShift’s published Privacy Policy describes possible transfers and references European Commission Standard Contractual Clauses when there is no adequacy decision. That policy states it was last modified on February 28, 2022, so verify its current status and the terms attached to your account. Review the live subprocessor annex and your signed DPA rather than assuming every deployment has an identical data path.
5. Is PDFShift GDPR compliant?
PDFShift’s FAQ states that it is GDPR compliant and directs readers to its DPA and privacy policy. The DPA provides a processor framework and commitments relevant to a customer’s assessment. These vendor statements do not establish that your own processing is compliant, nor do they amount to an independent finding about the service.
Assess the full use case: what personal data is in the HTML, why you need to convert it, whether you have an appropriate legal basis and notices, who can access the output, how long your system retains it, which regions and subprocessors are involved, and whether the relevant agreement and safeguards are in place. PDFShift’s privacy policy also says that, despite its efforts, it cannot guarantee full protection against an internet leak. Build your risk decision from the current contractual documents and evidence, not a broad “GDPR compliant” label alone.
6. Breach notification and deletion: check the document and scope
The current DPA says PDFShift will notify the customer without undue delay and no later than 36 hours after PDFShift or a subprocessor becomes aware of a personal data breach affecting covered Company Personal Data. It also describes information and cooperation to support the customer’s response.
A legacy PDFShift GDPR page says customers will be notified within 72 hours after a breach is detected. The two pages use different language and triggers, and the published materials do not reconcile the difference. For a contract, rely on the version that applies to your account and confirm the operational escalation path; do not assume one deadline describes every customer relationship.
Under the current DPA, after the specified service or processing condition and a written request, PDFShift must securely return or irreversibly delete or destroy covered data without delay and no later than five business days, subject to legal retention requirements. The legacy GDPR page also mentions two years of retention for expired trial and cancelled users, but it does not provide enough context to reconcile that claim with the DPA’s deletion clause. Ask which data categories, account status, and legal obligations apply. Do not treat the two-year statement as a universal retention schedule.
7. A practical review checklist before sending confidential HTML
- Map the payload. List personal data, secrets, account identifiers, embedded images, scripts, stylesheets, and any headers or footers the conversion requires.
- Choose the input path. If the page is private, determine whether your application can provide the needed HTML directly without exposing the page publicly. Minimize included data.
- Check output settings. Confirm whether
filenameor another output-storage option is enabled. For PDFShift S3 output, review the stated Paris location and up-to-two-day retention. For customer S3, review your own bucket controls and retention. - Review current contract terms. Confirm the DPA version, controller/processor roles, documented instructions, breach notice language, deletion process, and transfer safeguards that apply to your account.
- Review subprocessors and locations. Check the current annex and identify whether your organization has region or vendor restrictions.
- Protect your side of the flow. Use TLS, store API keys in a secrets manager or environment configuration, restrict access to generated PDFs, avoid logging request bodies and response content, and remove temporary copies under your retention policy.
- Test failure handling with non-sensitive content. Check how your integration logs errors, retries requests, and handles partially written files before processing confidential documents.
- Document the decision. Record the data categories, configuration, contract version, retention choices, and owner who approved the workflow.
8. Troubleshooting privacy and integration questions
| Symptom or question | Likely cause | What to do |
|---|---|---|
| “I thought PDFShift did not store PDFs, but my integration returns a link.” | The integration may request hosted output with filename, which the FAQ distinguishes from the basic binary-response flow. |
Inspect the request options and current API docs. Confirm whether output is hosted by PDFShift or delivered to your own S3, then apply the corresponding retention review. |
| The private page cannot be fetched by the service. | A URL-based conversion cannot access a page behind your login or network boundary. | Use the documented direct-HTML method if it suits your design, and include only the necessary rendered content and assets. Confirm the API schema in current docs. |
| A response or application log contains confidential content. | Your own HTTP client, proxy, tracing, error handler, or debug logging may record request or response bodies. | Disable body logging for the conversion route, redact sensitive headers, restrict log access, and set log retention deliberately. |
| The generated PDF is missing images or styles. | HTML that works in a browser may rely on relative asset URLs, authenticated resources, or browser state that is not included in the submitted source. | Use self-contained or accessible assets as supported by the current API; avoid embedding credentials in asset URLs. Verify with synthetic content before handling real data. |
| Your retention review finds both 36- and 72-hour breach language. | The current DPA and legacy GDPR page state different notification triggers and deadlines. | Confirm the signed DPA and current operational contact path with PDFShift. Record the applicable contractual term. |
| A deletion request appears inconsistent with a two-year statement. | The DPA deletion clause and legacy page retention statement may concern different data or account states; the public material does not reconcile them. | Ask PDFShift to identify the data category and retention basis, then preserve the answer with your contract records. |
| Security review asks for proof of controls or a certification. | Published commitments are not independent audit evidence. | Request the evidence available for your service and assess its scope, date, and relevance with your security team. Do not infer a certification from the DPA. |
9. Performance, reliability, and cost considerations
The sources summarized here do not provide benchmarks, service-level figures, or enough information to quantify conversion latency or availability. Measure the workload that matters to your application with representative, non-sensitive pages, and consult the current PDFShift service terms for applicable availability and support commitments. For reliability, set client timeouts, handle non-success responses, avoid retrying non-idempotent or billable operations blindly, and decide how your system detects incomplete or invalid output.
Cost depends on your PDFShift plan and usage; the reviewed privacy materials do not establish current prices. Check the live pricing and API terms before estimating a production workload. Include your own storage, logging, review, and retention costs in the total. For confidential workloads, also weigh whether the operational convenience of hosted output is worth the additional storage path.
10. ScreenshotNeo as an alternative for screenshot jobs
PDFShift converts HTML to PDF. If the job is a website screenshot rather than a PDF conversion, ScreenshotNeo is a separate website screenshot API and MCP server. It is not a PDFShift replacement for confidential PDF generation, and its stated product facts do not establish that it meets a particular privacy or compliance requirement. Review its documentation and decide whether sending your target URL and configured credentials fits your policy.
Or skip the browser setup
For a screenshot of a page you are authorized to access, ScreenshotNeo provides a one-call API. See the ScreenshotNeo API documentation for options and current usage details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo says it accepts cookie or consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. It says bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses identify page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. These product details do not make a privacy or compliance claim about your particular page or configuration.
Sign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.
11. Frequently asked questions
Can I convert a confidential page without making it public?
PDFShift’s FAQ says yes: submit the HTML directly. Check the current API guide for the required request format and consider whether the HTML and referenced assets contain only necessary data.
Does the “not stored” statement cover PDFShift’s optional S3 output?
No. The FAQ describes filename as storing the generated PDF in PDFShift’s S3; the current DPA annex states Paris, France, and retention up to two days. Treat that as a distinct flow.
Does a DPA mean my use is automatically GDPR compliant?
No. The DPA describes PDFShift’s processor commitments, while your organization remains responsible for its own processing and compliance decisions. Assess your use and the current signed terms.
Should I rely on the legacy 72-hour notification statement?
Use the breach clause in the DPA applicable to your account as the contractual reference, and confirm the actual escalation path. The current published DPA and legacy page differ.
Is ScreenshotNeo suitable for confidential PDFs?
ScreenshotNeo is a screenshot API and MCP server, while PDFShift is discussed here for HTML-to-PDF conversion. The supplied ScreenshotNeo facts do not establish confidential PDF processing or a compliance status; review its docs and your requirements before use.
Sources
- PDFShift FAQ — basic-flow privacy statement, direct HTML submission, and optional S3 behavior.
- PDFShift Data Processing Agreement — roles, security commitments, breach and deletion clauses, and subprocessor annex.
- PDFShift Privacy Policy — controller-context data practices, transfer information, and policy date.
- PDFShift legacy GDPR page — older breach-notification and retention statements that should be reconciled with applicable terms.


