ScreenshotNeo

BlogHow-to

How to Take Website Screenshots in Dify

Use Dify’s HTTP Request node to capture raw PNGs, pass them to vision models, and avoid binary, timeout, and credential mistakes.

By the ScreenshotNeo team1 October 20267 min read

Direct answer: Add an HTTP Request node to your Dify workflow, call a screenshot endpoint that returns raw image bytes with an image/png, image/jpeg, or image/webp content type, and pass the node’s Files output to a vision-capable LLM node or file output. Start with a bounded viewport and a longer read timeout. Store credentials in a custom header or a Secret-type environment variable.

This approach follows the response handling documented by Site-Shot’s Dify guide: Dify inspects Content-Disposition and MIME type, and samples the first 1,024 bytes when the type is ambiguous. A response served as image/png with PNG bytes becomes a file variable; text, JSON, XML, or HTML becomes regular response data.

1. Build the basic Dify workflow

  1. Create a Workflow or Chatflow in Dify.
  2. Add an HTTP Request node.
  3. Choose GET.
  4. Set the screenshot service URL. The example below uses Site-Shot’s endpoint.
  5. Add the target page as a URL query parameter and begin with bounded capture settings.
  6. Configure authentication with a custom header or a secret variable.
  7. Connect the HTTP node’s Files output to a vision-enabled LLM node.

Example HTTP Request configuration

Field Value Why
Method GET Most screenshot APIs expose a single capture request.
URL https://api.site-shot.com/ Example endpoint from the reference tutorial.
Query parameter url=https://example.com Page to render.
Query parameter full_size=1 Request a full-page image when needed.
Query parameter no_ads=1 Ask the service to omit advertisements.
Query parameter no_cookie_popup=1 Ask the service to suppress cookie popups.
Response handling File/binary Ensures Dify exposes a file variable.

Use a normal screenshot first. Enable full_size only when the workflow needs the entire document, because long pages produce larger files and take longer to render.

2. Pass the screenshot to a vision model

Connect the HTTP Request node to an LLM node that supports image or file input. Map the HTTP node’s Files field to the model’s image/file input, then provide an instruction such as:

Inspect this webpage screenshot. Summarize the page’s purpose, identify the primary call to action, and list any visible errors or broken layout elements.

Do not map Response Body when the endpoint returns binary data. Response Body is for text or structured responses; the screenshot should arrive through Files.

Raw binary versus base64

Format Use it when Main risk
Raw image bytes You want Dify to create a file variable automatically. Large full-page images can exceed binary limits.
Base64 in JSON or text An API only offers encoded data. Encoding increases size and can hit text or variable limits.
Hosted file URL A downstream node accepts a URL and the URL remains valid long enough. Signed URLs may expire; the cited Dify default validity is 300 seconds.

Site-Shot cites these Dify limits: a 1 MB HTTP Request text-response limit, a 200 KB maximum for one workflow variable, a 10 MB default binary response ceiling, a 600-second maximum HTTP read timeout, and a 10-second connect-timeout ceiling. Swagger-imported API Tool nodes are cited as having a 60-second default read timeout. Dify Cloud cannot raise the 10 MB binary ceiling.

3. Configure timeouts and image size

Heavy pages, client-side rendering, advertisements, and full-page scrolling increase capture time. Set the HTTP node’s read timeout high enough for the endpoint, while keeping the connect timeout within Dify’s limit. If a full-page PNG approaches 10 MB:

  • Reduce the requested viewport width.
  • Set a bounded maximum height if the service supports it.
  • Disable full-page mode and capture the visible viewport.
  • Use JPEG or WebP when photographic content does not require PNG’s lossless output.
  • Remove unnecessary page resources through the screenshot service’s options.

Keep the request bounded in production. A smaller image transfers faster, consumes less workflow memory, and gives the model less irrelevant page area to inspect.

4. Keep credentials out of the workflow

Prefer the HTTP node’s custom authorization/header fields when the screenshot provider accepts headers. If the provider only accepts a query-string key, put that key in a Dify Secret-type environment variable and reference the secret in the request. The reference guide states that secret values are masked in workflow and request logs.

Do not put a credential in a hidden field of a published web app. Values in URLs can appear in browser history, access logs, and network traffic.

5. Complete request examples

The following examples show the same capture pattern outside Dify. They are useful for validating the endpoint before wiring it into a workflow.

cURL

curl -G "https://api.site-shot.com/" \
  --data-urlencode "url=https://example.com" \
  -d "full_size=1" \
  -d "no_ads=1" \
  -d "no_cookie_popup=1" \
  -o screenshot.png

Python

import requests

params = {
    "url": "https://example.com",
    "full_size": 1,
    "no_ads": 1,
    "no_cookie_popup": 1,
}
response = requests.get("https://api.site-shot.com/", params=params, timeout=90)
response.raise_for_status()
with open("screenshot.png", "wb") as output:
    output.write(response.content)

Node.js

const query = new URLSearchParams({
  url: 'https://example.com',
  full_size: '1',
  no_ads: '1',
  no_cookie_popup: '1'
});

const response = await fetch(`https://api.site-shot.com/?${query}`);
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const image = Buffer.from(await response.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('screenshot.png', image));

6. Troubleshooting

Symptom Cause Fix
Dify returns text instead of a file The endpoint returned HTML, JSON, XML, or an ambiguous content type. Use an endpoint that returns image bytes and sets Content-Type: image/png (or JPEG/WebP). Map Files.
Vision node receives no image The HTTP node’s Response Body was mapped. Map the HTTP node’s Files output.
Request exceeds a size limit Base64 expansion or a very tall PNG crossed Dify’s limits. Use raw binary, reduce dimensions, cap height, disable full-page mode, or use JPEG/WebP.
Read timeout The page is slow, JavaScript-heavy, or full-page capture takes too long. Increase read timeout within Dify’s ceiling; reduce page size or capture scope.
Connect timeout The service was not reached within the connection limit. Check the endpoint URL and network access; Dify’s cited connect-timeout ceiling is 10 seconds.
Blank screenshot The target blocks automated browsers, fails to load, or requires authentication. Check the target directly, provide required headers/cookies where supported, and inspect the service response.
Cookie banner covers content The capture service did not remove or interact with the consent UI. Use a service with cookie-popup handling or configure a click/script step when available.
Signed file URL expired The downstream node ran after the URL validity window. Pass the file directly between nodes or fetch it before the cited 300-second validity expires.
Code node cannot fetch the image Dify’s Code node sandbox blocks outbound network and filesystem access. Make the request in an HTTP Request node, then pass its file output onward.

7. When Browserless or visual regression tooling fits better

Dify’s Marketplace lists Browserless as a verified tool. After obtaining a Browserless token and authorizing it under Dify Tools, you can add a Browserless tool to an Agent or Workflow node. The open-source plugin exposes browserless_smartscraper, browserless_export, browserless_function, and browserless_agent. Choose this route when the workflow needs navigation, form filling, clicks, or multi-step browser control.

For scheduled captures, alerts, cross-environment comparisons, masking, CI review, or visual diffs, use dedicated visual-regression software. Diffy documents breakpoints, browser engines, delays, cookies, headers, masking, CSS/JavaScript injection, Playwright upload, CI/CD integration, and scheduled screenshot comparisons. Those requirements differ from taking one image for model inspection.

8. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. It removes cookie and consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for request options. This Dify-compatible call returns the image directly:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo includes full-page and element capture, dark mode, device presets, retina scale, custom CSS and JavaScript, clicks, selector waits, network-idle waits, request blocking, custom headers/cookies/user agents, geolocation, timezone, transparent backgrounds, resizing, configurable caching, signed links, asynchronous jobs with signed webhooks, bulk capture for 100 URLs per call, a usage API, and an OpenAPI specification.

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is available on every plan. Create a free ScreenshotNeo account.

9. Performance, reliability, and cost checklist

  • Start with viewport capture and enable full-page mode only when required.
  • Set a read timeout that covers slow pages without masking persistent failures.
  • Use raw bytes instead of base64 whenever possible.
  • Keep credentials in headers or Dify secrets.
  • Pass files directly to the vision node to avoid signed-URL expiry.
  • Cache repeat captures when the page does not need to be current.
  • For high volume, bound image dimensions and process URLs in controlled batches.
  • Choose Browserless for interactive browser actions and visual-regression software for scheduled comparisons.

FAQ

Can Dify take a screenshot without a marketplace plugin?

Yes. The HTTP Request node can call a screenshot endpoint directly; the reference tutorial explicitly describes this approach.

Why does the screenshot show as JSON?

The endpoint likely returned structured data or an error document, or its MIME type was not an image type. Inspect the status, headers, and first bytes, then use an endpoint that returns raw image bytes.

Should I send the image to an LLM as base64?

Only when the provider requires it. Raw binary is smaller and lets Dify create a file variable without consuming text or variable limits.

Can a Dify Code node download the page?

No. The cited Dify Cloud sandbox blocks outbound network and filesystem access, so perform the download in an HTTP Request node.

What if I need a PDF instead?

Use a capture service that supports PDF output or ScreenshotNeo’s capture_pdf MCP tool and PDF options such as paper size, margins, landscape mode, and page ranges.