ScreenshotNeo

BlogHow-to

How to Use Apify Screenshots to Archive Web Pages as Evidence

Capture a public page with an Apify screenshot Actor, preserve its artifact and context, and understand what a screenshot can—and cannot—show as evidence.

By the ScreenshotNeo team4 October 202611 min read

To archive a public web page with an Apify screenshot Actor, submit its HTTP(S) URL, choose an image or PDF format, enable full-page capture if the entire scrollable page matters, run the Actor, and save the resulting file with its capture context. A screenshot records one rendered view at one moment. By itself, it does not establish who published the page, prove that the capture is complete or authentic, or guarantee legal admissibility.

Apify Actors are packaged tools that run on the Apify platform. Their inputs and outputs differ, so check the selected Actor’s own input schema and output documentation before building a repeatable workflow. The Apify-maintained Website Screenshot Generator input schema, for example, describes a list of URLs, PNG or PDF output, viewport and wait settings, scrolling, and selector hiding. The full-page example for another Actor lists output metadata such as requested and final URL, status, dimensions, truncation, and capture time. Those options and fields are Actor-specific.

1. Choose an Actor and define what you need to preserve

Open an Apify screenshot Actor that accepts page URLs and produces the artifact you need. Before capturing, decide whether you need the visible viewport or the full page, and whether an image or PDF best fits your archive. If the Actor accepts a list of URLs, you can submit a batch; verify how it reports individual successes and failures.

  1. Choose the Actor and read its current input schema, output description, pricing, and retention details.
  2. Prepare the public HTTP(S) URL. Record the URL you intend to capture before redirects.
  3. Select PNG or PDF if offered. Choose full-page capture when below-the-fold content is relevant.
  4. Set the viewport and wait behavior deliberately. Keep the settings for later captures if you need comparable records.
  5. Run the Actor and inspect the result for every requested URL.
  6. Download or copy the artifact into storage you control, then preserve the metadata alongside it.

Apify documents Actors as runnable tools whose input can be configured in the Console or supplied as JSON through the API. The exact capture fields are not universal. The following code examples use the documented input schema for the Apify-maintained apify/screenshot-url Actor; check its current schema before relying on them. The Actor page describes screenshots stored in a key-value store and supports a list of URLs, with PNG as the default format and PDF also listed.

2. Run the Apify Actor from code

These examples submit two URLs to apify/screenshot-url. They request PNG output, a 1280-pixel viewport, and full-page scrolling, then wait for the synchronous dataset-items response. Set APIFY_TOKEN to an Apify API token with permission to run the Actor. The response contains Actor output records; inspect those records for the artifact key or link exposed by the current Actor, and download and retain the actual file. A dataset response is not necessarily the image bytes themselves.

cURL

export APIFY_TOKEN='YOUR_APIFY_TOKEN'
curl --fail-with-body --silent --show-error \
  -X POST \
  'https://api.apify.com/v2/acts/apify~screenshot-url/run-sync-get-dataset-items?token='"$APIFY_TOKEN" \
  -H 'Content-Type: application/json' \
  --data '{
    "urls": [
      "https://example.com/",
      "https://example.org/"
    ],
    "format": "png",
    "waitUntil": "domcontentloaded",
    "delay": 1000,
    "viewportWidth": 1280,
    "scrollToBottom": true,
    "delayAfterScrolling": 2500,
    "waitUntilNetworkIdleAfterScroll": false
  }' \
  -o results.json

Python

import json
import os
import requests

api_token = os.environ["APIFY_TOKEN"]
endpoint = (
    "https://api.apify.com/v2/acts/"
    "apify~screenshot-url/run-sync-get-dataset-items"
)
payload = {
    "urls": ["https://example.com/", "https://example.org/"],
    "format": "png",
    "waitUntil": "domcontentloaded",
    "delay": 1000,
    "viewportWidth": 1280,
    "scrollToBottom": True,
    "delayAfterScrolling": 2500,
    "waitUntilNetworkIdleAfterScroll": False,
}
response = requests.post(
    endpoint,
    params={"token": api_token},
    json=payload,
    timeout=600,
)
response.raise_for_status()
records = response.json()
with open("results.json", "w", encoding="utf-8") as output:
    json.dump(records, output, indent=2)
print(json.dumps(records, indent=2))

Node.js

const token = process.env.APIFY_TOKEN;
if (!token) throw new Error('Set APIFY_TOKEN first');

const endpoint = new URL(
  'https://api.apify.com/v2/acts/apify~screenshot-url/run-sync-get-dataset-items'
);
endpoint.searchParams.set('token', token);

const response = await fetch(endpoint, {
  method: 'POST',
  headers: { 'content-type': 'application/json' },
  body: JSON.stringify({
    urls: ['https://example.com/', 'https://example.org/'],
    format: 'png',
    waitUntil: 'domcontentloaded',
    delay: 1000,
    viewportWidth: 1280,
    scrollToBottom: true,
    delayAfterScrolling: 2500,
    waitUntilNetworkIdleAfterScroll: false
  }),
  signal: AbortSignal.timeout(600000)
});

if (!response.ok) {
  throw new Error(`Apify request failed: ${response.status} ${await response.text()}`);
}
const records = await response.json();
console.log(JSON.stringify(records, null, 2));

Use a secret manager or environment variable for the token; do not commit it to source control or include it in a public archive. A synchronous run is convenient for small jobs, but can take a long time or time out for larger batches. Apify’s general Actor input and output documentation describes configuring an Actor in Console or passing its JSON input through the API. For large or long-running work, use the asynchronous run flow and retrieve the resulting dataset or stored artifact after completion.

3. Select capture settings carefully

Setting When it helps Record or watch for
URL list Capturing several public pages in one run, if that Actor accepts multiple URLs. Keep per-URL success, failure, and error details. A partial batch is not a complete archive.
Format PNG preserves a raster view; PDF may fit a document-oriented workflow. Both are static outputs, not replayable copies of the page’s resources or interactions.
Viewport width Making the rendering context explicit and repeatable. A different viewport can change responsive layout and visible content. Save its dimensions.
Wait condition domcontentloaded can capture sooner; load or a network-idle option may suit pages that populate later. Network-idle waits can be slow or unsuitable for pages with persistent network activity. Actor-specific supported values vary.
Delay Allowing a client-rendered page additional time to settle. A fixed delay can still be too short or waste time. Use the smallest delay that produces the needed state and note it.
Scroll before capture Triggering lazy-loaded content on long pages. Scrolling and post-scroll waiting are Actor-specific. Check for truncation or missing content in the output.
Selectors to hide Removing page elements that obscure the material being documented, where the Actor supports it. Hiding content alters the rendered record. Preserve the original capture and document any alteration if evidence context matters.
Proxy configuration Only when the Actor supports it and the use is permitted. It can affect what the server returns. Save relevant configuration and do not treat a different network path as proof of authenticity.

For the documented apify/screenshot-url schema, waitUntil options include load, domcontentloaded, networkidle2, and networkidle0; its schema lists a delay, viewport width, scroll-to-bottom option, post-scroll delay or network-idle wait, proxy configuration, and selectors to hide. It lists PNG and PDF formats. Consult the current input schema for accepted values and limits; do not assume those names exist on another Actor.

4. Preserve the artifact and its context

Save the image or PDF outside temporary run storage if you need it for long-term retention. The reviewed Actor documentation says unnamed run storage is removed according to the applicable plan’s retention period. Check the selected Actor and plan’s current storage and retention rules, then copy the artifact to storage with a retention period that matches your use.

Store a sidecar record with the artifact. Include, where available:

  • Requested URL and final URL after redirects.
  • Capture timestamp, including timezone or a clear UTC convention.
  • Page title and HTTP status.
  • Output format, image dimensions, viewport/device settings, and whether capture was full-page or truncated.
  • Actor name and version/build if exposed, plus the input settings used.
  • Run status, error details, and any limitations, selector hiding, or other adjustments.
  • A stable artifact location and, if independently calculated, a cryptographic hash of the saved file.

Not every Actor returns all these fields. A separately maintained evidence-pack Actor documents a bundle containing a PDF, full-page PNG, visible text, optional rendered HTML, URL and status metadata, timestamp, and SHA-256 fingerprint. Those are capabilities of that particular Actor, not a platform-wide promise. A timestamp or hash can support reproducibility and change detection; neither proves the original publisher or guarantees admissibility.

5. What a screenshot can establish—and what it cannot

A screenshot is a visual record of a browser-rendered state at capture time. A PDF is also a static output. Neither is a replayable copy of every resource, script, response, or interaction behind the page. A saved image does not, on its own, prove who published the content, what the page showed earlier, that nothing was omitted, or that a court will admit it.

Automated browser output can differ from a reader’s local browser. Sites may vary content by location, account state, cookies, viewport, or time. Some sites block automated browsers; the reviewed Actor does not log in, solve captchas, or bypass bot protection, and a blocked page or failed run is not a successful capture. Retain the error or block result as such rather than describing it as an archived page.

If the purpose is a formal legal or compliance record, follow the applicable evidence-handling requirements and obtain qualified advice about the process. The research sources do not establish that an Apify screenshot alone is sufficient evidence. For a broader preservation workflow, an ABI industry review describes Pagefreezer WebPreserver exports including image PDF, searchable PDF with OCR, WARC, MHTML, JPEG, and a collection report. That is a specialized comparison point, not a requirement for every screenshot; the cited review’s claims should be considered in that context.

6. Batch captures, reliability, and cost

Batch capture is useful for a set of pages, but treat each URL as its own outcome. Keep a manifest of requested URLs and reconcile it against the Actor’s output records. Retry only failed or incomplete items, and preserve each attempt’s result so retries do not erase the record of a block or timeout. For repeated captures, use stable settings and an explicit schedule, and compare artifacts alongside metadata rather than relying on filenames alone.

Full-page captures can be tall and may be truncated. One marketplace Actor’s documentation reports a maximum full-page screenshot height of 16,384 pixels; this is an Actor-specific limit and may change. Check the particular Actor’s output fields for a truncation indicator and inspect the actual file. A PDF should likewise be checked for missing pages or clipped content.

Performance depends on page load behavior, selected wait conditions, full-page scrolling, batch size, and the Actor’s run resources. Network-idle waits and large pages can extend run time. For routine records, a modest concurrency and measured wait strategy make failures easier to diagnose; do not select a delay merely to make a capture appear complete.

Actor and platform charges vary by Actor, usage model, and plan. The research does not establish a general Apify price for screenshot archiving. Check current pricing before running a large batch, estimate the number of URLs and retries, and account for storage and retention separately. Pricing and Actor options can change.

7. Troubleshooting

Symptom Likely cause What to do
Run fails or reports an error for one URL Bad URL, page/network failure, unsupported page behavior, or Actor error. Check the URL and per-item error, retry once if appropriate, and keep the failed status. Do not label it a successful capture.
Image shows a CAPTCHA or bot-check page The site blocks automated browsing. Do not claim the intended page was archived. Preserve the returned status and follow the site’s access rules.
Only the top of the page appears Viewport capture was used, scrolling was disabled, or the Actor truncated the page. Enable its full-page/scroll option, inspect the output’s full-page and truncation fields, and compare the artifact dimensions.
Images or sections are missing Lazy loading or client rendering had not completed before capture. Use the Actor’s supported scroll and wait controls, then recapture. Record changed timing settings.
Layout differs from the expected view Viewport, device settings, redirect, region, cookies, or time-dependent content changed the rendered state. Record final URL and viewport; repeat with consistent settings where possible. Describe the variation instead of silently editing the image.
API returns an authorization error Missing, invalid, or insufficiently scoped token. Set a valid Apify API token with run permission and keep it secret. Confirm the API request uses the correct Actor identifier.
API request times out The synchronous request outlasted the client timeout, often due to a large batch or slow page. Use an asynchronous Actor run and retrieve results when it finishes; reduce batch size and check the run status.
Artifact link no longer works Run storage or a temporary file link expired under its retention rules. Download the artifact promptly and copy it to durable storage under your retention policy.

8. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF. Its clean-shot steps accept cookie/consent banners and remove 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and responses report page verdict and billing status in headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents and MCP clients.

For an archive, retain the returned artifact together with its requested URL, capture time, and any relevant context. A ScreenshotNeo capture is still a static screenshot or PDF, not proof of legal admissibility or a full replayable web archive. See the ScreenshotNeo API documentation for options and response details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 screenshots. Sign up for ScreenshotNeo’s free plan.

FAQ

Can I take full-page screenshots of many URLs at once?

Some Actors accept a URL list and support full-page capture. Confirm both capabilities in the specific Actor’s current input schema, then check success, failure, and truncation for every result.

How do I convert a URL to PDF?

Select an Actor that supports PDF output, provide the URL, and choose PDF in its input. The apify/screenshot-url schema lists PDF as an available format; other Actors may not.

Does a screenshot prove what a page said?

It records a rendered view produced during a particular capture. It does not alone prove authorship, historical state, completeness, or legal admissibility.

Should I keep only the image?

No. Keep the file with its source and final URL, timestamp, capture settings, status, and any limitations. Copy it to storage whose retention you control.