ScreenshotNeo

BlogHow-to

How to Run a Bulk Screenshot Job in Azure Functions

Process many URLs reliably with Azure Functions, Playwright, bounded concurrency, durable storage, and asynchronous job tracking.

By the ScreenshotNeo team4 October 20269 min read

Run a bulk screenshot job in Azure Functions as an asynchronous workload: accept the URL list, assign a stable job ID, queue individual screenshot tasks, and return promptly. Bounded-concurrency workers use Playwright to capture each page and save images and per-item results to durable storage. Track and aggregate results separately. Do not keep one HTTP request open while the entire batch runs: Azure documents a 230-second maximum response time for HTTP-triggered functions, even when the configured function timeout is longer. See Azure Functions plan and timeout limits.

This guide uses Node.js and Playwright for the worker example. The architecture applies to other supported Functions languages, but browser packaging and APIs differ. The sample illustrates the capture operation; it is not a complete deployable Azure Functions project. Queue bindings, storage clients, identity, and job aggregation depend on your selected Azure resources and Functions programming model.

1. Choose the job architecture

Keep request acceptance, screenshot execution, and job status as separate steps:

  1. Accept: validate the batch, persist a job record and input manifest, and return a job ID.
  2. Dispatch: enqueue one task per URL, or a small bounded chunk. Include the job ID, item ID, URL, and capture options.
  3. Capture: workers run Playwright with explicit navigation and readiness limits.
  4. Persist: save the image and item outcome to durable storage using stable names.
  5. Aggregate: update job progress and mark it complete once every item is in a terminal state.
Approach Choose it when Consider
Queue-triggered workers Items are independent and you need straightforward distribution and retries. You must implement job-level progress and fan-in tracking.
Durable Functions fan-out/fan-in You need workflow state, coordinated fan-out, and aggregation. Orchestrators replay. Keep I/O and screenshot work in activity functions.

Durable Functions orchestrators should remain deterministic and must not perform I/O, blocking, or CPU-intensive work. Put browser navigation and capture in activities. Activity functions follow the same timeout rules as other functions. See Microsoft’s Durable Functions performance and scale guidance.

2. Accept the batch and return a job ID

The HTTP endpoint should validate and persist the request, then return 202 Accepted with a job identifier. Persist the URL manifest or a durable reference to it so a process restart does not lose the batch. A response might look like this:

HTTP/1.1 202 Accepted
Content-Type: application/json

{"jobId":"job-7f32","status":"accepted","statusUrl":"/api/jobs/job-7f32"}

Use a generated job ID and stable item IDs rather than deriving identity only from the URL: the same URL can appear more than once with different options. Store a record with creation time, requested item count, state, and manifest location. A status endpoint can report counts for pending, running, succeeded, and failed items without keeping the original request open.

3. Capture one URL with Playwright

Each worker should perform one bounded unit of work. The following Node.js function shows the core browser-to-image sequence. Supply task from your queue or activity input and saveImage from your chosen durable storage client. Configure browser installation and system dependencies for the exact Azure operating system and runtime you deploy.

const { chromium } = require('playwright');

async function captureTask(task, saveImage) {
  const browser = await chromium.launch({ headless: true });
  try {
    const context = await browser.newContext({
      viewport: { width: 1440, height: 900 },
      deviceScaleFactor: 1
    });
    const page = await context.newPage();
    page.setDefaultNavigationTimeout(45_000);

    const response = await page.goto(task.url, {
      waitUntil: 'domcontentloaded',
      timeout: 45_000
    });

    if (!response) {
      throw new Error('Navigation did not return a main-document response');
    }

    // Use a task-specific readiness rule where the target page needs one.
    if (task.readySelector) {
      await page.locator(task.readySelector).waitFor({
        state: 'visible',
        timeout: 15_000
      });
    }

    const image = await page.screenshot({
      type: 'png',
      fullPage: true,
      animations: 'disabled'
    });

    await saveImage(task.outputKey, image, 'image/png');
    return {
      jobId: task.jobId,
      itemId: task.itemId,
      state: 'succeeded',
      outputKey: task.outputKey,
      status: response.status()
    };
  } finally {
    await browser.close();
  }
}

Playwright’s Page API documents navigation and screenshot options. Adapt readiness and status handling to your target sites. A page can return an HTTP error status and still render useful content; decide whether that counts as a successful capture for your application. Always close browser resources in a finally path.

4. Set browser concurrency and retries

Start with a conservative, explicitly bounded number of active captures per worker instance. Browser processes share instance memory, CPU, and connections with other work. There is no universal safe browser count: measure representative pages in the chosen plan and deployment, then raise concurrency only while memory, CPU, latency, and failure rates remain acceptable. Azure explains shared resources and scaling in its concurrency guidance and Functions best practices.

  • Use queue or host concurrency settings to cap active work; do not start an unbounded promise for every URL in a large batch.
  • Give each item a navigation timeout and, where needed, a separate selector or application-readiness timeout.
  • Retry transient navigation or storage failures with a limit and backoff. Do not retry permanent invalid URLs indefinitely.
  • Make retries safe: use a stable output key per job item, and ensure repeated writes or status updates do not create duplicate logical results.
  • Record attempt count and terminal error category for each item.

Queues can deliver an item again after failures or lost acknowledgements, so design task processing to tolerate duplicate delivery. Treat storage failure separately from navigation failure: a successful capture that cannot be persisted is not a completed item.

5. Store results and expose job status

Store image bytes in durable object storage and keep a small per-item record with job ID, item ID, state, output key, timestamps, attempt count, and error details. Update the aggregate job state after each item completes, or compute it from item records. Serve results only through the authorization mechanism your application requires, and define an explicit retention policy for both images and manifests.

Useful terminal item states include succeeded and failed; keep transient work states such as queued and running distinct. A batch can finish with both successful and failed items. Report partial completion clearly instead of presenting the whole job as successful when some captures failed.

6. Select an Azure Functions plan and timeout

Choose a plan based on per-item duration, browser resource use, cold-start tolerance, and the required completion time. A batch should not depend on one function execution spanning the whole workload.

Plan or hosting choice Documented timeout detail Design implication
Consumption Five-minute default and ten-minute maximum function timeout. Keep each task bounded; do not make the batch one execution.
Flex Consumption and Premium Thirty-minute default and no enforced maximum execution timeout in the current scale documentation. Scale-in and platform update grace periods are also documented. Longer execution allowances do not replace checkpointing or asynchronous job tracking.
Dedicated (App Service) Timeout configuration still matters; Always On is needed for Functions on an App Service plan to run correctly. Consider when a stable or manually/autoscaled allocation suits the workload.
Container Apps Check the current Functions hosting comparison and trigger and replica behavior for the specific design. May be relevant when the browser runtime needs a custom container image.

These plan details and the separate 230-second HTTP response ceiling are described in Azure Functions scale and hosting. Configure functionTimeout deliberately for the unit of work. For long requests, Microsoft points to the Durable Functions async pattern or deferring work and returning promptly. The legacy Linux Consumption plan documentation lists a planned retirement date of 30 September 2028; recheck the lifecycle notice before choosing a plan for a new deployment.

7. Package and operate Playwright in Azure

Playwright requires a compatible browser binary and runtime libraries. Install or package the browser for the same operating system and runtime used by the deployed Function App, and verify the browser can launch there. Microsoft’s Ceruleoscope repository is an example combining Playwright with scheduled Azure Functions and Application Insights; treat it as an example to inspect, and confirm current package and browser installation behavior against current documentation.

For visual comparisons, standardize the capture environment. Playwright notes that screenshots can vary with operating system, browser and version, fonts, settings, hardware, power source, and headless mode. Keep these consistent before treating pixel differences as application changes; see the Playwright project README.

Monitor queue age, completed and failed item counts, browser startup failures, duration percentiles, memory, CPU, and storage errors. No source cited here establishes a universal throughput figure or per-instance browser count, so size from representative runs in your own deployment.

8. Troubleshooting

Symptom Likely cause Fix
The caller times out while the job is still running. The request waits for the whole batch; HTTP-triggered responses have a 230-second ceiling. Persist the job, return its ID promptly, and expose a status endpoint.
Function execution ends before capture completes. The plan’s function timeout is shorter than the task or batch. Make work per item, set a suitable timeout for that unit, and select a plan that fits. Do not use one execution for the entire batch.
Browser launch fails in Azure. Browser binaries or required system libraries are missing or incompatible with the deployed runtime. Package the matching browser and dependencies, then validate launch in the same OS and runtime configuration.
Workers become slow or fail under load. Too many simultaneous browser processes compete for instance resources. Lower the concurrency cap, observe resource use and duration, then tune with representative pages.
Captures are blank or incomplete. The capture ran before the page or required content was ready, or the target blocked automation. Use a task-specific selector or readiness condition, inspect navigation status and logs, and classify blocked targets as item outcomes.
Items appear twice or overwrite unexpected output. Duplicate queue delivery or unstable output identity. Use stable job and item IDs, deterministic output keys, and idempotent status updates.
Capture succeeded but the item is marked failed. Image persistence failed after the browser work completed. Record storage errors distinctly and retry persistence safely using the same item identity.
Visual diffs change between runs. Browser, OS, fonts, rendering settings, or headless environment changed. Pin and standardize the capture environment and browser version.

9. Cost and reliability notes

Azure cost depends on the selected hosting plan, execution and resource use, queue and storage choices, and how long outputs are retained. The research sources do not establish a cost estimate for this workload. Measure representative batches and include storage and monitoring in your estimate. Keep images only as long as the product needs them.

Reliability comes from treating each URL as an independent, observable task: persist inputs before dispatch, checkpoint outcomes, cap retries, distinguish transient from permanent errors, and make writes safe to repeat. A job status endpoint should expose partial failures so callers can retry only the items that need attention.

Or skip the browser setup

If you want the screenshot API to handle browser capture, ScreenshotNeo provides one GET request for an image or PDF. See the ScreenshotNeo API documentation for request options. For a single capture:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

For bulk work, call the API per URL from your own queue and keep the same asynchronous job and result tracking pattern. ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, and failed loads are not billed. Its MCP server lets AI agents take screenshots, and bulk capture supports up to 100 URLs per call. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Learn more about ScreenshotNeo and its API options, then sign up free for 1,000 screenshots a month, with no card.

FAQ

Should I use a queue or Durable Functions?

Use a queue for independent tasks when you can implement job aggregation separately. Use Durable Functions when workflow tracking and fan-out/fan-in coordination are central requirements.

Can I use networkidle as the readiness rule?

Choose readiness based on the target page. Some sites keep network connections open, so waiting for a specific visible element or a deliberate delay may be more appropriate.

Does a successful navigation mean the page is suitable for a screenshot?

No. Navigation completion and application readiness are separate. Define what content must be present, and record HTTP status and capture outcome independently.

How many URLs can one function process?

There is no universal number. It depends on page weight, browser resource use, plan, configured timeout, and required completion time. Measure in the deployed environment and bound concurrency.