ScreenshotNeo

BlogHow-to

How to Convert Web Pages to Video with an API

Learn how to render JavaScript pages, capture browser video, encode MP4, and choose between Playwright, Puppeteer, and cloud video APIs.

By the ScreenshotNeo team1 October 202610 min read

To convert a web page into an MP4, open the URL in a real browser, wait until its data and visual assets are ready, record the browser output, and encode the recording in the format your destination requires. JavaScript-heavy pages cannot be captured reliably with an HTTP request alone.

This guide shows how to record a JavaScript-rendered page with Playwright or Puppeteer, how to use a cloud rendering API, how to control timing and output quality, and how to troubleshoot the failures that make web-to-video jobs unreliable.

Choose an implementation

Approach Best for Output and operations
Playwright Self-hosted browser recordings with explicit viewport and video size WebM examples; the file is finalized when the page or browser context closes
Puppeteer Chrome-specific recording and DevTools Protocol control Page.record() can output MP4; screencast defaults to WebM/VP9 at 30 FPS and needs ffmpeg
Shotstack Cloud composition and media rendering JSON edit API with an HTML5/CSS3/JS asset; asynchronous rendering
Creatomate Putting page content into a designed template REST renders from a template or RenderScript, with polling or webhooks
Browserless Managed browser rendering before another capture or composition stage Returns fully rendered HTML from a URL or HTML input

For pages whose visible content is generated by JavaScript, use a real browser. A plain HTTP client only receives the initial HTML and misses data loaded after scripts run. Browserless documents this distinction for its rendered-content API, and Playwright and Puppeteer expose browser capture controls.

Pipeline: URL to video

  1. Render: launch a pinned browser engine and navigate to the URL.
  2. Prepare: set the viewport, timezone, locale, authentication state, and any required cookies or headers.
  3. Wait: wait for navigation, a selector, application data, fonts, images, and animations to reach the state you intend to record.
  4. Capture: record a browser video or a sequence of frames for the required duration.
  5. Encode: produce MP4, WebM, or another container and codec accepted by your destination.
  6. Deliver: upload the file or return it from your API, and report failures with enough metadata to reproduce the run.

There is no universal readiness signal. networkidle can be misleading on pages with analytics, polling, or open sockets. Prefer an application-specific selector or data attribute, then add a short delay only when transitions or late fonts require it.

Record a page with Playwright

Playwright records video in a browser context. Its documentation states that the video is written when the browser context closes and supports an explicit recording size. Install the package and browser first:

npm install playwright
npx playwright install chromium

Create capture.mjs:

import { chromium } from 'playwright';

const browser = await chromium.launch();
const context = await browser.newContext({
  viewport: { width: 1280, height: 720 },
  recordVideo: {
    dir: 'recordings',
    size: { width: 1280, height: 720 }
  },
  colorScheme: 'light',
  deviceScaleFactor: 1
});

const page = await context.newPage();
await page.goto('https://example.com', { waitUntil: 'domcontentloaded', timeout: 60000 });
await page.locator('body').waitFor({ state: 'visible', timeout: 30000 });
await page.waitForLoadState('networkidle').catch(() => {});
await page.waitForTimeout(1000);

// Record a five-second interaction or animation window.
await page.waitForTimeout(5000);

await page.close();
await context.close();
await browser.close();
console.log('Video finalized in recordings/');

Run it with node capture.mjs. The resulting file is normally WebM. The filename is assigned by Playwright, so inspect the page’s video path before closing it if your application needs to rename or upload the file:

const videoPath = await page.video().path();
// Close the context before reading or uploading the completed file.
await context.close();
console.log(videoPath);

Use the same browser context for pages that share login state. For a private page, load a storage state created by a secure login flow instead of placing credentials in source code. Keep recording directories outside publicly served paths.

Convert Playwright WebM to MP4

Use ffmpeg when the consumer requires MP4. The exact codec settings depend on your compatibility target; H.264 video with AAC audio is widely accepted:

ffmpeg -i recordings/VIDEO.webm -c:v libx264 -pix_fmt yuv420p -c:a aac -movflags +faststart output.mp4

A browser recording may contain no audio. In that case, omit -c:a aac or add an audio track deliberately; do not assume a silent track exists.

Record with Puppeteer

Puppeteer’s current documentation describes Page.record() as an experimental Chrome DevTools Protocol API that outputs an MP4 stream. Its screencast API records WebM with VP9 at 30 FPS by default and requires ffmpeg. The older Page.screencast() API is marked obsolete, so evaluate Page.record() first for new work.

npm install puppeteer
import puppeteer from 'puppeteer';

const browser = await puppeteer.launch({ headless: 'new' });
const page = await browser.newPage();
await page.setViewport({ width: 1280, height: 720, deviceScaleFactor: 1 });
await page.goto('https://example.com', { waitUntil: 'networkidle2', timeout: 60000 });
await page.waitForSelector('body', { visible: true, timeout: 30000 });
await new Promise(resolve => setTimeout(resolve, 1000));

// Page.record is experimental and depends on the installed Chromium version.
if (typeof page.record !== 'function') {
  throw new Error('This Puppeteer/Chromium version does not expose Page.record()');
}
const recording = await page.record({
  path: 'output.mp4',
  speed: 1
});
await new Promise(resolve => setTimeout(resolve, 5000));
await recording.stop();
await browser.close();

Pin Puppeteer and Chromium versions in production. Experimental protocol methods can change with browser revisions. If you use screencast instead, plan for WebM output and an ffmpeg conversion step.

Python option with Playwright

The Python package exposes the same browser-context video model. Install it and its browser:

pip install playwright
playwright install chromium
from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch()
    context = browser.new_context(
        viewport={"width": 1280, "height": 720},
        record_video_dir="recordings",
        record_video_size={"width": 1280, "height": 720},
    )
    page = context.new_page()
    page.goto("https://example.com", wait_until="domcontentloaded", timeout=60000)
    page.locator("body").wait_for(state="visible", timeout=30000)
    try:
        page.wait_for_load_state("networkidle", timeout=10000)
    except Exception:
        pass
    page.wait_for_timeout(1000)
    page.wait_for_timeout(5000)
    page.close()
    context.close()  # finalizes the video
    browser.close()

Cloud rendering APIs

Shotstack

Shotstack accepts a JSON edit and renders media through its Edit API. Its Html5Asset renders full HTML5, CSS3, and JavaScript, which makes it a direct fit when a webpage is one asset in a larger composition. The documented process validates the edit, downloads and caches assets, preprocesses media, renders, and stores the final file.

Keep the API key on your server, submit the edit, store the returned render identifier, and poll the render status or consume the documented completion mechanism. Put only publicly reachable assets in the edit unless the service supports the authentication method your page requires.

Creatomate

Creatomate creates video, image, or GIF renders from a template or a JSON RenderScript. Use it when the webpage needs to be mapped into a designed composition with titles, timing, overlays, or multiple sources rather than captured as an unmodified browser viewport. Submit the render from a server, then poll or use its webhook callback to obtain the finished file URL.

Browserless

Browserless is useful as the browser-rendering layer. Its Content API accepts a URL or HTML and returns fully rendered HTML, including JavaScript-generated content. Pass that rendered result to a capture or composition stage when you need managed browser infrastructure but want to control the final media pipeline yourself.

Settings that determine quality and reproducibility

Setting What to decide Failure it prevents
Viewport Set width and height to the delivery dimensions; use the same device scale factor on every run. Unexpected responsive layouts and blurry scaling
Readiness Wait for a page-specific selector, loaded data, fonts, images, and animation state. Skeleton screens, missing fonts, and partial charts
Duration Define the exact recording window and whether timers should be frozen or allowed to run. Different clips on every retry
Frame rate Choose the rate required by the destination and keep it consistent through encoding. Judder or unnecessarily large files
Codec/container Use MP4/H.264 for broad compatibility or WebM/VP9 where supported. Upload rejection or browser playback failures
Fonts and assets Make cross-origin images, scripts, video, and font files reachable to the rendering browser. Blank boxes, fallback fonts, and missing media
Authentication Provide cookies, headers, or a storage state through a secret-managed server process. Login redirects and unauthorized responses
Geography and time Fix timezone, locale, geolocation, browser version, and service region where possible. Date, currency, consent, and regional content drift

Reliability checklist

  • Pin the browser and ffmpeg versions.
  • Use a per-job temporary directory and clean it after upload.
  • Set navigation, selector, recording, and overall job timeouts.
  • Retry transient navigation or provider errors with exponential backoff, but do not blindly retry deterministic authorization or bot-check failures.
  • Record the URL, viewport, browser version, wait condition, start time, render provider, output format, and error classification.
  • Use an idempotency key or your own job identifier so a retry cannot create duplicate publishing records.
  • Validate the finished file before delivery: container, duration, dimensions, codec, and nonzero size.
  • Check that you have permission to reproduce the page, fonts, images, video, and other media.

Performance and cost

Browser startup is often the expensive part of a self-hosted job. Reuse a browser process when isolation allows it, create a fresh context per job, and avoid loading resources that cannot affect the recording. Caching static assets can reduce latency, while overly broad blocking can remove fonts, CSS, or application data and produce a different page.

Cloud services usually charge for rendering, output duration, or both according to their current pricing. Verify current limits and prices before publishing an estimate. Self-hosting shifts the cost to browser CPU, memory, storage, egress, and ffmpeg time. Measure your own pages: a page with WebGL, large images, long animations, or many third-party requests behaves differently from a static document.

Troubleshooting

Symptom Likely cause Fix
Video shows a blank or loading page Capture started before application data or fonts arrived. Wait for a page-specific ready selector and verify it in a headed run.
JavaScript content is missing The implementation fetched HTML without executing a browser. Use Playwright, Puppeteer, Browserless, or an HTML5-capable rendering API.
Video is cut off or corrupt The page/context was not closed, so Playwright did not finalize the file. Close the page and context in a finally block before reading or uploading.
Output is WebM but the platform requires MP4 The capture API’s default container differs from the destination. Transcode with ffmpeg and validate H.264/AAC compatibility.
Layout changes between runs Responsive viewport, timezone, locale, random data, or live timers differ. Fix viewport and environment values; stub or freeze nondeterministic data where permitted.
Images or video are missing Cross-origin requests, signed URLs, or resource blocking prevent access. Allow the required origins, refresh expiring URLs, and inspect browser network logs.
Login page is recorded Cookies or storage state were not supplied, or the session expired. Create a fresh authenticated context and keep credentials server-side.
Bot-check or CAPTCHA appears The site challenged the rendering browser. Do not attempt to bypass access controls; use an authorized session or obtain permission from the site owner.
Cloud job never completes Polling code ignores failed terminal states or the provider callback cannot reach your endpoint. Handle every terminal status, verify webhook authentication, and retain the provider job ID.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server when you need a reliable still image of the page before assembling a video or when a frame sequence is sufficient. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. Its cleanup steps accept cookie and consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

See the ScreenshotNeo API documentation for all options. This runnable cURL request captures a frame:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

It also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. You can set full-page capture, a CSS element, dark mode, device or custom viewport, retina scale, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agent, authorization, timezone, geolocation, transparency, resizing, caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, and usage reporting.

The Free plan includes 1,000 screenshots each month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account.

FAQ

Can an HTTP request alone make an MP4?

No. An HTTP request can download source HTML, but a browser must execute JavaScript and paint the page before you can record what a visitor sees.

Should I capture frames or record a browser video?

Record a browser video for continuous motion and interaction. Capture frames when you need deterministic slides, screenshots, or a composition pipeline that controls timing itself.

Which format should I publish?

Use MP4/H.264 when compatibility is the priority. Keep WebM/VP9 when your destination supports it and you want to avoid a transcode.

How do I make a recording repeatable?

Fix the browser version, viewport, scale factor, timezone, locale, authentication state, data snapshot, wait condition, duration, and output settings. Live pages can still change outside your control.

Can I record any website?

No. Login walls, bot defenses, cross-origin media, timers, animations, WebGL, changing data, and access permissions can change or prevent the result. Capture only content you are authorized to reproduce.