ScreenshotNeo

BlogEngineering

WebRTC Screen Capture API

Learn how to capture a user-selected screen, tab, or window with getDisplayMedia(), preview it, send it over WebRTC, and handle permissions safely.

By the ScreenshotNeo team1 October 202610 min read

Use navigator.mediaDevices.getDisplayMedia() from a user-initiated action in a secure context. The browser opens its own picker, the user chooses a display surface, and the promise resolves to a MediaStream. You can preview that stream, record it, or add its tracks to an RTCPeerConnection for WebRTC transmission.

A web page cannot silently select a monitor or remember display-capture permission for a later session. The browser must let the user choose a surface every time. Capture options can influence the resulting track, but they cannot be used to remove choices from the picker. See the W3C Screen Capture specification and MDN API documentation.

Minimal screen-sharing example

Save this as index.html and serve it from https:// or from http://localhost. Opening the file directly with file:// is not a reliable deployment method because display capture requires a secure context.

<!doctype html>
<html lang="en">
<head>
  <meta charset="utf-8">
  <meta name="viewport" content="width=device-width, initial-scale=1">
  <title>WebRTC screen capture</title>
  <style>
    body { font: 16px system-ui, sans-serif; max-width: 900px; margin: 2rem auto; padding: 0 1rem; }
    video { width: 100%; background: #111; border-radius: 8px; }
    button { margin: .5rem .5rem .5rem 0; padding: .6rem 1rem; }
    #status { min-height: 1.5rem; }
  </style>
</head>
<body>
  <h1>Share your screen</h1>
  <button id="start">Start sharing</button>
  <button id="stop" disabled>Stop sharing</button>
  <p id="status" role="status">Not sharing</p>
  <video id="preview" autoplay playsinline muted></video>

  <script>
    const startButton = document.querySelector('#start');
    const stopButton = document.querySelector('#stop');
    const preview = document.querySelector('#preview');
    const status = document.querySelector('#status');
    let stream;

    startButton.addEventListener('click', async () => {
      try {
        stream = await navigator.mediaDevices.getDisplayMedia({
          video: true,
          audio: false
        });

        preview.srcObject = stream;
        startButton.disabled = true;
        stopButton.disabled = false;
        status.textContent = 'Sharing is active';

        const [videoTrack] = stream.getVideoTracks();
        videoTrack.addEventListener('ended', stopSharing);
      } catch (error) {
        status.textContent = `${error.name}: ${error.message || 'capture was not started'}`;
      }
    });

    stopButton.addEventListener('click', stopSharing);

    function stopSharing() {
      if (stream) {
        for (const track of stream.getTracks()) track.stop();
        stream = undefined;
      }
      preview.srcObject = null;
      startButton.disabled = false;
      stopButton.disabled = true;
      status.textContent = 'Not sharing';
    }
  </script>
</body>
</html>

Run a local server, for example python3 -m http.server 8000, then visit http://localhost:8000. The call must remain directly reachable from the click handler; browsers may reject calls made after an unrelated asynchronous step or without recent user activation.

How the WebRTC screen-capture flow works

  1. The user clicks a clearly labeled share button.
  2. Your code calls getDisplayMedia().
  3. The browser presents its picker for a tab, window, or display, depending on the browser and platform.
  4. The user approves or cancels. Approval returns a MediaStream.
  5. You attach the stream to a <video> element, a recorder, or an RTCPeerConnection.
  6. You listen for the track’s ended event because the user can stop sharing from browser chrome at any time.

Send the captured stream over WebRTC

Screen capture only produces media. WebRTC still needs signaling to exchange an offer, answer, and ICE candidates between peers. The following sender-side code shows where the captured video track enters the connection; your application must transport the signaling messages through its own server or channel.

const peer = new RTCPeerConnection({
  iceServers: [{ urls: 'stun:stun.l.google.com:19302' }]
});

const stream = await navigator.mediaDevices.getDisplayMedia({ video: true });
const [videoTrack] = stream.getVideoTracks();
peer.addTrack(videoTrack, stream);

videoTrack.addEventListener('ended', () => {
  // Notify the remote peer or update your application state.
  console.log('The user stopped sharing');
});

peer.onicecandidate = event => {
  if (event.candidate) sendSignal({ candidate: event.candidate });
};

const offer = await peer.createOffer();
await peer.setLocalDescription(offer);
sendSignal({ description: peer.localDescription });

// Call this when a remote description arrives:
async function acceptAnswer(description) {
  await peer.setRemoteDescription(description);
}

On the receiving side, listen for track and attach the incoming stream to a video element:

const receiver = new RTCPeerConnection();
const remoteVideo = document.querySelector('#remote-video');

receiver.addEventListener('track', event => {
  remoteVideo.srcObject = event.streams[0];
});

Capture options and constraints

The options object controls the requested media types and some track behavior. It does not silently choose a source for the user.

Option Purpose Practical note
video Required in normal screen-capture calls; can be true or a constraints object. The browser still decides which surfaces appear in its picker.
audio Requests system or tab audio when the browser and selected surface support it. Audio support varies. Treat it as optional and verify the resulting stream.
cursor Requests whether the captured cursor is always shown, never shown, or shown only while moving. Support and exact behavior depend on the browser.
displaySurface Expresses a preference such as monitor, window, or browser. It is not a permission bypass and cannot force a source.
preferCurrentTab Can help a browser present the current tab as a convenient choice. The user still makes the final selection.
selfBrowserSurface Controls whether the current tab may be offered as a surface where implemented. Use this to reduce accidental hall-of-mirrors sharing, but do not assume universal support.

After capture starts, inspect the actual settings rather than assuming that requested values were honored:

const track = stream.getVideoTracks()[0];
console.log(track.getSettings());
console.log(track.getCapabilities());

You may apply supported track constraints later, for example to request a frame size or frame rate:

await track.applyConstraints({ frameRate: { ideal: 15, max: 30 } });

Constraint application can fail with OverconstrainedError when the selected surface cannot satisfy the request. Use ideals and maximums for adaptive behavior, and keep a fallback path.

Capturing audio safely

const stream = await navigator.mediaDevices.getDisplayMedia({
  video: true,
  audio: true
});

const hasAudio = stream.getAudioTracks().length > 0;
console.log(hasAudio ? 'Audio captured' : 'Video only');

Requesting audio: true does not guarantee an audio track. Availability depends on the browser, operating system, selected surface, and user choice. Explain exactly what may be shared, give audio its own visible control where appropriate, and do not infer that silence means the API failed.

Permissions Policy and embedded applications

The top-level document normally has display-capture available through the default self allowlist. If you use a restrictive Permissions-Policy header or an iframe, explicitly permit the feature.

Permissions-Policy: display-capture=(self "https://app.example")
<iframe
  src="https://app.example/share.html"
  allow="display-capture">
</iframe>

Policy permission is only a technical prerequisite. It does not replace the browser picker or the user’s consent. A disabled policy commonly produces NotAllowedError. See MDN’s display-capture Permissions Policy reference.

Security and privacy checklist

  • Use HTTPS in production and a secure local development origin.
  • Start capture only from a recent, intentional user action.
  • Show a persistent sharing indicator in your application and explain whether video, audio, or both are sent.
  • Provide a stop button and respond to the browser’s native stop action.
  • Do not promise that a page can capture without asking or select a source silently.
  • Warn users that private messages, credentials, notifications, and other cross-origin content visible on the selected surface can be exposed.
  • Stop tracks when the call ends: stream.getTracks().forEach(track => track.stop()).
  • Keep signaling authenticated and authorize who may receive the media.

The W3C specification states that the user agent must let the end user choose a display surface from all available choices every time, and must not use media constraints to limit that choice. The specification is a Working Draft dated 27 August 2026 and warns that it is incomplete and subject to change, so verify the status when publishing production guidance.

Browser support and deployment decisions

MDN currently labels getDisplayMedia() as limited availability and not Baseline. Do not claim identical support across browsers. Check the current compatibility table for every browser and version you serve, especially for system audio, tab audio, cursor controls, and picker behavior.

Decision What to check
Surface choice Whether the target browser offers tabs, windows, displays, or platform-specific variants.
Audio Whether the selected surface can provide audio and whether the returned stream contains an audio track.
Policy Whether the response header and iframe allow attribute permit display-capture.
UX How the browser communicates active sharing and how users stop it.
Fallback What your application does when capture is unsupported, denied, or stopped.

Performance and reliability

  • Choose a sensible frame rate. Presentation sharing often needs less bandwidth than video. A 15–30 fps target can reduce CPU and network use when supported.
  • Adapt to the track. Use getSettings() and monitor WebRTC sender statistics instead of assuming a fixed resolution.
  • Handle renegotiation. If users stop and restart sharing, replace or add tracks and renegotiate according to your peer-connection design.
  • Expect interruption. Laptop sleep, browser controls, display changes, and OS permissions can end a track. Treat ended as a normal state transition.
  • Keep the preview lightweight. A muted local preview avoids sending the stream back through your network and helps users confirm what is visible.
  • Measure the real stream. Use RTCRtpSender.getStats() to watch frames sent, bytes sent, packet loss, and available bitrate.

Troubleshooting

Symptom or error Likely cause Fix
NotAllowedError User cancelled, the call was not user initiated, the origin is insecure, or Permissions Policy blocks capture. Call from a button click, use HTTPS or localhost, inspect policy headers and iframe permissions, and explain the prompt.
NotFoundError No capturable display source is available. Check OS permissions and test on a supported desktop browser.
NotReadableError The selected source cannot be read because of an OS or browser problem. Try another surface, close conflicting capture software, and retry after reporting a clear error.
OverconstrainedError Requested constraints cannot be satisfied by the chosen surface. Relax exact values, use ideal/max, or apply constraints after capture.
TypeError Invalid options, such as video: false or malformed constraints. Pass a valid video request and validate option construction.
No audio track Browser or selected surface does not support audio capture, or the user did not share it. Check getAudioTracks(), describe audio as optional, and provide a video-only path.
Preview is black The stream was not attached, autoplay policy blocked playback, or the track ended. Set video.srcObject = stream, use autoplay playsinline muted, call video.play() after the click if needed, and inspect track.readyState.
Sharing stops unexpectedly The user stopped it in browser chrome, the source disappeared, or the OS revoked access. Listen for ended, clean up state, and offer a restart button.
Works locally but not in production Production is not a secure context or its policy/iframe configuration differs. Serve over HTTPS and compare response headers and embedding attributes.

Or skip the browser setup

If your goal is a server-side image or PDF of a URL rather than a live user-selected display, ScreenshotNeo provides a single HTTP request. Its capture flow removes cookie banners, newsletter popups, and chat widgets before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and responses identify the result with X-Page-Verdict and X-Billed headers. ScreenshotNeo also includes an MCP server so Claude, Cursor, and other MCP clients can call take_screenshot, get_page_info, and capture_pdf.

See the ScreenshotNeo API documentation for all options. The following examples use the API base endpoint and save the returned image.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const file = await res.arrayBuffer();
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(file)));

ScreenshotNeo supports full-page captures with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF output, custom CSS and JavaScript, click and wait actions, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Existing parameter names used by other screenshot APIs also work, which helps when migrating.

There is a free plan with 1,000 screenshots per month and no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account.

Cost and architecture notes

WebRTC screen sharing itself does not impose a ScreenshotNeo-style per-shot charge: your costs come from signaling, TURN relay traffic, bandwidth, storage, recording, and infrastructure. Direct peer-to-peer paths are usually cheaper than relaying, but TURN may be necessary when networks cannot establish a direct connection. Screen content can change rapidly, so encoding and bitrate settings affect both CPU and bandwidth.

For static URL screenshots, a capture API avoids maintaining browser automation, display sessions, consent handling, and rendering workers. Cache stable pages with an appropriate TTL, use asynchronous jobs for slow or large captures, and inspect verdict and billing headers when reconciling usage.

FAQ

Can a website capture a screen without asking?

No. The browser must let the user choose a display surface every time and cannot persist silent display-capture permission.

Does getDisplayMedia() capture only a browser tab?

No. Depending on the browser and platform, the picker may offer a tab, application window, or entire display. Your code must handle the surface the user selects.

Why does screen audio work in one browser but not another?

Audio capture support depends on the browser, operating system, and selected surface. Check the current compatibility information and inspect the returned audio tracks.

Can constraints force a specific monitor?

No. Constraints may affect the resulting track, but they cannot silently narrow the user’s source choices.

What should I use for screenshots of a website URL?

Use a rendering service such as ScreenshotNeo when you need a clean PNG, JPEG, WebP, or PDF from a URL without asking an end user to share their display.