ScreenshotNeo

BlogScreenshots on your device

How to Capture Screenshots with the Screen Capture API

Learn how to turn getDisplayMedia() into a downloadable PNG, handle permissions and errors, capture elements, and automate screenshots with ScreenshotNeo.

By the ScreenshotNeo team29 September 20268 min read

How to Capture Screenshots with the Screen Capture API

Direct answer: navigator.mediaDevices.getDisplayMedia() returns a live MediaStream, not an image file. To save one screenshot, request a stream from a user gesture, read its video track, call new ImageCapture(track).grabFrame(), draw the resulting ImageBitmap on a canvas, and export the canvas with toBlob(). Always handle cancellation and errors, stop every track when finished, and close the bitmap.

This guide covers the complete browser workflow, element and region capture, permissions, security policy, browser limitations, performance, troubleshooting, and an automated alternative. The authoritative references are MDN’s getDisplayMedia() documentation and its guide to Element Capture and Region Capture.

1. The screenshot pipeline

The Screen Capture API has four distinct stages:

  1. A user clicks a button, providing transient user activation.
  2. getDisplayMedia() opens a browser-controlled chooser for a display, window, or tab.
  3. The selected source arrives as a live video track inside a MediaStream.
  4. ImageCapture.grabFrame() produces one ImageBitmap, which a canvas encodes as PNG, JPEG, or another supported format.

The chooser is intentionally controlled by the browser and the user. Options can suggest preferences after selection, but they cannot silently force a particular screen, window, or tab. A new permission prompt is expected for each request; permission is not a reusable silent grant.

2. Complete runnable browser example

Save this as an HTML file and serve it from a secure context such as HTTPS or localhost. The request must be made directly from the button handler.

The Screen Capture API turns a user-selected stream into one canvas image.
The Screen Capture API turns a user-selected stream into one canvas image.
<!doctype html>
<meta charset="utf-8">
<title>Screen Capture Screenshot</title>
<button id="capture">Capture screenshot</button>
<a id="download" hidden>Download PNG</a>
<img id="preview" alt="Captured screenshot preview">
<pre id="status"></pre>
<script>
const button = document.querySelector('#capture');
const download = document.querySelector('#download');
const preview = document.querySelector('#preview');
const status = document.querySelector('#status');
let previousUrl = null;

button.addEventListener('click', async () => {
  let stream;
  let bitmap;
  try {
    status.textContent = 'Choose a screen, window, or tab…';
    stream = await navigator.mediaDevices.getDisplayMedia({
      video: true,
      audio: false,
      preferCurrentTab: true
    });

    const [track] = stream.getVideoTracks();
    if (!track) throw new Error('The selected source has no video track.');

    bitmap = await new ImageCapture(track).grabFrame();
    const canvas = document.createElement('canvas');
    canvas.width = bitmap.width;
    canvas.height = bitmap.height;
    const context = canvas.getContext('2d');
    context.drawImage(bitmap, 0, 0);

    const blob = await new Promise((resolve, reject) => {
      canvas.toBlob(result => result ? resolve(result) : reject(new Error('PNG encoding failed')), 'image/png');
    });

    if (previousUrl) URL.revokeObjectURL(previousUrl);
    previousUrl = URL.createObjectURL(blob);
    preview.src = previousUrl;
    download.href = previousUrl;
    download.download = `screenshot-${Date.now()}.png`;
    download.hidden = false;
    status.textContent = `${bitmap.width} × ${bitmap.height} PNG ready.`;
  } catch (error) {
    status.textContent = `${error.name || 'Error'}: ${error.message}`;
  } finally {
    if (stream) stream.getTracks().forEach(track => track.stop());
    if (bitmap) bitmap.close();
  }
});
</script>

The finally block matters. A stopped track releases the capture session, while bitmap.close() releases the image’s graphics resources. Revoke an old object URL before replacing it so repeated captures do not retain unnecessary blobs.

3. Request options and what they actually do

Option Purpose Important limitation
video Required capture constraint; usually true or a constraint object. false is invalid because screenshots require a video track.
audio Optionally requests system or tab audio. Audio is unnecessary for a still and may add another permission choice.
preferCurrentTab Hints that the current tab should be prominent in the chooser. It does not select the tab for the user.
selfBrowserSurface Controls whether the current tab may appear as a source. It changes chooser behavior; it does not bypass consent.
surfaceSwitching Hints whether the user can switch shared surfaces during a session. Support varies by browser.
monitorTypeSurfaces Hints whether entire monitors should be offered. The browser can ignore unsupported or unsafe combinations.
displaySurface Constraint describing a preferred monitor, window, or browser tab. It cannot force a source before the user chooses.

Use conservative options first. Feature-detect optional constraints and test the target browsers because chooser hints and newer capture extensions change independently of the basic API.

4. Capturing one element or a region

The normal flow captures the entire selected surface. If the requirement is a DOM element, two newer approaches have different privacy behavior:

Element Capture isolates the DOM tree, while Region Capture crops a rectangle.
Element Capture isolates the DOM tree, while Region Capture crops a rectangle.
  • Element Capture restricts the stream to an element and its descendants. Other overlapping page content is excluded.
  • Region Capture crops the tab to the element’s bounding rectangle. Content that overlaps that rectangle can remain visible.

Element Capture is the better fit for isolating a component; Region Capture is a geometric crop. Both are optional capabilities and are currently documented as desktop-only. Confirm support for your exact browser versions before making them a production dependency.

async function captureElement(element) {
  const stream = await navigator.mediaDevices.getDisplayMedia({
    video: true,
    audio: false,
    preferCurrentTab: true
  });
  const [track] = stream.getVideoTracks();
  let bitmap;
  try {
    if (!('RestrictionTarget' in window) || !track.restrictTo) {
      throw new Error('Element Capture is not supported in this browser.');
    }
    const target = await RestrictionTarget.fromElement(element);
    await track.restrictTo(target);
    bitmap = await new ImageCapture(track).grabFrame();
    return bitmap;
  } finally {
    stream.getTracks().forEach(t => t.stop());
    // The caller owns the returned bitmap and must call bitmap.close().
  }
}

The target element must be rendered and available when RestrictionTarget.fromElement() runs. Layout changes after restriction can affect the resulting frame. If you only need a stable rectangle, take a full frame and crop it on a canvas instead, while remembering that overlapping content remains part of that rectangle.

5. Permissions, policy, and security

Call getDisplayMedia() from a click, pointer, or keyboard handler with transient activation. Calling it later from a timer commonly causes InvalidStateError. The browser must show the chooser so the user can see exactly what may be exposed. Treat captured pixels as sensitive: they can include private documents, notifications, credentials, or other applications.

Screen capture requires a secure context. If a site embeds the capture code in an iframe, the parent may need:

<iframe src="/capture.html" allow="display-capture"></iframe>

The documented Permissions Policy default allowlist is self. An allow attribute does not replace the chooser or grant silent access. Review the Screen Capture API security guidance when deploying across origins.

6. Error handling and troubleshooting

Error or symptom Cause Fix
InvalidStateError The call was not made during transient user activation, or the document is not in a valid focused state. Invoke it directly inside the button handler; do not defer with a timer.
NotAllowedError The user cancelled or denied sharing, the browser blocked it, or policy disallows it. Explain that a source must be selected, check iframe policy, and let the user retry.
NotFoundError No capturable display source is available. Check operating-system permissions and retry after a source is available.
NotReadableError The source was selected but an OS, hardware, or browser failure prevented reading it. Stop old tracks, retry, and check system screen-recording permissions.
Black or blank image The track ended, the source changed, the page was hidden, or the frame was grabbed before a usable video frame arrived. Check track.readyState, listen for ended, and retry after a short readiness check.
toBlob() returns null The requested encoder is unavailable or encoding failed. Use image/png, check the canvas context, and handle the rejection.
Element restriction fails Element Capture is unsupported or the element is not rendered. Feature-detect RestrictionTarget; fall back to full capture plus canvas cropping.
Unexpected source in the image The user selected a monitor or window rather than the intended tab. Describe the chooser step clearly; options cannot force a source.

For a long-lived capture, listen for the track ending:

track.addEventListener('ended', () => {
  // The user stopped sharing or the source disappeared.
  captureButton.disabled = false;
});

7. Image format, size, and performance

PNG is lossless and preserves text well, but it can produce large files. JPEG is smaller for photographic content but introduces compression artifacts; WebP may offer a useful compromise where your downstream tools accept it. Pass the desired MIME type and quality to canvas.toBlob(callback, type, quality), then verify the returned blob type.

Capture resolution follows the selected source and browser constraints. A full monitor at high pixel density can create a large canvas and consume substantial memory during drawImage() and encoding. To reduce cost in the browser, scale into a second canvas before encoding:

const maxWidth = 1600;
const ratio = Math.min(1, maxWidth / bitmap.width);
const output = document.createElement('canvas');
output.width = Math.round(bitmap.width * ratio);
output.height = Math.round(bitmap.height * ratio);
output.getContext('2d').drawImage(bitmap, 0, 0, output.width, output.height);

Do not capture continuously when one frame is enough. Stop tracks immediately, close bitmaps, revoke object URLs, and avoid keeping full-resolution canvases in application state. For a stream recorder, use the stream directly; grabFrame() is intended for still images.

8. Testing checklist

  • Serve from HTTPS or localhost.
  • Trigger the request from a visible user gesture.
  • Test monitor, window, and tab selections separately.
  • Test cancel, deny, source-ended, and OS-permission cases.
  • Verify PNG dimensions and file size on standard and high-DPI displays.
  • Test iframe deployments with an explicit allow="display-capture".
  • Feature-detect Element Capture and keep a fallback.
  • Confirm every track is stopped and every bitmap is closed.

9. Or skip the browser setup

If your application needs screenshots of public URLs rather than a user’s own screen, ScreenshotNeo provides a single GET request that returns PNG, JPEG, WebP, or PDF. Its API accepts the URL directly, so there is no chooser, stream, canvas, or browser permission flow. See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are never billed, and response headers report the page verdict and billing result. You can also use its MCP server tools—take_screenshot, get_page_info, and capture_pdf—from Claude, Cursor, or another MCP client. The free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

10. FAQ

Can getDisplayMedia() save a PNG by itself?

No. It returns a live stream. Use ImageCapture.grabFrame(), a canvas, and toBlob() to create the file.

Can I choose a specific monitor without asking the user?

No. The browser controls source selection and requires user consent for every request.

Why is my screenshot different from the browser viewport?

You may have selected a window or monitor, or the source may have a different pixel density. The captured dimensions describe the selected surface.

Should I use Element Capture or Region Capture?

Use Element Capture when overlapping content must be excluded. Use Region Capture when a rectangular crop is sufficient.

Can I use this API on a server?

getDisplayMedia() is a browser API tied to a user’s display and permissions. For server-side URL screenshots, use a rendering service such as ScreenshotNeo.