ScreenshotNeo

BlogHow-to

How to Capture Webpages as PNG Images in Ruby

Capture webpage screenshots as PNG files in Ruby with Ferrum, Selenium, Cuprite or Watir, then compare a hosted ScreenshotNeo workflow.

By the ScreenshotNeo team1 October 20261 min read

Direct answer: use a Ruby browser automation library to load the page, wait until the content you need is rendered, and save a PNG screenshot. Ferrum gives a compact direct API for Chrome or Chromium:

require "ferrum"

browser = Ferrum::Browser.new
begin
  browser.go_to("https://example.com")
  browser.screenshot(path: "page.png")
ensure
  browser.quit
end

Ferrum drives Chrome or Chromium through the Chrome DevTools Protocol and does not require Selenium, WebDriver or ChromeDriver. A Chrome or Chromium binary is still required; it must be on PATH, available through BROWSER_PATH, or configured with browser_path. See the Ferrum documentation for current installation details.

1. Install Ruby, Ferrum and a browser

Add Ferrum to your project:

bundle add ferrum
# or
 gem install ferrum

Install Chrome or Chromium using your operating system’s package manager or the browser instructions for your deployment image. Verify that the executable is available:

google-chrome --version
# or
chromium --version

If the binary is not on PATH, configure its location:

require "ferrum"

browser = Ferrum::Browser.new(
  browser_path: "/usr/bin/chromium"
)

2. Capture a viewport screenshot

A normal screenshot captures the browser viewport. Set the viewport explicitly when you need repeatable dimensions:

require "ferrum"

browser = Ferrum::Browser.new(
  window_size: [1440, 900]
)
begin
  browser.go_to("https://example.com")
  browser.screenshot(path: "viewport.png", format: "png")
ensure
  browser.quit
end

PNG is Ferrum's default format. Supplying format: "png" makes the intent clear and avoids surprises if capture settings are later shared with JPEG or WebP output.

3. Capture a complete page

Use full: true to capture the full document rather than only the visible viewport:

require "ferrum"

browser = Ferrum::Browser.new(window_size: [1440, 900])
begin
  browser.go_to("https://example.com/article")
  browser.screenshot(path: "article-full.png", full: true, format: "png")
ensure
  browser.quit
end

Full-page capture is different from viewport capture: the resulting image can be much taller, and very long pages may consume substantial memory. If a page uses lazy loading, scroll through it first so images that appear only after scrolling are requested before the screenshot.

4. Capture one element or a rectangular area

Ferrum can capture an element selected with CSS:

require "ferrum"

browser = Ferrum::Browser.new
begin
  browser.go_to("https://example.com")
  browser.screenshot(
    path: "hero.png",
    selector: "main .hero",
    format: "png"
  )
ensure
  browser.quit
end

For a fixed rectangle, pass an area:

browser.screenshot(
  path: "region.png",
  area: { x: 0, y: 0, width: 800, height: 500 },
  format: "png"
)

Do not combine capture modes accidentally. If full: true is combined with selector: or area:, the selector or area is ignored. If both selector and area are supplied, the area is ignored.

5. Control output format, scale and encoding

Option Use
path: Write binary image data directly to a file.
format: Choose png, jpeg/jpg or webp.
full: Capture the full document.
selector: Capture an element selected by CSS.
area: Capture a rectangular region.
encoding: Return Base64 or binary data instead of saving a file.
scale: Change capture scale for higher or lower pixel density.
background_color: Set the page background through Ferrum's RGBA type.

To keep the image in memory as Base64:

png_base64 = browser.screenshot(format: "png", encoding: :base64)
File.write("page.base64", png_base64)

Use PNG for lossless text and interface screenshots. JPEG can be smaller for photographic pages but introduces compression artifacts. WebP is useful when your downstream system supports it.

6. Wait for dynamic pages before capturing

Calling screenshot immediately after navigation can capture a loading state. Choose a wait condition that matches the page instead of relying on one universal delay. For example, wait for a selector that marks the finished view:

require "ferrum"

browser = Ferrum::Browser.new
begin
  browser.go_to("https://example.com/dashboard")
  browser.at_css("[data-rendered='true']", wait: 15)
  browser.screenshot(path: "dashboard.png", format: "png")
ensure
  browser.quit
end

For content that appears after a known short transition, a bounded sleep can supplement a selector wait, but it should not replace a condition that proves the required content exists. For lazy images, scroll in stages and wait for the image elements to finish loading before the final capture.

7. Complete reusable Ruby script

#!/usr/bin/env ruby
require "ferrum"

url = ARGV.fetch(0, "https://example.com")
output = ARGV.fetch(1, "page.png")

browser = Ferrum::Browser.new(
  window_size: [1440, 900],
  timeout: 30
)
begin
  browser.go_to(url)
  browser.at_css("body", wait: 15)
  browser.screenshot(path: output, format: "png")
  puts "Saved #{output}"
ensure
  browser.quit
end

Run it with:

ruby capture.rb https://example.com example.png

8. Choosing another Ruby library

Cuprite with Capybara

Use Cuprite when your application already uses Capybara. Cuprite is a Ferrum-based headless Chrome or Chromium driver, so it fits Capybara's session and test APIs.

Selenium

Use Selenium when your codebase already relies on WebDriver. Selenium's Ruby API supports browser and selected-element screenshots:

require "selenium-webdriver"

driver = Selenium::WebDriver.for :chrome
begin
  driver.navigate.to "https://example.com"
  driver.save_screenshot("page.png")
  driver.find_element(css: "main").take_screenshot("main.png")
ensure
  driver.quit
end

The Selenium full-page options and defaults are Selenium-specific; do not assume they map directly to Ferrum.

Watir

Watir's screenshot API can save a file or return PNG/Base64 data:

require "watir"

browser = Watir::Browser.new(:chrome)
begin
  browser.goto "https://example.com"
  browser.screenshot.save "page.png"
ensure
  browser.close
end

9. Troubleshooting

Symptom Likely cause Fix
Browser executable not found Chrome/Chromium is missing or not on PATH. Install a browser, set BROWSER_PATH, or pass browser_path.
Screenshot is blank or incomplete Capture ran before client-side rendering finished. Wait for a meaningful selector or other page-specific readiness condition.
Lazy images are missing Images load only after scrolling into view. Scroll through the document, wait for image loads, then use full: true.
Element selector fails The selector is wrong, the element is inside an iframe, or it never appears. Check the selector in browser developer tools; wait for it; handle the iframe context explicitly.
Full-page option has no effect A selector or area was supplied with full: true. Choose one capture mode; full-page capture ignores selector and area.
Out-of-memory process A very tall page or high scale creates a large bitmap. Capture sections, lower scale, reduce viewport size, or use a hosted capture service.
Fonts differ from local output The runtime lacks the page's fonts or loads them late. Install required fonts and wait for font-dependent content before capture.
Navigation timeout The site is slow, blocked, or waiting on an external resource. Set a bounded timeout, inspect network dependencies, and retry only transient failures.

10. Performance, reliability and cost considerations

  • Reuse a browser when capturing many URLs. Starting Chrome for every image adds process startup overhead; isolate sessions when cookies or authentication must not leak between jobs.
  • Keep waits specific. Waiting for the selector that represents the finished view is usually more predictable than a large fixed delay.
  • Limit page size. Full-page, high-scale screenshots require more memory than viewport captures.
  • Handle cleanup. Always close the browser in an ensure block so failed jobs do not leave orphaned processes.
  • Retry selectively. Retry transient navigation or network failures with a cap; do not hide deterministic selector or configuration errors behind retries.
  • Budget the runtime. Self-hosted capture costs compute, browser maintenance and engineering time. The reviewed sources do not establish a universal performance ranking among Ferrum, Cuprite, Selenium and Watir.

11. Or skip the browser setup

ScreenshotNeo is a hosted website screenshot API. One GET request returns a PNG, JPEG, WebP or PDF, so your Ruby process does not need to install or manage Chrome.

Ruby can call the endpoint with any HTTP client:

require "net/http"
require "uri"

uri = URI("https://api.screenshotneo.com/v1/shot")
uri.query = URI.encode_www_form(
  access_key: "YOUR_API_KEY",
  url: "https://stripe.com"
)
response = Net::HTTP.get_response(uri)
raise "Screenshot failed: #{response.code}" unless response.is_a?(Net::HTTPSuccess)
File.binwrite("shot.webp", response.body)

See the ScreenshotNeo API documentation for request options. The same endpoint can be called from cURL, Python or Node.js:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie and consent banners, newsletter popups and chat widgets before capture; each cleanup step can be disabled. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. It also provides full-page and element capture, dark mode, device presets, custom CSS and JavaScript, waits, request blocking, headers, cookies, user agents, geolocation, caching, signed links, asynchronous jobs, bulk capture and an MCP server with take_screenshot, get_page_info and capture_pdf.

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is available on every plan. Create a free ScreenshotNeo account.

12. FAQ

Can Ruby save a screenshot without Selenium?

Yes. Ferrum controls Chrome or Chromium directly through CDP and does not depend on Selenium, WebDriver or ChromeDriver. You still need a browser binary.

What is the difference between viewport and full-page PNGs?

A viewport image contains the currently visible browser area. A full-page image extends across the document's entire scroll height and can be substantially larger.

Can I capture only a CSS element?

Yes. Ferrum supports selector:; Selenium and other libraries have their own element screenshot APIs.

Why is my screenshot missing content loaded by JavaScript?

Navigation completion does not guarantee that application rendering is complete. Wait for a selector or page-specific readiness signal before capturing.

When should I use a hosted API?

Use one when installing browsers, managing fonts, handling dynamic pages and operating capture workers would cost more time than the screenshot feature itself. ScreenshotNeo provides that hosted path and a free monthly tier.