ScreenshotNeo

BlogHow-to

How to Run Selenium Headless Chrome Screenshots on an AWS Lightsail Server in India

Set up a Mumbai Lightsail instance to capture web pages with Selenium and headless Chrome, save screenshots, and troubleshoot common server issues.

By the ScreenshotNeo team4 October 202610 min read

To take Selenium screenshots on an AWS Lightsail server in India, create an Ubuntu instance in the Mumbai Region (ap-south-1), install Chrome and Selenium, run Chrome with --headless=new, wait for the page content you need, and save a screenshot. Headless Chrome runs without a visible desktop; you do not need Xvfb for this workflow. [AWS Lightsail Regions] [Selenium Chrome documentation] [Chrome Headless documentation]

This guide uses Python for the complete capture example. It also includes a minimal Selenium JavaScript example, since Selenium bindings exist in both languages. The server steps are Ubuntu-specific; package names and browser installation methods can vary on other Linux distributions.

1. Create an Ubuntu Lightsail instance in Mumbai

  1. In the AWS Lightsail console, create an instance and select the Asia Pacific (Mumbai) Region, ap-south-1. Choose an Ubuntu Linux/Unix image and an instance size suitable for your expected concurrency and page complexity.
  2. Connect through the browser-based SSH terminal or your SSH client. AWS documents instance creation and SSH access in its Lightsail getting-started guide.
  3. Update the operating system packages before installing browser software:
sudo apt-get update
sudo apt-get upgrade -y

Use Mumbai if an India-based deployment is a requirement. Region selection alone does not guarantee low latency to every target website: the target’s own location, routing, response time, and content affect each capture. Lightsail instances run in one Availability Zone within the selected Region.

2. Install Chrome and Selenium

Install Python and the system libraries commonly needed by Chrome. Then install Google Chrome from Google’s official Linux package. The following Ubuntu shell commands download Google’s current stable Debian package rather than pinning a browser version:

sudo apt-get install -y python3 python3-venv python3-pip wget ca-certificates
wget -O /tmp/google-chrome-stable_current_amd64.deb https://dl.google.com/linux/direct/google-chrome-stable_current_amd64.deb
sudo apt-get install -y /tmp/google-chrome-stable_current_amd64.deb

mkdir -p ~/selenium-shot
cd ~/selenium-shot
python3 -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install selenium

These commands assume an Ubuntu x86-64 instance. Confirm the browser is installed and record its version:

google-chrome --version
python --version
python -c 'import selenium; print(selenium.__version__)'

Modern Selenium includes Selenium Manager, which can obtain a compatible browser driver in common supported setups. This requires the instance to have outbound network access when a driver download is needed. Selenium’s driver documentation describes Selenium Manager and driver troubleshooting, including availability from Selenium 4.6. [Selenium Manager] [Driver location troubleshooting]

Driver management choices

Approach Use it when Trade-off
Selenium Manager You use a current Selenium release and can allow driver downloads. Less manual version maintenance; first startup may need network access and take longer.
Explicit ChromeDriver You need controlled, repeatable browser and driver versions or downloads are restricted. You must keep Chrome and ChromeDriver compatible and update them together.

If you manage ChromeDriver yourself, ensure its major version matches Chrome’s, put it on PATH or set its path explicitly, and verify both versions after every browser update. Selenium’s Chrome guidance describes compatibility and Chrome options. [Selenium Chrome documentation]

3. Capture a page with Python

Save this as capture.py in the project directory. Pass the target URL as an argument and choose the output path. The script sets a fixed viewport, waits for the document load event, then waits briefly for client-side rendering before taking a viewport screenshot. For a page with known dynamic content, replace the short delay with a wait for a meaningful element, as shown below.

#!/usr/bin/env python3
import argparse
from pathlib import Path

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait


def main():
    parser = argparse.ArgumentParser(description="Capture a page with headless Chrome")
    parser.add_argument("url", help="Full URL, including https://")
    parser.add_argument("--output", default="screenshot.png", help="Output PNG path")
    parser.add_argument("--width", type=int, default=1365)
    parser.add_argument("--height", type=int, default=900)
    args = parser.parse_args()

    options = Options()
    options.add_argument("--headless=new")
    options.add_argument("--window-size={},{}".format(args.width, args.height))
    # Useful in restricted container or low-shared-memory environments.
    options.add_argument("--disable-dev-shm-usage")

    driver = webdriver.Chrome(options=options)
    try:
        driver.set_page_load_timeout(60)
        driver.get(args.url)
        WebDriverWait(driver, 20).until(
            lambda browser: browser.execute_script("return document.readyState") == "complete"
        )
        # A site may render more content after the load event. Prefer a selector wait
        # when you know which element must appear before the capture.
        output = Path(args.output)
        output.parent.mkdir(parents=True, exist_ok=True)
        driver.save_screenshot(str(output))
        print("Saved {} ({}x{})".format(output, args.width, args.height))
    finally:
        driver.quit()


if __name__ == "__main__":
    main()

Run it from the activated virtual environment:

python capture.py https://example.com --output /tmp/example.png --width 1365 --height 900

save_screenshot captures the current viewport. For a page that renders content only after a particular element appears, use an explicit condition rather than assuming a fixed delay:

from selenium.webdriver.common.by import By

WebDriverWait(driver, 30).until(
    lambda browser: browser.find_element(By.CSS_SELECTOR, "main article")
)
driver.save_screenshot("article.png")

That condition checks that the element exists. If it must also be visible, use Selenium’s expected conditions, such as visibility_of_element_located. Choose a selector that is stable for the target site. A timeout should be treated as a failed or incomplete capture, not silently saved as a valid screenshot.

Full-page capture

Selenium’s standard screenshot method captures the viewport. For a basic full-page image, measure the document height, resize the browser window, and capture again. Very tall pages can consume substantial memory or exceed practical image dimensions; for those, capture sections or use a tool with native full-page capture.

height = driver.execute_script(
    "return Math.max(document.body.scrollHeight, document.documentElement.scrollHeight)"
)
driver.set_window_size(1365, height)
driver.save_screenshot("full-page.png")

Some sites load lazy images only as they approach the viewport. Scroll through the page before measuring and capturing if those images matter, then allow each area to render. A fixed-height full-page resize may not trigger every site’s lazy-loading logic.

4. Minimal Node.js Selenium example

If your application is already in Node.js, install the Selenium WebDriver package with npm install selenium-webdriver and make sure Chrome is installed as above. With compatible ChromeDriver available through Selenium Manager or your configured environment, this script saves a viewport screenshot:

const { Builder, By, until } = require('selenium-webdriver');
const chrome = require('selenium-webdriver/chrome');
const fs = require('node:fs/promises');

(async () => {
  const url = process.argv[2];
  if (!url) throw new Error('Usage: node capture.js https://example.com');

  const options = new chrome.Options()
    .addArguments('--headless=new', '--window-size=1365,900', '--disable-dev-shm-usage');
  const driver = await new Builder().forBrowser('chrome').setChromeOptions(options).build();
  try {
    await driver.manage().setTimeouts({ pageLoad: 60000 });
    await driver.get(url);
    await driver.wait(until.elementLocated(By.css('body')), 20000);
    const png = await driver.takeScreenshot();
    await fs.writeFile('screenshot.png', png, 'base64');
    console.log('Saved screenshot.png');
  } finally {
    await driver.quit();
  }
})().catch((error) => {
  console.error(error);
  process.exitCode = 1;
});

5. Make captures reliable

  • Wait for the page you need. A navigation completing does not mean every asynchronous widget, image, or client-rendered section is ready. Wait for a site-specific selector or condition. Avoid arbitrary long sleeps as the only readiness check.
  • Set timeouts. Bound page-load and element waits so a hung target cannot occupy a worker indefinitely. On timeout, record the URL and error, close the driver, and retry only when the failure may be transient.
  • Always close the browser. Call quit() in a finally block. Leaked Chrome processes gradually consume memory and file descriptors.
  • Use one browser per isolated job or worker. Reusing a session can save startup time, but it also carries cookies, local storage, tabs, and page state between jobs. For unrelated or untrusted URLs, isolate sessions and control concurrency.
  • Expect target-side variation. Sites may redirect, require authentication, block automation, show consent dialogs, or return different content by region. A successful driver call does not prove that the intended page content appeared.
  • Keep outputs and logs bounded. Use unique filenames for concurrent jobs, check that the resulting file exists and has nonzero size, and avoid logging credentials embedded in URLs.

Headless mode needs no display server. Chrome still uses CPU, memory, disk, and network, and complex pages can need more resources than a static page. Start with modest parallelism and observe the instance under your own workload; this dossier contains no performance benchmarks for a particular Lightsail size.

6. Configure Lightsail access and recovery

Allow inbound SSH only from trusted source addresses where practical. Lightsail’s IPv4 and IPv6 firewall rules are configured independently, so review both if the instance has both address types. The screenshot worker needs outbound access to the target pages and, when using Selenium Manager, possibly to download a driver. Do not expose a public WebDriver port for this local workflow. AWS explains Lightsail firewall rules.

If you stop and restart an instance without a static IP, its public IP can change. Attach a static IP if callers need a stable address. Lightsail snapshots provide a point-in-time recovery copy and have storage charges; a restored snapshot also restores its older installed software, so apply security updates after recovery. See AWS guidance on static IPs and instance snapshots.

7. Troubleshooting

Symptom Likely cause Fix
SessionNotCreatedException or Chrome exits at startup Chrome and ChromeDriver are incompatible, browser dependencies are missing, or Chrome cannot start with the current options. Check google-chrome --version and the driver version; align their major versions. Update Selenium and system packages, then inspect ChromeDriver’s full error output.
Driver executable not found or Selenium Manager download fails Old Selenium, no outbound network, DNS/TLS trouble, or a manually installed driver not on PATH. Upgrade the Selenium binding; allow required outbound access or install a compatible ChromeDriver and set its path explicitly.
Chrome reports “DevToolsActivePort file doesn’t exist” Chrome failed to initialize, often due to environment restrictions, resource pressure, or a stale/incompatible browser setup. Check Chrome/driver compatibility and system logs, ensure the process can write to its temporary profile, reduce concurrent browsers, and try --disable-dev-shm-usage where shared memory is constrained.
Screenshot is blank or missing content Capture ran before client rendering, content is below the fold and lazy-loaded, or the site returned a bot check or error page. Wait for a meaningful selector, scroll to trigger lazy loading, check the final URL and page title, and inspect the screenshot before treating the job as successful.
Navigation timeout The target is slow, has long-running requests, or never reaches the browser’s load condition. Set a bounded page-load timeout and use an appropriate site-specific readiness condition. Treat repeated timeouts as target or network failures rather than increasing limits without bound.
Permission denied writing screenshot The process user cannot write to the output directory. Choose a writable path, create its parent directory, and run the capture as the intended service user rather than relying on an interactive shell’s permissions.
Works over SSH but fails as a service The service uses a different Python environment, PATH, user, working directory, or environment variables. Use absolute paths for the virtual environment, script, and output; configure the service user and environment explicitly; capture stderr and exit status.

8. Performance, reliability, and cost

Each Selenium job starts or uses a real browser process, so the meaningful cost is the Lightsail instance and its ongoing operation, not a per-screenshot API charge. Select an instance based on measured memory and CPU needs at your intended concurrency; pages with large scripts, images, or long waits increase resource use. This guide does not claim a throughput figure or recommend a specific plan as universally sufficient.

Improve reliability by limiting concurrent Chrome sessions, restarting workers cleanly after repeated browser failures, using bounded retries for transient network errors, and making jobs idempotent with unique output paths. Keep Chrome, Selenium, the operating system, and any pinned ChromeDriver maintained. A pinned stack helps reproducibility but adds update work; Selenium Manager reduces manual driver setup in common cases but requires it to resolve a compatible driver.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF. Its [API documentation] describes the options and parameters.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners are accepted and removed, along with known newsletter popups and chat widgets, before the shot.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers report the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients.
  • The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. All features are available on every plan.

Sign up for ScreenshotNeo and get 1,000 free screenshots a month with no card.

FAQ

Do I need Xvfb to run headless Chrome?

No. Chrome headless runs without a visible UI or desktop display server. Xvfb is relevant only if you choose to run a non-headless graphical browser in a virtual display.

Does choosing Mumbai mean every screenshot request is served from India?

The browser runs on the Lightsail instance you created in Mumbai. Network routes and target infrastructure still determine how quickly individual sites respond.

Can Selenium capture a PDF?

This guide’s Selenium examples save PNG screenshots. Chrome has separate printing-to-PDF capabilities, while ScreenshotNeo’s API also returns PDFs with configurable PDF options.