How to Use Browser Automation from Any Programming Language
Learn how Selenium, Playwright, Puppeteer and WebDriver let you automate browsers from nearly any language, with setup, code and troubleshooting.
Short answer: choose a browser automation framework that has bindings for your language, or call a language-neutral protocol such as WebDriver. Selenium is the broadest protocol-oriented option; Playwright provides JavaScript/TypeScript, Python, Java and .NET APIs; Puppeteer is a JavaScript library for Chrome and Firefox. Your choice depends on language ecosystem, browser engines, protocol features and how you will run the automation.
This guide shows the setup model, runnable examples, cross-language design patterns, browser coverage, scaling choices, failure handling and a hosted screenshot alternative.
1. Choose the automation model first
| Need | Good starting point | Why |
|---|---|---|
| Many programming languages and browser vendors | Selenium WebDriver | WebDriver is a language-neutral interface and protocol with separate language bindings, browsers and browser-specific drivers. Selenium setup documentation |
| One API across Chromium, WebKit and Firefox | Playwright | Official APIs are available for JavaScript/TypeScript, Python, Java and .NET, with matching browser binaries. Playwright languages |
| JavaScript automation focused on Chrome or Firefox | Puppeteer | Puppeteer uses CDP and WebDriver BiDi; Chrome defaults to CDP and Firefox defaults to BiDi. Puppeteer protocol documentation |
| Only a rendered image or PDF | ScreenshotNeo | It returns a screenshot or PDF from one HTTP request, without maintaining browser setup. |
Keep these questions separate:
- Language: Which language and test runner does your team already use?
- Browser: Do you need Chromium, branded Chrome or Edge, WebKit, Firefox, or several?
- Execution: Will the browser run locally, in CI, on Selenium Grid or through another remote service?
- Protocol: Do you need request/response WebDriver commands or bidirectional browser events?
- Output: Are you driving an interactive workflow, running tests, scraping permitted content, or only producing screenshots and PDFs?
2. Selenium WebDriver from different languages
Selenium describes WebDriver as “an API and protocol that defines a language-neutral interface for controlling the behaviour of web browsers.” The binding, browser and driver are separate setup components, so install and version all three for your target environment.
Python
from selenium import webdriver
from selenium.webdriver.common.by import By
options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
options.add_argument("--window-size=1440,1000")
driver = webdriver.Chrome(options=options)
try:
driver.get("https://example.com")
print(driver.title)
heading = driver.find_element(By.TAG_NAME, "h1")
print(heading.text)
driver.save_screenshot("example.png")
finally:
driver.quit()
JavaScript (Node.js)
import { Builder, By } from "selenium-webdriver";
const driver = await new Builder().forBrowser("chrome").build();
try {
await driver.manage().window().setRect({ width: 1440, height: 1000 });
await driver.get("https://example.com");
console.log(await driver.getTitle());
console.log(await (await driver.findElement(By.css("h1"))).getText());
await driver.takeScreenshot().then(data => require("fs").writeFileSync("example.png", data, "base64"));
} finally {
await driver.quit();
}
Java
import org.openqa.selenium.By;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.chrome.ChromeDriver;
import org.openqa.selenium.chrome.ChromeOptions;
public class Main {
public static void main(String[] args) {
ChromeOptions options = new ChromeOptions();
options.addArguments("--headless=new", "--window-size=1440,1000");
WebDriver driver = new ChromeDriver(options);
try {
driver.get("https://example.com");
System.out.println(driver.getTitle());
System.out.println(driver.findElement(By.tagName("h1")).getText());
((org.openqa.selenium.TakesScreenshot) driver).getScreenshotAs(
org.openqa.selenium.OutputType.FILE).renameTo(new java.io.File("example.png"));
} finally { driver.quit(); }
}
}
C# (.NET)
using OpenQA.Selenium;
using OpenQA.Selenium.Chrome;
var options = new ChromeOptions();
options.AddArgument("--headless=new");
options.AddArgument("--window-size=1440,1000");
using IWebDriver driver = new ChromeDriver(options);
driver.Navigate().GoToUrl("https://example.com");
Console.WriteLine(driver.Title);
Console.WriteLine(driver.FindElement(By.TagName("h1")).Text);
((ITakesScreenshot)driver).GetScreenshot().SaveAsFile("example.png");
Remote execution and Grid
When local execution is insufficient, Selenium Server and Grid provide remote browser control and parallel execution. Keep the test code unchanged and replace the local driver with a remote command executor configured for your Grid endpoint. Verify the browser, driver and server versions together.
3. Playwright in JavaScript, Python, Java and .NET
Playwright lists JavaScript/TypeScript, Python, Java and .NET language APIs. Core browser automation features are available across them, while test-runner integration differs by ecosystem. Playwright supports Chromium, WebKit and Firefox, plus branded Chrome and Edge; install the browser binaries that match your Playwright version. See the browser support and installation documentation.
JavaScript
import { chromium } from "playwright";
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });
try {
await page.goto("https://example.com", { waitUntil: "domcontentloaded" });
await page.locator("h1").waitFor();
console.log(await page.title());
await page.screenshot({ path: "example.png", fullPage: true });
} finally { await browser.close(); }
Python
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page(viewport={"width": 1440, "height": 1000})
try:
page.goto("https://example.com", wait_until="domcontentloaded")
page.locator("h1").wait_for()
print(page.title())
page.screenshot(path="example.png", full_page=True)
finally:
browser.close()
Java
import com.microsoft.playwright.*;
public class Main {
public static void main(String[] args) {
try (Playwright pw = Playwright.create()) {
Browser browser = pw.chromium().launch(new BrowserType.LaunchOptions().setHeadless(true));
Page page = browser.newPage(new Browser.NewPageOptions().setViewportSize(1440, 1000));
page.navigate("https://example.com", new Page.NavigateOptions().setWaitUntil(WaitUntilState.DOMCONTENTLOADED));
page.locator("h1").waitFor();
System.out.println(page.title());
page.screenshot(new Page.ScreenshotOptions().setPath(java.nio.file.Paths.get("example.png")).setFullPage(true));
browser.close();
}
}
}
.NET
using Microsoft.Playwright;
using var playwright = await Playwright.CreateAsync();
await using var browser = await playwright.Chromium.LaunchAsync(new() { Headless = true });
var page = await browser.NewPageAsync(new() { ViewportSize = new() { Width = 1440, Height = 1000 } });
await page.GotoAsync("https://example.com", new() { WaitUntil = WaitUntilState.DOMContentLoaded });
await page.Locator("h1").WaitForAsync();
Console.WriteLine(await page.TitleAsync());
await page.ScreenshotAsync(new() { Path = "example.png", FullPage = true });
Playwright setup checklist
- Install the language package.
- Install the Playwright browser binaries for that package version.
- Pin package and browser versions in CI.
- Choose the engine explicitly when cross-browser coverage matters.
- Use locator waits and assertions instead of arbitrary sleeps.
4. Puppeteer and protocol choices
Puppeteer is a JavaScript library for high-level automation of Chrome and Firefox. Current documentation says Chrome uses CDP by default because some CDP features are not yet available through BiDi, while Firefox uses BiDi by default. Check protocol support before depending on a browser-specific feature.
import puppeteer from "puppeteer";
const browser = await puppeteer.launch({ headless: true });
const page = await browser.newPage();
try {
await page.setViewport({ width: 1440, height: 1000 });
await page.goto("https://example.com", { waitUntil: "domcontentloaded" });
await page.waitForSelector("h1");
console.log(await page.title());
await page.screenshot({ path: "example.png", fullPage: true });
} finally { await browser.close(); }
5. A language-neutral architecture
Design your automation behind a small interface so the rest of your application does not depend on one library:
interface BrowserSession {
open(url: string): Promise<void>;
waitFor(selector: string): Promise<void>;
text(selector: string): Promise<string>;
screenshot(path: string): Promise<void>;
close(): Promise<void>;
}
Implement that interface with Selenium, Playwright or Puppeteer in the language your service uses. Keep selectors, timeouts, navigation policy and output naming in shared configuration. This makes a later browser or framework change a contained adapter change.
6. Waiting, state and browser context
- Wait for a meaningful selector, URL condition or network state; do not assume a fixed delay means the page is ready.
- Create isolated contexts or profiles for independent users and tests.
- Set viewport, locale, timezone, user agent and permissions explicitly when they affect rendering.
- Load authentication state through the framework’s supported storage mechanism; never hard-code credentials in source.
- For infinite-scroll pages, define a stopping condition such as item count, height growth or a maximum scroll count.
- For downloads, popups and new tabs, register the expected event before triggering the action.
7. Reliability and security checklist
- Use a bounded navigation timeout and a separate assertion timeout.
- Capture the URL, browser version, console errors and a screenshot or trace when a run fails.
- Retry only transient failures such as a connection reset; do not retry deterministic selector or authorization errors blindly.
- Close pages and browsers in a finally/try-with-resources block.
- Restrict outbound network access when automating untrusted pages, and avoid exposing internal services through user-controlled URLs.
- Pin framework and browser versions, then update them deliberately because browser behavior changes.
8. Troubleshooting common errors
| Symptom | Likely cause | Fix |
|---|---|---|
| Driver executable not found | Selenium binding, browser or driver is missing from PATH. | Install the browser-specific driver, expose it to the process, and verify compatible versions. |
| Playwright launches but cannot find a browser | Matching browser binaries were not installed. | Run the package’s browser installation command in the same build image used at runtime. |
| Element not found | The page has not rendered it, the selector changed, or it is inside a frame or shadow root. | Wait for a stable locator, inspect frames and shadow DOM, and avoid brittle generated class names. |
| Click intercepted | A modal, consent banner or overlay covers the target. | Dismiss the overlay, wait for it to disappear, or click the intended element using a locator that verifies visibility. |
| Navigation timeout | Slow origin, blocked resource, redirect loop or an overly short timeout. | Inspect redirects and network logs, set a bounded longer timeout, and choose the correct wait condition. |
| Different screenshot in CI | Fonts, viewport, device scale factor, browser version or timezone differs. | Pin the image, install required fonts, set viewport and timezone, and compare browser versions. |
| WebSocket or BiDi feature missing | The selected browser/protocol does not implement that feature yet. | Check the framework’s current support matrix and use a supported protocol or browser. |
| Remote session drops | Grid node capacity, network interruption or browser crash. | Collect server logs, limit parallel sessions, add bounded retries for transient disconnects and clean up sessions. |
9. Performance, parallelism and cost
- Reuse a browser process and create new contexts or pages when isolation allows; launching a browser for every URL adds startup work.
- Run independent cases in parallel only within the CPU, memory and browser capacity of the worker or Grid.
- Block unnecessary assets only when the test does not depend on them; blocking fonts, scripts or images can change behavior.
- Prefer locator-based waits over long global sleeps, which increase runtime and still fail on slower pages.
- Cache installed browser binaries in CI, but invalidate the cache when the framework version changes.
- Track browser minutes, worker resources, Grid nodes and artifact storage as separate cost drivers. The research sources do not provide comparative speed or price benchmarks.
10. Or skip the browser setup
If your result is a clean screenshot or PDF rather than an interactive browser session, ScreenshotNeo accepts one GET request. It removes cookie and consent banners, newsletter popups and chat widgets before capture; bot checks, blank pages, failed loads, timeouts and cache hits are not billed. Responses include X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients.
See the ScreenshotNeo API documentation for all options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Options include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets or custom viewports, retina scale, PDF paper and page settings, custom CSS and JavaScript, clicks, waits, hidden selectors, request blocking, headers, cookies, user agent, Authorization, timezone, geolocation, transparent backgrounds, resizing, selectable cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture for 100 URLs per call, a usage API and an OpenAPI specification. The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
11. FAQ
Can I automate a browser from a language without an official framework?
Use a language binding for Selenium WebDriver, call a remote WebDriver endpoint, or expose a small service written in a supported language.
Is WebDriver the same as Selenium?
WebDriver is the interface and protocol; Selenium provides implementations, bindings, drivers and tools around it.
Should I use Playwright or Puppeteer for Firefox?
Both document Firefox support. Compare the exact API and protocol features you need, then pin and verify the browser version in your environment.
When is a screenshot API enough?
Use one when you need rendered images or PDFs and do not need to interact with a session across multiple steps. Use browser automation when the workflow itself is the product or test.


