ScreenshotNeo

BlogHow-to

How to Automate Browser Actions in Chrome

Learn when to use Chrome extensions, Playwright, Selenium, or CDP, with runnable examples for clicks, forms, waits, testing, security, and screenshots.

By the ScreenshotNeo team1 October 202610 min read

Use a Chrome extension when automation must run in the user’s open browser. Use Playwright or Selenium for repeatable end-to-end workflows in a separate process. Use the Chrome DevTools Protocol (CDP) when you need low-level network, debugging, or performance control.

This guide shows complete implementations for each route, explains permissions and privacy boundaries, and covers the reliability problems that make browser automation fail.

1. Choose the right Chrome automation route

Route Runs where Best for Main trade-off
Chrome extension Inside a user’s installed Chrome Page enhancements and workflows triggered by the user Permission, messaging, and Web Store rules
Playwright Separate Node.js, Python, Java, or .NET process Reliable end-to-end tests, automation jobs, screenshots Requires browser lifecycle and test infrastructure
Selenium Separate WebDriver client and browser Existing WebDriver grids and broad language support More driver and synchronization setup
CDP Directly against a Chromium debugging endpoint Network inspection, tracing, profiling, and specialized controls Low-level protocol complexity

Chrome’s extension guidance describes content scripts as files that run in web-page context. Playwright and Selenium control a browser from an external process, while CDP is the low-level interface for instrumenting, inspecting, debugging, and profiling Chromium browsers.

2. Define the workflow before writing automation

  1. Write the exact user-visible sequence: open page, enter values, click control, and expected result.
  2. Choose a stable success condition, such as a visible heading, URL change, downloaded file, or enabled button.
  3. Select the least powerful route that can complete it.
  4. Identify authentication, iframes, popups, downloads, dialogs, and permission prompts.
  5. Decide what data may leave the machine and document it.

Reliable automation asserts what a user can see. Avoid assertions that depend on private extension state or implementation details.

3. Automate with a Chrome extension (Manifest V3)

An extension is the right choice when a user clicks your extension and the workflow should operate on the tab they are viewing. Keep page interaction in a content script. Send privileged work to the service worker with message passing.

3.1 Minimal project

chrome-workflow/
  manifest.json
  content.js
  service-worker.js

3.2 manifest.json

{
  "manifest_version": 3,
  "name": "Chrome Workflow Example",
  "version": "1.0.0",
  "description": "Fills a form on the active page after a user action.",
  "permissions": ["activeTab", "scripting"],
  "background": {"service_worker": "service-worker.js"},
  "action": {"default_title": "Run workflow"},
  "content_scripts": [{
    "matches": ["https://example.com/*"],
    "js": ["content.js"],
    "run_at": "document_idle"
  }]
}

activeTab grants temporary access after an explicit user action. Narrow the matches pattern to the sites your feature actually needs. Request the narrowest permissions necessary, as required by Chrome Web Store policy.

3.3 content.js

function waitForSelector(selector, timeoutMs = 10000) {
  return new Promise((resolve, reject) => {
    const existing = document.querySelector(selector);
    if (existing) return resolve(existing);

    const observer = new MutationObserver(() => {
      const element = document.querySelector(selector);
      if (element) {
        observer.disconnect();
        clearTimeout(timer);
        resolve(element);
      }
    });

    const timer = setTimeout(() => {
      observer.disconnect();
      reject(new Error(`Timed out waiting for ${selector}`));
    }, timeoutMs);

    observer.observe(document.documentElement, {childList: true, subtree: true});
  });
}

async function runWorkflow() {
  const email = await waitForSelector('input[type="email"]');
  const submit = await waitForSelector('button[type="submit"]');

  email.focus();
  email.value = 'person@example.com';
  email.dispatchEvent(new Event('input', {bubbles: true}));
  email.dispatchEvent(new Event('change', {bubbles: true}));

  submit.click();
  await waitForSelector('[role="status"], .success, .confirmation');

  chrome.runtime.sendMessage({type: 'workflow-complete'});
}

runWorkflow().catch(error => {
  chrome.runtime.sendMessage({type: 'workflow-error', message: error.message});
});

Dispatching input and change matters for frameworks that listen for events instead of reading the DOM only after submission.

3.4 service-worker.js

chrome.action.onClicked.addListener(async (tab) => {
  if (!tab.id) return;
  try {
    await chrome.scripting.executeScript({
      target: {tabId: tab.id},
      files: ['content.js']
    });
  } catch (error) {
    console.error('Could not inject workflow:', error);
  }
});

chrome.runtime.onMessage.addListener((message, sender) => {
  if (message.type === 'workflow-complete') {
    console.log('Workflow completed in tab', sender.tab?.id);
  }
  if (message.type === 'workflow-error') {
    console.error('Workflow failed:', message.message);
  }
});

3.5 Load and debug the extension

  1. Open chrome://extensions.
  2. Enable Developer mode.
  3. Choose Load unpacked and select the project directory.
  4. Open a matching page and click the extension action.
  5. Use the page DevTools console for content-script errors and the extension’s service-worker inspector for background errors.

3.6 Extension limitations

  • A content script cannot directly use every privileged extension API; message passing bridges the page and service-worker contexts.
  • Pages with strict frame or origin boundaries may require a content script in the correct frame.
  • Cross-origin iframes cannot be treated as ordinary DOM descendants.
  • Do not ship remote executable code in the extension package unless a documented permitted API covers it.
  • Browsing activity and page content are user data. Disclose collection and transmission, even when processing is local, and transmit sensitive data securely.

4. Automate Chrome with Playwright

Playwright is usually the simplest choice for a new repeatable workflow. It provides browser launch, locators, auto-waiting, assertions, screenshots, downloads, network controls, and connections to Chromium-based browsers.

4.1 Install

mkdir chrome-playwright
cd chrome-playwright
npm init -y
npm install -D playwright
npx playwright install chromium

4.2 Complete Node.js example

import { chromium } from 'playwright';

const browser = await chromium.launch({headless: true});
const context = await browser.newContext({
  viewport: {width: 1440, height: 900},
  locale: 'en-US',
  timezoneId: 'UTC'
});
const page = await context.newPage();

try {
  await page.goto('https://example.com/login', {waitUntil: 'domcontentloaded', timeout: 30000});
  await page.getByLabel('Email').fill('person@example.com');
  await page.getByLabel('Password').fill(process.env.TEST_PASSWORD ?? 'replace-me');
  await page.getByRole('button', {name: 'Sign in'}).click();
  await page.getByRole('heading', {name: 'Dashboard'}).waitFor({state: 'visible', timeout: 15000});
  await page.screenshot({path: 'dashboard.png', fullPage: true});
} finally {
  await browser.close();
}

Prefer user-facing locators such as getByRole, getByLabel, and getByText. Use CSS or test IDs when the accessible interface is unavailable. Replace fixed sleeps with waits for a visible state, URL, response, or network condition.

4.3 Useful Playwright patterns

// Wait for a specific application request
await Promise.all([
  page.waitForResponse(response => response.url().endsWith('/api/save') && response.ok()),
  page.getByRole('button', {name: 'Save'}).click()
]);

// Handle a download
const downloadPromise = page.waitForEvent('download');
await page.getByRole('button', {name: 'Export'}).click();
const download = await downloadPromise;
await download.saveAs('export.csv');

// Handle a new tab
const popupPromise = page.waitForEvent('popup');
await page.getByRole('link', {name: 'Open report'}).click();
const popup = await popupPromise;
await popup.waitForLoadState('domcontentloaded');

// Work inside an iframe
const frame = page.frameLocator('iframe[title="Payment"]');
await frame.getByLabel('Card number').fill('4242424242424242');

// Accept a JavaScript dialog
page.on('dialog', async dialog => {
  if (dialog.type() === 'confirm') await dialog.accept();
  else await dialog.dismiss();
});

5. Automate Chrome with Selenium

Selenium is a strong fit when your organization already uses WebDriver, a remote browser grid, or several programming languages. Selenium documents WebDriver and WebDriver BiDi, the W3C bidirectional protocol for browser automation.

5.1 Install

python -m venv .venv
source .venv/bin/activate
pip install selenium

5.2 Complete Python example

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

options = webdriver.ChromeOptions()
options.add_argument('--headless=new')
options.add_argument('--window-size=1440,900')

driver = webdriver.Chrome(options=options)
wait = WebDriverWait(driver, 20)

try:
    driver.get('https://example.com/login')
    email = wait.until(EC.visibility_of_element_located((By.LABEL, 'Email')))
    password = wait.until(EC.visibility_of_element_located((By.LABEL, 'Password')))
    email.send_keys('person@example.com')
    password.send_keys('replace-me')
    wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, 'button[type="submit"]'))).click()
    wait.until(EC.visibility_of_element_located((By.CSS_SELECTOR, 'h1.dashboard')))
    driver.save_screenshot('dashboard.png')
finally:
    driver.quit()

If a label locator is unavailable, use a stable ID, role-oriented CSS selector, or a test-specific attribute. Avoid selectors based on generated class names or DOM position.

6. Use the Chrome DevTools Protocol directly

CDP is appropriate when you need browser-level controls such as request interception, console events, tracing, performance instrumentation, or protocol commands unavailable in a higher-level framework. It is more verbose and couples your code to Chromium protocol details.

6.1 Start Chrome with a debugging port

google-chrome --headless=new --remote-debugging-port=9222 --user-data-dir=/tmp/chrome-automation

6.2 Connect with Playwright’s CDP support

import { chromium } from 'playwright';

const browser = await chromium.connectOverCDP('http://127.0.0.1:9222');
const context = browser.contexts()[0] ?? await browser.newContext();
const page = context.pages()[0] ?? await context.newPage();
await page.goto('https://example.com', {waitUntil: 'domcontentloaded'});
console.log(await page.title());
await browser.close();

Use direct protocol commands only when a framework abstraction cannot express the required operation. Keep CDP sessions isolated and protect the debugging endpoint.

7. Selectors, waits, and state that survive UI changes

  • Prefer accessible roles, labels, and visible names.
  • Wait for the state you need: visible, enabled, attached, URL changed, response received, or download started.
  • Use an explicit timeout that matches the environment; do not make every timeout unlimited.
  • For dynamic lists, locate the item by its text or data attribute, then act within that item.
  • After navigation, wait for the next page’s meaningful element instead of assuming the load event means the app is ready.
  • When an action changes state asynchronously, assert the resulting state immediately.

8. Security and privacy checklist

  • Use the narrowest extension host permissions and prefer activeTab or optional permissions where possible.
  • Never hard-code production passwords, API keys, or session cookies in source control.
  • Use a dedicated test account with the minimum required privileges.
  • Mask secrets in logs, traces, screenshots, and downloaded artifacts.
  • Disclose what page data is read, stored, or transmitted.
  • Keep browser and driver versions controlled and update them deliberately.
  • Do not expose a remote CDP port to an untrusted network.

9. Common errors and fixes

Error Likely cause Fix
Element not found Wrong selector, frame, or page state Inspect the DOM, switch to the correct frame, and wait for a visible condition.
Element is covered Cookie banner, modal, or overlay intercepts the click Close the overlay through a user-visible control, then retry.
Click does nothing Button is disabled or app event has not initialized Wait for enabled state and dispatch the correct input/change events when filling fields.
Timeout during navigation Slow server, blocked resource, or page never reaches the chosen load state Use a realistic timeout, wait for a specific readiness element, and capture console/network logs.
Works headed but fails headless Viewport, timing, GPU, or hidden-dialog difference Set the same viewport, handle dialogs explicitly, and record a trace or screenshot.
Iframe controls unavailable Locator searches the top document Use Playwright’s frame locator or Selenium’s frame switch.
Popup or download missed Listener registered after the click Start waiting for the event before triggering the action.
Extension permission error Host pattern or API permission is too narrow Request only the required permission and verify the active tab matches it.
Stale element reference Framework re-rendered the node Locate the element again immediately before acting.
Authentication loop Expired session, blocked third-party cookie, or wrong origin Use a clean, controlled context, authenticate through the supported flow, and verify the resulting account state.

10. Performance, reliability, and cost

Performance

  • Reuse a browser process and create isolated contexts for independent tasks.
  • Run independent pages in parallel only when the target service and machine have capacity.
  • Block unnecessary images, fonts, analytics, and advertising resources in test environments when they are irrelevant to the assertion.
  • Use a targeted readiness selector instead of waiting for every network request to become idle on applications with long-lived connections.

Reliability

  • Retry only transient failures such as a failed navigation; do not blindly retry a non-idempotent purchase or form submission.
  • Record the URL, browser version, console errors, failed requests, and a screenshot on failure.
  • Keep test data isolated so retries do not collide with earlier runs.
  • Run a small smoke flow before a large batch to detect authentication or deployment failures.

Cost

Self-hosted extensions, Playwright, Selenium, and CDP consume your own browser and compute resources. Hosted browser grids add provider charges and usually charge for runtime or concurrency. Minimize unnecessary browser launches, reuse contexts, and clean up every browser in a finally block.

11. Or skip the browser setup

If your goal is a clean screenshot after a Chrome workflow, ScreenshotNeo provides a single GET request. See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account and get 1,000 screenshots each month with no card.

12. FAQ

Can I automate Chrome without an extension?

Yes. Playwright and Selenium launch or connect to Chrome from a separate process, and CDP connects directly to a Chromium debugging endpoint.

Should I choose Playwright or Selenium?

Choose Playwright for a new project that benefits from modern locators and built-in waiting. Choose Selenium when WebDriver infrastructure, an existing grid, or a particular language is already standard in your organization.

When should I use CDP instead?

Use CDP for network inspection, debugging, profiling, tracing, or protocol features that higher-level libraries do not expose conveniently.

Why does an extension need a service worker?

The service worker handles privileged extension APIs and long-lived extension events. Content scripts interact with the page and communicate with the worker through messages.

How do I make automation safe for real users?

Request minimal permissions, disclose page-data handling, keep secrets out of logs, use secure transport, and make the workflow’s purpose clear to users and reviewers.