ScreenshotNeo

BlogHow-to

How to Write Automation Scripts for Browser Tasks

Build reliable browser scripts by choosing a framework, using stable locators, waiting for real conditions, and verifying each task’s result.

By the ScreenshotNeo team4 October 202610 min read

A reliable browser automation script makes the task explicit: open a known page, find the intended control, perform an action, and verify the resulting state. Choose a framework based on the browsers, language, test runner, debugging tools, and execution environment your project needs. Prefer locators tied to accessible names or labels, and use condition-based waits and assertions instead of fixed pauses.

This guide uses Playwright with JavaScript for a complete example, then outlines the equivalent framework choices and reliability practices for Selenium and Puppeteer. The sample automates a controlled test page; use a site and account you are authorized to automate.

1. Define the task and choose a framework

Write down both the action and the evidence that proves it succeeded. For example: “Submit the contact form, then confirm the page displays a success message.” A click by itself does not prove that a task completed.

Choose a framework by matching the project’s needs rather than assuming one tool is best for every job:

Framework Consider it when Useful documented capabilities
Playwright You want a test-oriented workflow in JavaScript or another supported language, browser engine projects, and integrated debugging. Semantic locators, actionability checks, retrying web assertions, code generation, and trace viewing. Locators · Assertions · Best practices
Selenium You need WebDriver-based browser control, an established language binding, or distributed runs through a grid. Selenium Manager can manage browser drivers by default when a driver is not otherwise supplied; Selenium Grid routes commands to remote browser instances. Selenium Manager · Grid
Puppeteer You want to control a browser with its JavaScript API by launching or connecting to a browser and creating pages. Its locator APIs provide readiness checks before actions. Page interactions · Getting started

Compare the browser engines and operating systems you must support, language and API fit, whether this is a one-off task or a repeatable test suite, built-in waits and assertions, debugging facilities, and whether execution must be distributed. The documentation establishes these capabilities; it does not establish a universal speed or reliability winner.

2. Set up and run a complete Playwright example

The following Node.js test opens the Playwright website, clicks its “Get started” link, and verifies that the Installation heading appears. It demonstrates the complete task pattern without depending on a third-party form or account.

Install

mkdir browser-task
cd browser-task
npm init -y
npm install --save-dev @playwright/test
npx playwright install chromium

Save this as browser-task.spec.js:

const { test, expect } = require('@playwright/test');

test('opens the getting started guide', async ({ page }) => {
  await page.goto('https://playwright.dev/');
  await page.getByRole('link', { name: 'Get started' }).click();
  await expect(
    page.getByRole('heading', { name: 'Installation' })
  ).toBeVisible();
});

Run it with:

npx playwright test browser-task.spec.js

Playwright Test supplies the page fixture, runs the test, and reports whether the assertion passed. The asynchronous visibility assertion retries until its condition is met or the configured timeout expires. See the official writing tests guide and assertion reference.

Translate the pattern to your own task

  1. Navigate: open a stable starting URL and make the expected starting state clear.
  2. Find: locate the control using its role and accessible name, label, or an explicit test identifier.
  3. Act: use the framework action that matches the task: click, fill, check, select, or another supported interaction.
  4. Verify: assert a visible confirmation, changed state, expected title, URL, or other outcome that demonstrates success.

For a team-owned form, the task may look like this. Replace the URL and accessible labels with ones from your controlled test application:

test('submits a contact form', async ({ page }) => {
  await page.goto('http://localhost:3000/contact');
  await page.getByLabel('Email').fill('dev@example.test');
  await page.getByLabel('Message').fill('Please contact me.');
  await page.getByRole('button', { name: 'Send message' }).click();
  await expect(page.getByRole('status')).toContainText('Message sent');
});

The example.test address is illustrative. Use test data that your application accepts and that can be safely repeated.

3. Find elements that survive page changes

A locator is a description of the element to find. Prefer selectors that express the user-facing interface or a deliberate test contract:

Target Preferred locator idea Example
Button or link Role and accessible name page.getByRole('button', { name: 'Save' })
Form field Associated label page.getByLabel('Email address')
Known test hook Explicit test identifier page.getByTestId('save-profile')
Repeated control Narrow by meaningful container Find the named button inside the relevant dialog or list item
Element with no semantic hook Use a short CSS selector tied to a stable attribute page.locator('[data-state="ready"]')

Playwright recommends user-facing locators such as roles and labels, or explicit test IDs. Locators are resolved against the current page when used, which helps when a page rerenders between steps. A long chain of layout-dependent selectors or a generated class name is usually an incidental detail and can break when markup changes. Playwright locator guidance

If several controls share the same name, scope the locator to the meaningful region first, such as a dialog with a unique accessible name. Check uniqueness when ambiguity would make the action unsafe. A locator that matches the wrong one of several “Delete” buttons is not made reliable by an automatic wait.

4. Wait for conditions and assert outcomes

Use the framework’s action and assertion APIs to synchronize with page state. Playwright actions check actionability, and its web-first assertions retry while checking the expected state. Puppeteer locators likewise perform checks before interaction. These features reduce common timing races, but they cannot fix an incorrect locator, an unexpected page state, or a failing external service.

Avoid treating a fixed delay as proof that a page is ready:

// Fragile: assumes the page is ready after exactly two seconds
await page.waitForTimeout(2000);
await page.getByRole('button', { name: 'Continue' }).click();

Prefer a condition tied to the task:

await expect(page.getByRole('button', { name: 'Continue' })).toBeEnabled();
await page.getByRole('button', { name: 'Continue' }).click();
await expect(page.getByRole('heading', { name: 'Review order' })).toBeVisible();

For custom page behavior, wait for a specific selector, text, URL, or application state that reflects readiness. Choose timeouts based on the task and environment; extending a timeout can accommodate a genuinely slower operation, but it will also make a real failure take longer to report.

5. Keep runs reproducible and debug failures

  • Control the starting state. Use an explicit URL, known account, and repeatable fixture data. Reset state between tests where practical.
  • Keep tests focused. Make each script verify behavior your team controls. Third-party websites can change markup, rate-limit traffic, or change behavior independently.
  • Handle authentication deliberately. Use an approved test account and the framework’s supported state or secret-management approach. Do not hard-code production passwords or commit credentials.
  • Inspect evidence. Review the failure message, current URL, locator match, page state, and relevant network or browser errors. Playwright’s trace viewer can show action timelines, DOM snapshots, and network activity. Trace viewer
  • Review generated code. Code generation can help discover candidate locators, but check that they are unique, meaningful, and followed by an assertion for the real goal.

For Playwright Test, you can run with a trace enabled using npx playwright test --trace on, then inspect the generated trace with the documented viewer. See running and debugging tests.

6. Adapt the approach to Selenium or Puppeteer

The same task structure applies across frameworks. The examples below show the shape of the interaction; follow the selected framework’s current installation guidance and use its APIs and assertions consistently.

Selenium with Python

Install the Python binding with python -m pip install selenium. Selenium Manager is included with Selenium bindings and can locate or manage a driver when one is not supplied. The explicit wait below waits for a condition instead of sleeping an arbitrary duration.

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

options = webdriver.ChromeOptions()
driver = webdriver.Chrome(options=options)
try:
    driver.get("https://playwright.dev/")
    link = WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.LINK_TEXT, "Get started"))
    )
    link.click()
    heading = WebDriverWait(driver, 10).until(
        EC.visibility_of_element_located(
            (By.XPATH, "//h1[normalize-space()='Installation']")
        )
    )
    assert heading.is_displayed()
finally:
    driver.quit()

Here the Selenium locator uses the link’s visible text and the expected heading text. Where the application exposes accessible attributes or stable IDs, use a locator strategy that reflects those contracts. Consult the official WebDriver documentation and waits guidance.

Puppeteer with Node.js

Install Puppeteer with npm install puppeteer. Its getting-started workflow launches a browser, creates a page, and uses the browser API. This example uses a locator action and checks the resulting heading:

const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch({ headless: true });
  try {
    const page = await browser.newPage();
    await page.goto('https://playwright.dev/', {
      waitUntil: 'domcontentloaded'
    });
    await page.locator('a').filter(button => button.innerText === 'Get started').click();
    await page.waitForFunction(() =>
      [...document.querySelectorAll('h1, h2')]
        .some(el => el.textContent.trim() === 'Installation')
    );
    console.log('Installation heading is visible');
  } finally {
    await browser.close();
  }
})();

Locator APIs evolve, so confirm the exact locator operations for the installed Puppeteer version in its page interaction guide. Keep the assertion tied to the task rather than treating a successful click as completion.

7. Troubleshooting common failures

Symptom Likely cause What to do
Locator times out or finds nothing The page is different than expected, content is inside a frame, the locator is wrong, or the content has not appeared. Inspect the current URL and page state; verify the visible name or label; wait for a task-specific condition; use a frame locator if the control is inside an iframe.
Strict or ambiguous locator error Several elements match the same role, name, or text. Scope to a dialog, row, or other meaningful container. Use a more specific accessible name or stable test ID. Avoid choosing the first match unless order is part of the requirement.
Element is not actionable The target is hidden, disabled, covered, moving, or not the control users interact with. Check the rendered page and expected state. Wait for the correct control to become visible and enabled; do not use forced clicks to conceal an incorrect state.
Assertion fails after a click The click did not trigger the assumed transition, the application returned an error, or the assertion checks the wrong outcome. Inspect the resulting page, URL, and network activity. Assert the application’s actual success indicator and diagnose failed requests separately.
Works locally but fails in CI Different browser version, viewport, permissions, environment variables, data, or resource constraints. Make browser installation and test data explicit, use consistent viewport and configuration, and capture traces or logs from the failed run.
Selenium cannot start the browser Browser or driver is missing, incompatible, or inaccessible; a proxy or firewall may block Selenium Manager downloads. Check installed browser and Selenium versions, network and proxy access, and Selenium Manager diagnostics. Configure a driver explicitly if your environment requires it. Selenium Manager limitations
Browser closes before evidence is inspected Cleanup runs immediately after an error. Save logs, screenshots, or framework traces before teardown. Use headed or step-debugging modes during local diagnosis.
Task changes real data unexpectedly The script ran against a production account or non-repeatable state. Use a controlled staging environment and dedicated test data; make destructive actions explicit and authorized.

8. Performance, reliability, and cost

Performance: Keep scripts focused and avoid redundant page reloads, repeated broad DOM queries, and arbitrary delays. A fixed sleep makes fast runs slower while still failing when the page takes longer than expected. Browser startup, page weight, network conditions, and the number of independent runs all affect elapsed time; the cited framework documentation does not support a universal speed ranking.

Reliability: Semantic locators, condition-based waits, meaningful assertions, isolated state, and captured failure evidence address common sources of flaky automation. They do not guarantee success against changing third-party pages, unstable networks, or services outside your control. Keep third-party dependencies out of assertions when the behavior under test belongs to your own application.

Execution cost: Local browser automation has no per-screenshot API charge, but it uses machine time and requires browser installation, maintenance, and execution infrastructure. Distributed execution adds infrastructure and operational choices; Selenium documents Grid for routing WebDriver scripts to remote browser instances. Budget for debugging and upkeep as well as successful runs.

9. Or skip the browser setup

If the browser task is to capture a page rather than interact with its controls, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
    f.write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', image));

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

10. FAQ

Should browser scripts be tests or standalone automation?

Use a test runner when you need repeatable checks, assertions, and reports. A standalone script can suit a one-off controlled task, but it still needs a clear success condition, safe credentials, and predictable cleanup.

Can browser automation handle any website?

Technically possible interactions depend on the page and browser, but the site’s policies, authentication, and permissions matter. Automate only pages and accounts you are authorized to use, and expect external sites to change independently.

When should I use a screenshot API instead?

Use a screenshot API when the output you need is a rendered image or PDF of a URL. Use browser automation when the task must interact with controls, submit information, or verify an application workflow.

Do automatic waits remove the need for assertions?

No. Waiting helps an action happen at an appropriate time; an assertion confirms that the intended result actually occurred.