ScreenshotNeo

BlogHow-to

How to Generate Playwright Tests and Scale Browser Automation

Record a first Playwright test with codegen, review it for reliable coverage, then scale execution with projects, workers, shards, and CI retries.

By the ScreenshotNeo team4 October 20269 min read

Use Playwright’s code generator to record a representative browser journey, then review and strengthen the generated test before running it across projects, workers, or CI shards. Record with npx playwright codegen https://your-site.example. For reliable scale, make tests independent, keep saved sign-in state private, start CI with one worker when stability is the priority, and use shards to distribute work across jobs.

Codegen helps you get a test started; it does not decide whether the test checks the right outcome. Review its locators, assertions, setup, and expected behavior before treating the result as coverage. Playwright’s codegen guide describes the recorder and its locator choices.

1. Set up a Playwright project

If you do not already have a Playwright Test project, initialize one with the official package setup:

npm init playwright@latest

Follow the prompts to choose JavaScript or TypeScript, whether to add a GitHub Actions workflow, and whether to install browsers. The setup creates a Playwright configuration and sample tests. If you already have a project, use its existing package scripts and configuration instead of initializing it again.

Make sure the browser you intend to use is installed. The setup wizard can install browsers; for an existing project, install the required browser binaries with:

npx playwright install

2. Record a representative journey

Run codegen with the page where the journey begins:

npx playwright codegen https://your-site.example
  1. Playwright opens a browser and the Playwright Inspector.
  2. Perform one realistic user journey, such as searching for an item and opening its details.
  3. Stop when the behavior you want to cover is complete.
  4. Copy the generated code into a test file in your project.
  5. Replace example URLs and data, then review every locator and assertion.

The URL is optional. Without it, codegen opens a browser so you can navigate to the site yourself. Use the recorder to discover a practical starting point, not as proof that the test covers the requirement.

Record signed-in journeys safely

If you already have a Playwright storage-state file, load it for the recording session:

npx playwright codegen --load-storage=auth.json https://your-site.example

Storage state can include cookies, local storage, and IndexedDB data. It can contain credentials or session tokens. Keep it out of source control, use it locally, and delete it when it is no longer needed. See Playwright’s authenticated-state guidance and authentication documentation.

3. Turn recorded actions into a useful test

Codegen prioritizes role, text, and test-id locators. When a locator matches more than one element, it attempts to make it unique. Still, check that each locator identifies the intended control and that the assertions verify the behavior users depend on. A test that only clicks through a page can pass without checking that the feature worked.

For example, a reviewed test might look like this in TypeScript:

import { test, expect } from '@playwright/test';

test('search results show the requested product', async ({ page }) => {
  await page.goto('https://your-site.example');

  await page.getByRole('searchbox', { name: 'Search products' }).fill('desk lamp');
  await page.getByRole('button', { name: 'Search' }).click();

  await expect(page.getByRole('heading', { name: 'Search results' })).toBeVisible();
  await expect(page.getByRole('link', { name: /desk lamp/i }).first()).toBeVisible();
});

Use the actual accessible names, roles, and expected outcomes from your application. The sample is a shape to adapt, not a claim about any particular site.

Review checklist

  • Locator meaning: Does the role, label, text, or test ID identify the intended element?
  • Uniqueness: Does the locator match only the target? If not, narrow it by a meaningful parent or accessible name.
  • Behavior: Does the test assert the result, not only the action?
  • Independence: Can it run without relying on another test’s state or order?
  • Data: Are test records and accounts isolated or reset so parallel runs do not collide?
  • Sensitive state: Are credentials and storage-state files excluded from version control?

Playwright’s best practices explain locator guidance and test design. Prefer locators that reflect how a user or assistive technology finds an element. Use test IDs when the application needs a stable testing contract.

4. Organize browser and environment coverage with projects

A Playwright project is a logical group of tests that shares configuration. Projects can represent browser or device variants, environments, or distinct test groups. The configuration below defines Chromium and Firefox projects and a setup project that runs before them:

import { defineConfig, devices } from '@playwright/test';

export default defineConfig({
  testDir: './tests',
  projects: [
    {
      name: 'setup',
      testMatch: /.*\.setup\.ts/,
    },
    {
      name: 'chromium',
      use: { ...devices['Desktop Chrome'] },
      dependencies: ['setup'],
    },
    {
      name: 'firefox',
      use: { ...devices['Desktop Firefox'] },
      dependencies: ['setup'],
    },
  ],
});

Projects with a dependency wait for the dependency project to finish before they run. Browser projects still use the configured worker limit, so adding projects increases coverage and can also increase total work. Choose a matrix that matches the browsers, devices, and environments you need to support. See Playwright projects for project configuration and dependencies.

Run a single project by name:

npx playwright test --project=chromium

5. Scale execution with workers and shards

By default, Playwright runs test files in parallel; tests within one file run in order. Parallel work happens in separate worker processes, and workers do not share in-memory state. Treat every test as independently runnable: avoid depending on another test’s data, browser session, or side effects.

Strategy Where work runs Useful when Check first
Workers Multiple worker processes in one job The machine has resources for concurrent browsers and tests are independent CPU, memory, service capacity, and collisions in shared test data
Shards Separate CI jobs or machines You want to distribute a suite across jobs That each shard is configured, run, and reported as part of the same CI workflow
Projects Logical browser, device, environment, or test groups You need a defined coverage matrix How project count and worker limits affect total concurrent work

Set a worker limit explicitly when you need to control concurrency. For example:

npx playwright test --workers=4

The value is an example, not a universal recommendation. Actual concurrency depends on the machine, available browser memory, CI limits, and the load your application and test services can handle. Increase it gradually and watch for resource pressure and data conflicts.

To run one shard of a suite, use a fraction such as:

npx playwright test --shard=2/3

This selects shard 2 of 3. Configure your CI workflow to run the other shard fractions in separate jobs and collect their results appropriately. Sharding distributes work across machines or jobs; it does not make tests independent for you. Consult the official sharding guide and CI guide for workflow-specific setup.

Choose a scaling path

  1. Start with a stable suite. Fix order dependencies and shared-data collisions before adding concurrency.
  2. Choose coverage projects. Add only the browser, device, environment, and test-group variants that serve a requirement.
  3. Establish a CI baseline. Playwright recommends setting workers to 1 in CI when prioritizing stability and reproducibility.
  4. Distribute with shards when useful. If wider parallelization is needed, split the suite across CI jobs and monitor the result.
  5. Increase concurrency with evidence. Check available CPU and memory, application capacity, failure patterns, and total CI time at each step.

Worker concurrency and shard count solve different placement problems: workers add processes within a job; shards divide the suite across jobs. Projects add configured coverage dimensions. The best combination depends on resources, test independence, coverage requirements, and the effort needed to diagnose failures.

6. Configure retries and diagnostics deliberately

Retries are disabled by default. A test that fails on its first attempt and passes on retry is reported as flaky; a test that continues to fail through its retries remains failed. Retries can expose intermittent behavior in reports, but a retry pass does not remove the underlying cause.

For example, a CI-oriented configuration can set retries and collect a trace when a retry occurs:

import { defineConfig } from '@playwright/test';

export default defineConfig({
  retries: process.env.CI ? 2 : 0,
  use: {
    trace: 'on-first-retry',
  },
});

This is an example policy, not a mandatory setting. Choose retry count and trace collection to fit your CI cost and investigation needs. When a test is flaky, examine its trace and logs, then investigate timing assumptions, unstable data, environmental dependencies, or resource contention. See the retry documentation and configuration reference.

7. Troubleshoot common failures

Symptom Likely cause What to do
Codegen cannot launch a browser The browser binary is not installed, or the environment cannot launch it. Run npx playwright install for the required browsers and check the environment’s browser dependencies.
A generated locator is ambiguous or selects the wrong element The page has repeated text or controls, or the chosen locator does not identify the intended element. Inspect the page and generated locator. Narrow it with a meaningful role, accessible name, parent, or test ID; keep the locator tied to the intended behavior.
A test passes alone but fails in the suite It may rely on another test’s order, shared state, or shared data. Make it independently runnable, isolate records and accounts, and remove cross-test dependencies.
Tests become flaky after increasing workers Concurrent workers may compete for machine resources or modify shared application data. Reduce worker count, isolate test data, and check resource limits and service capacity before increasing concurrency again.
A shard does not run the expected tests The selected fraction, CI job matrix, or command may not match the intended shard setup. Check the shard fraction in every job and follow the Playwright sharding example for your CI provider.
Authentication is missing during recording The recording session did not load the intended storage state, or that state has expired. Check the file path and sign-in state, regenerate the local state if needed, and keep it out of source control.
A test passes only after retry The first attempt exposed intermittent behavior; the retry has not fixed it. Mark it as flaky in your investigation, inspect trace and logs, and address the underlying race, data, timing, or environment issue.
CI is less stable than local runs CI has different resources or environmental limits, or runs more concurrent work. Use one worker as a stability-oriented starting point, compare CI and local settings, and add shard jobs when you need broader distribution.

8. Performance, reliability, and cost considerations

  • Performance: More workers can run files concurrently, while shards can distribute a suite across jobs. Neither guarantees a faster end-to-end run: machine resources, browser memory, application capacity, and job startup all matter.
  • Reliability: Independent tests, isolated data, and a deliberate project matrix help keep results interpretable. More concurrency can reveal hidden shared state and resource contention.
  • Retries: Retries can make intermittent failures visible in reports, but they also mean additional attempts. Track retry-passing tests as investigation items rather than treating them as clean passes.
  • CI cost: Projects and shards can increase the amount of browser work or the number of concurrent jobs. Choose coverage and parallelism based on the value of faster feedback and the limits of your CI environment.
  • Diagnostics: Trace collection and reports help explain failures but add artifacts and work. Configure them to give enough evidence for debugging without collecting more than the team can use.

Or skip the browser setup

If your immediate task is to capture a website for review, documentation, or an AI workflow rather than test browser interactions, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF. Its API also accepts the parameter names used by other screenshot APIs, which can make switching easier. See the ScreenshotNeo API docs.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners are accepted and removed, along with known consent platforms, newsletter popups, and chat widgets, before the shot; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

FAQ

Does codegen create a complete test automatically?

No. It records interactions and proposes locators. Review the assertions, setup, and expected outcomes to make the test meaningful.

Can Playwright tests share state between workers?

Workers run in separate processes and do not share in-memory state. Set up tests so they do not depend on another worker’s state.

Should I use more workers or more shards?

Workers add concurrency within one job; shards distribute the suite across jobs or machines. Choose based on available resources, test independence, and CI setup.

Do retries fix flaky tests?

No. A retry can help identify intermittency, but a test that passes only after retry still warrants investigation.