ScreenshotNeo

BlogGuides

Using GitHub Copilot and AI to Automate Tests

Use GitHub Copilot to draft tests, check edge cases, and automate parts of your test workflow. Learn a review-first process that keeps tests tied to your requirements.

By the ScreenshotNeo team4 October 20267 min read

GitHub Copilot can help draft unit, integration, mock-based, and end-to-end tests. Give Copilot the code, test framework, expected behavior, and edge cases; review the generated assertions, then run the tests with your project’s test runner. Generated tests are suggestions, not proof that the code is correct or that every important scenario is covered.

1. Choose the right Copilot workflow

For a focused function or file, use Copilot Chat to generate tests in the context of the code and nearby test files. GitHub documents the /tests command for existing code. For test-driven development, ask Copilot to write tests without using /tests, then implement code against those tests.

Need Workflow
Tests for existing code Select the function or file and ask Copilot Chat; try /tests with framework and case requirements.
Test-first development Describe the intended behavior and ask for tests without /tests.
Changes across several files or command execution Use IDE agent mode where available and configured; give it steps and review edits and commands.
Repeatable repository tasks Consider Copilot app automations or GitHub Agentic Workflows, after checking access, policy, and preview availability.

Use the test framework already established in the repository. Copilot’s capabilities and controls vary by IDE, plan, organization policy, and feature availability. GitHub describes Agentic Workflows as public preview, so verify current status before relying on them.

2. Prepare useful repository context

  1. Open the implementation and relevant neighboring tests so the assistant can see naming, fixtures, and framework conventions.
  2. Select the smallest useful unit of code, or identify the file and behavior precisely.
  3. State the expected behavior, not just “write tests.” Include representative inputs and outputs, boundaries, invalid inputs, and expected errors.
  4. Name the test framework and any conventions such as fixture use, mocking rules, or test file location.

For example, a useful request says which empty input is expected, what result should be returned, and whether a dependency should be mocked. If a suggestion misses the point, provide a concrete input and expected output rather than repeating a vague instruction.

3. Generate tests for existing code

In Copilot Chat, select the relevant code and request tests. GitHub documents this pattern:

/tests Write pytest tests for this function. Cover normal inputs, empty input, boundary values, and invalid input. Follow the existing fixtures and do not change production code.

Adapt the framework and wording to your project. Copilot’s documentation includes examples with Python unittest, while its IDE guidance also shows frameworks such as Jest. The generated syntax and command must match your repository.

When the behavior is unclear, first resolve the specification from requirements or product decisions. Do not let the implementation alone define expected behavior: that can produce tests that simply preserve an existing bug.

4. Use Copilot for test-driven development

For TDD, ask for tests describing the desired behavior without the /tests command. Review and save the tests, run them to see the expected failure, implement the behavior, and run them again. If a generated test encodes an assumption you did not make, correct the test before treating its failure as meaningful.

5. Ask for edge cases and meaningful assertions

Make edge cases explicit. Depending on the function, consider:

  • Empty, missing, null, or whitespace-only values, where the language and contract allow them.
  • Minimum and maximum valid values, just-inside and just-outside boundaries, and zero or negative values.
  • Malformed input, unsupported values, and expected exceptions or error results.
  • Duplicate items, ordering, case sensitivity, Unicode, time zones, rounding, and floating-point tolerances when relevant.
  • Dependency failures, retries, timeouts, and cleanup for integration or asynchronous code.
  • Authorization and access boundaries for code that handles identities or permissions.

Ask for observable behavior. Prefer assertions on outputs, public side effects, and documented errors over assertions about private helper calls or incidental implementation details. Mocks are useful for isolating external dependencies, but excessive mocking can make a test pass while real integration behavior is broken.

6. Review and run the generated tests

  1. Read every test and verify each expected value against the requirements.
  2. Check that the test would fail if the behavior regressed. Watch for assertions that only check a value exists, repeat the implementation’s logic, or accept too many outcomes.
  3. Confirm mocks and fixtures model the dependency behavior the test claims to cover.
  4. Run the project’s usual test command from the repository, with the same relevant configuration used in development or CI.
  5. Investigate failures. A failure can point to a test defect, environment setup, or a production defect; do not automatically edit code until the cause is clear.
  6. Review coverage alongside assertion quality. Coverage can show unexecuted lines, but it does not establish that assertions catch regressions.

GitHub’s examples show commands such as python -m unittest; use the command and flags documented by your project. A generated test file does not run itself unless you explicitly invoke the test runner or configure an agent to do so.

7. Extend work with agents and automation

IDE agent mode can handle multi-step work across project files and run commands, subject to the IDE’s support and controls. Give an agent a bounded task, name the files or behavior in scope, ask it to show a summary of edits and command results, and inspect its changes before accepting them.

Copilot app automations can save tasks to run on demand or on a schedule. Cloud automation has repository and organization prerequisites, including cloud-agent access and policy settings. GitHub Agentic Workflows provide a public-preview GitHub Actions approach for natural-language repository tasks, including test-coverage work; availability and behavior can change. For scheduled or cloud work, confirm permissions, review gates, and what credentials or repository access the task needs.

Approach Best fit Review focus
Single Chat request A focused test file or function Assertions, framework fit, missing cases
IDE agent Multi-file edits and command execution Diff, commands run, generated side effects
Saved or scheduled automation Recurring repository maintenance Access prerequisites, policy, approval controls, task output

8. Measure whether automation helps

Try the workflow on a small set of representative changes. Compare the resulting tests with written requirements, review the missing and incorrect cases, and track outcomes that matter to your team, such as review effort, useful regression coverage, and corrections needed. Product examples show how a workflow can be used; they are not independent evidence of correctness, defect reduction, or time saved.

9. Troubleshoot common problems

Symptom Likely cause What to do
Tests use the wrong framework or location The prompt lacked repository conventions or nearby examples. Open a matching test, name the framework and destination, and regenerate or edit the tests.
Tests compile but do not catch a bug Assertions are weak, implementation-shaped, or omit the relevant case. State the regression scenario and expected observable result; add a test that fails for that regression.
Copilot invents behavior The requirement was ambiguous or absent from context. Check the specification, then give explicit examples and expected outcomes.
Mock-based test passes while integration fails The mock does not represent actual dependency behavior. Verify the mock contract and add an integration test at the appropriate boundary.
Agent changes unrelated files The task scope was broad or the agent inferred extra work. Review the diff, revert unrelated edits, and retry with file and behavior boundaries.
Automated task cannot access a repository or run Cloud-agent access, organization policy, or repository prerequisites are missing. Check current product prerequisites and organization settings; use an interactive workflow if access is unavailable.
Tests fail locally or in CI only Environment, configuration, timing, or external-service assumptions differ. Compare runner versions and configuration; remove uncontrolled dependencies or make setup explicit.

10. Reliability, performance, and cost

Generated tests need the same review as other code. Keep a human approval point for changes and commands, especially in multi-file or scheduled workflows. Treat preview features as changeable and avoid unattended workflows whose access or failure consequences have not been reviewed.

Generation speed is not the main measure of value: a fast suggestion still costs review time if it contains incorrect assumptions or weak assertions. Limit prompts to relevant code and context, and start with focused tasks; expand to agent workflows only when cross-file work or repeated execution justifies the additional review and setup. Copilot plan costs and feature entitlements are not specified here, so check GitHub’s current plan and organization details before adopting a workflow.

11. Or skip the browser setup

If one of your tests or monitoring jobs needs a website screenshot, you can capture it with your own browser automation or use ScreenshotNeo, a website screenshot API and MCP server from Yorker Media. One GET request returns an image or PDF. The examples below use PNG output; see the ScreenshotNeo API documentation for supported formats and parameters.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed; response headers report the page verdict and billing status. An MCP server gives AI agents tools to take screenshots, inspect page information, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Every feature is on every plan. Sign up for 1,000 free screenshots a month, no card required.

12. FAQ

Can Copilot generate unit tests?

Yes. It can draft tests for existing code when given code context, a framework, and expected behavior. Review and run them before relying on them.

Can Copilot run tests automatically?

An IDE agent may run commands as part of a multi-step task, and separate automation features can run saved tasks. Availability, access, and controls depend on the configured product and organization.

Does a passing generated test prove the code is correct?

No. It shows that the implementation passed those particular checks. Requirements may still be missing, and assertions may be too weak.

Should Copilot choose my test framework?

Usually use the framework already adopted by the project. Name it in the prompt so the generated tests fit the repository.