ScreenshotNeo

BlogHow-to

How to Test Your Website for Black Friday and Cyber Monday Traffic

Build a realistic peak traffic test, exercise checkout, find bottlenecks, and turn results into capacity and incident plans before Black Friday and Cyber Monday.

By the ScreenshotNeo team4 October 202610 min read

Direct answer: Use your own traffic history and campaign forecast to define a plausible peak workload, then test realistic shopping journeys—especially checkout—against performance and error limits your team sets. Increase load in stages, test sudden surges and sustained demand when relevant, monitor the application and the load generators, fix bottlenecks, and repeat the same scenarios before the event. There is no universal Black Friday traffic multiplier, concurrency target, or latency threshold that fits every retailer.

Google Cloud recommends load-testing critical workloads three to five months before an event, reviewing quota and cost implications, and using results to inform forecasts and quota requests. Its peak-event guidance also includes monitoring, post-event analysis, and disaster-recovery simulations. Google Cloud peak-event guidance and Google Cloud load-testing guidance are useful planning references.

1. Define what “ready” means for your site

Start with a workload hypothesis and measurable pass criteria. Use prior Black Friday and Cyber Monday traffic, other campaign peaks, planned promotions, expected growth, and marketing forecasts. Record the assumptions so the team can revise them when the forecast changes.

  • Workload: requests per second, concurrent sessions, completed shopping transactions, or a combination. For checkout, completed transactions are more meaningful than raw requests.
  • Scope: pages and dependencies included, such as catalog, product detail, inventory, cart, checkout, payment, and third-party services.
  • Cache behavior: specify whether the run represents a warm cache, a cold cache, or a deliberate mix. Otherwise, you may test a path unlike the one you intend to understand.
  • Pass criteria: set your own maximum acceptable latency and error rate for browsing and purchase completion. Also state the minimum successful transaction rate or volume the business needs.
  • Guardrails: agree on test windows, stop conditions, test accounts and data, and how the team will avoid sending unintended orders or real payment requests.

Do not pick a traffic multiplier because it is commonly repeated. The sources here do not establish a standard holiday multiplier or universal response-time target. Your forecast, architecture, user expectations, and service objectives should determine the test values.

2. Model complete shopping journeys

A useful performance scenario represents what a shopper does, including request dependencies and state. A typical path might be:

  1. Open a landing or category page.
  2. Browse or search, then open a product detail page.
  3. Select a variant or quantity and add the item to the cart.
  4. Review the cart and submit checkout details.
  5. Complete a safe purchase flow in a test environment, or stop before the irreversible payment step if your test setup requires it.

Carry cookies, session tokens, cart identifiers, and other state between dependent requests. Use representative products, query parameters, and request bodies. Include more than one journey if users can arrive through different campaign landing pages or use different checkout paths.

Keep a low-load functional check separate from capacity tests. It answers whether the key pages and purchase actions work at all. AWS describes a shopping website functional test as a full transaction such as browsing a catalog, selecting products, checking out, and purchasing. See AWS Prescriptive Guidance on load testing.

3. Choose test shapes based on the question

Test shape Question it answers What to observe
Smoke or baseline Does the journey work at modest load? Functional failures, baseline latency, and successful transactions.
Expected-load test Can the forecast workload meet the agreed service objectives? Latency and errors by journey step, plus completed transactions.
Stepped capacity or stress test At what load does a pass criterion fail, and does scaling respond? Each load step, time for autoscaling to act, saturation signals, and the failure point.
Spike test Can the system handle a sudden increase, such as a promotion or campaign send? Immediate queueing, errors, recovery time, and scale-up behavior.
Soak or sustained-load test Does prolonged demand expose degradation that a short run misses? Latency drift, resource growth, back pressure, and failures over time.

Write down the hypothesis, load curve, duration, and pass/fail conditions for each run. A stepped test should hold each stage long enough for the configured scaling signals to respond. A spike and a soak test answer different questions; neither replaces an expected-load test. Grafana k6 documents performance test types, including spike, stress, and soak approaches.

4. Prepare the generator and test environment

Before a large run, verify the load generator can produce the intended traffic and that your account, region, and target environment have enough quota and capacity. For distributed tests, document the regions, generator or task count, users per task, ramp-up, and hold duration.

  • Calibrate gradually: increase users or request rate in controlled steps while watching generator CPU and memory.
  • Keep generator placement and network path in mind; they affect what traffic reaches the target and what latency includes.
  • Check quotas and capacity for the generator, target services, databases, caches, and relevant third-party dependencies.
  • Run one calibration test at a time when generators share a cluster, so resource readings can be attributed to the right run.
  • Use test accounts, isolated data where possible, and safeguards for email, fulfillment, payment, and other side effects.

A saturated generator can make a healthy website look ready because it is unable to apply the planned load. AWS advises gradually raising users per test container while monitoring generator CPU and memory, and notes that task capacity depends on quotas and available resources. See AWS Distributed Load Testing solution documentation. Its example values apply to its documented setup; they are not universal generator limits.

5. Run a test and monitor the whole path

Capture user-facing and system-side evidence during each run. At a minimum, track:

  • Latency distributions and errors for each important page or transaction step.
  • Successful end-to-end purchases per time interval, alongside attempted transactions.
  • Achieved request rate or concurrency, compared with the planned load curve.
  • Application and infrastructure saturation signals relevant to your architecture.
  • Scaling activity, queues, retries, and downstream dependency health.
  • Load-generator CPU and memory, to check that the test itself remains capable.

Trace or correlate requests across services so a slow checkout can be connected to the component causing delay. Review the relevant load balancer, application, database, cache, API, payment, and external-service signals for your system. A concurrency-based generator may send fewer requests as the target slows, making the achieved request rate fall at the moment you need to understand degradation. For that investigation, a constant-request-rate approach can help maintain the intended offered load. AWS discusses this distinction and transaction-level measurement in its load-testing guidance.

6. Turn findings into changes and repeat

After every run, save the scenario and configuration, planned and achieved load, transaction success, latency and error results, scaling behavior, and evidence for the suspected bottleneck. Change one meaningful thing where practical, then repeat the same scenario to compare results.

  1. Identify the first pass criterion that failed and the component evidence around that time.
  2. Form a specific hypothesis, such as a database limit, a slow dependency, ineffective caching, or delayed scaling.
  3. Make an operational or code change and record it with the run.
  4. Repeat the same workload and compare both user outcomes and system signals.
  5. Revisit capacity, quota, and test costs before increasing the scale or duration.

Google Cloud recommends using load-test results to evaluate cost implications and inform quota requests. Cloud quotas, regional generator capacity, and product behavior can change, so verify the limits for your account and the date of the planned test.

7. Rehearse event operations and recovery

Capacity is only one part of event readiness. Assign who watches dashboards, who can make operational changes, how incidents are escalated, and how the team will communicate during the event. Rehearse the response to a degraded checkout, unavailable dependency, or unexpectedly high traffic level. Confirm how to roll back risky changes and restore service.

Google Cloud’s peak-event framework recommends active monitoring, post-event analysis, and disaster-recovery simulations one to three months before an event. Schedule these alongside load tests, leaving time to act on what they reveal.

Do it yourself with k6

The following runnable example illustrates a staged, constant-arrival-rate API test for a single endpoint. It is a starting point, not a complete checkout test: add the actual dependent requests, session handling, safe test data, and transaction success checks for your site. Set the URL and thresholds to match your own system. The values below are illustrative, not recommended holiday targets.

import http from 'k6/http';
import { check } from 'k6';
import { Rate } from 'k6/metrics';

const failed = new Rate('failed_requests');
const baseUrl = __ENV.BASE_URL;

if (!baseUrl) {
  throw new Error('Set BASE_URL to the test environment URL');
}

export const options = {
  scenarios: {
    staged_browse: {
      executor: 'ramping-arrival-rate',
      startRate: 2,
      timeUnit: '1s',
      preAllocatedVUs: 20,
      maxVUs: 200,
      stages: [
        { target: 10, duration: '2m' },
        { target: 25, duration: '5m' },
        { target: 25, duration: '5m' },
        { target: 0, duration: '2m' },
      ],
    },
  },
  thresholds: {
    http_req_failed: ['rate<0.01'],
    http_req_duration: ['p(95)<800'],
    failed_requests: ['rate<0.01'],
  },
};

export default function () {
  const response = http.get(`${baseUrl}/products`, {
    tags: { journey: 'browse' },
  });
  const ok = check(response, {
    'browse returned success': (r) => r.status >= 200 && r.status < 400,
  });
  failed.add(!ok);
}

Save it as peak-test.js, install k6 according to the official installation instructions, then run:

BASE_URL=https://staging.example.com k6 run peak-test.js

Replace the example staging host and path with an environment you control. A 2xx-to-3xx check is only a basic availability check; for a real purchase flow, validate the expected content and business outcome. Configure scenario stages, rate, virtual-user capacity, and thresholds from your forecast and service objectives. For browser-level user experience, k6 also documents browser metrics and combining browser tests with protocol-level performance tests; see its browser testing documentation.

Or skip the browser setup

For screenshot checks of campaign pages, product pages, or key states during your preparation, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. It does not generate load or replace a performance test; it can help capture visual evidence of pages at the tested state.

For load testing, keep your own runner and choose its traffic profile. To capture a page, the API call is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for setup and options. Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents use its screenshot, page-info, and PDF tools. The free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000.

Sign up for 1,000 free screenshots a month, with no card required.

Troubleshooting common load-test problems

Symptom Likely cause What to do
Planned request rate is not reached Generator CPU, memory, virtual-user capacity, network, or account quota is limiting the run. Check generator metrics and quota; calibrate gradually; add generator capacity or distribute the run where appropriate.
Target slows but achieved request rate falls A concurrency-based workload sends less traffic when responses get slower. Inspect achieved load and consider a constant-arrival-rate test to study behavior under a maintained offered rate.
Browsing passes but checkout fails The test covers isolated endpoints, omits state, or does not model dependent checkout requests. Carry session and cart state through the full journey, and count completed transactions.
Results vary widely between runs Cache state, data, background work, deployment changes, or generator contention differs. Record the environment and cache assumptions; control test data and background jobs; run calibration separately.
Autoscaling does not appear to help The ramp is too fast or the stages are too short for configured signals and scaling delays. Use gradual stages with sufficient hold time; inspect scaling metrics and service quotas.
Test creates real orders or side effects The flow reaches production integrations or uses live accounts and payment behavior. Use a controlled test environment and test credentials; stub or safely disable fulfillment, email, payment, and other irreversible actions.
High latency appears only from one region Network path, regional capacity, or a region-specific dependency may differ. Compare generator location, target region, and dependency signals; run a controlled regional comparison.

Performance, reliability, and cost considerations

  • Performance: Judge the experience by latency and completion at each meaningful step, not only aggregate request volume. Keep the offered load and achieved load distinct in reports.
  • Reliability: Repeat runs after changes, test the failure and recovery paths that matter, and rehearse alerting and escalation. A single successful run does not demonstrate recovery readiness.
  • Test cost: Generator runtime, distributed regions, target infrastructure, data transfer, and downstream services can all contribute to cost. Estimate before increasing duration or scale.
  • Operational cost: Run tests early enough that teams can interpret results, request quota changes, and implement fixes. Google Cloud’s recommended three-to-five-month window allows room for iteration.
  • Scope: A protocol-level test can apply substantial backend load efficiently, while browser tests provide user-experience signals at higher per-user resource cost. Choose the mix based on the question and validate that generators remain healthy.

Frequently asked questions

How far ahead should I start?

Google Cloud recommends testing critical workloads three to five months before the event. Start early enough to repeat runs, address findings, and verify quotas; schedule recovery simulations in the one-to-three-month window its peak-event guidance describes.

Is there a standard Black Friday traffic multiplier?

No universal multiplier is established by the sources used here. Derive the workload from your historical traffic, forecast, campaign plans, and architecture.

Does passing a load test guarantee the site will stay available?

No. The test represents selected scenarios and conditions. Production traffic, dependencies, failures, and operational response can differ, so pair testing with monitoring, recovery exercises, and an event plan.

Should I test production?

Use an environment and integrations designed for safe load generation. If production testing is necessary for a particular question, define explicit guardrails and coordinate the test so it cannot create unintended purchases or disrupt customers.

Can screenshots tell me whether the site can handle load?

No. Screenshots document rendered appearance. Use a load generator and transaction metrics to measure capacity; screenshots can support visual checks of important pages and states.

Sources and historical context

Google Cloud reported that Apigee customer API calls rose 95% over the comparable 2017 Black Friday period, with peak traffic increasing from 48,000 to 108,000 transactions per second and 99.999% platform availability during the 2018 period. These are historical, vendor-reported Apigee platform figures, not a current benchmark or a prediction for another retailer. See Google Cloud’s 2018 account.