How to Perform Hypothesis Testing in Python
Choose a statistical test that fits your question and data, run it with Python, and interpret the p-value, effect estimate, and uncertainty.
To perform hypothesis testing in Python, define a null and alternative hypothesis, identify your outcome and study design, choose a test that matches them, check its assumptions, and interpret the test statistic and p-value alongside an effect estimate and uncertainty interval. For two independent groups with a numeric outcome, a Welch two-sample t-test is a common starting point; in SciPy, use scipy.stats.ttest_ind(..., equal_var=False).
This guide focuses on selecting and running the right analysis, not just calling a function. The examples use SciPy and Statsmodels, whose official documentation describes the available tests and their options: SciPy hypothesis testing tutorial, SciPy statistics reference, and Statsmodels statistics reference.
1. Define the question before choosing a test
Write down the quantity or relationship you want to learn about. Then specify:
- Null hypothesis (H₀): the reference claim, such as equal population means or no association.
- Alternative hypothesis (H₁): the effect or relationship you are looking for.
- Direction: two-sided (a difference in either direction) or one-sided (a difference in a prespecified direction).
Choose the direction and your significance threshold before examining the test result. Switching to a one-sided test or changing the threshold after seeing the data can make the reported evidence misleading.
2. Match the test to outcome and design
Identify the outcome type, how observations were collected, and the target quantity. The following is a starting guide, not a substitute for checking assumptions and the details of the design.
| Question and data | Possible method | Key consideration |
|---|---|---|
| Compare a numeric mean with a fixed value | One-sample t-test | Check whether the sampling model and observations support inference about the mean. |
| Compare numeric outcomes in two independent groups | Welch t-test | Use independent units; Welch’s test does not assume equal population variances. |
| Compare numeric outcomes measured on the same units twice, or in matched pairs | Paired t-test | Analyze within-pair differences; do not treat the paired values as independent groups. |
| Compare numeric outcomes in more than two groups | ANOVA or an appropriate alternative | A significant overall result does not identify which groups differ; plan follow-up comparisons and account for multiplicity. |
| Test association between categorical variables in a count table | Chi-square test of independence | Use counts and check whether the approximation is suitable for the table and sample. |
| Test a small 2 × 2 count table where an exact method is appropriate | Fisher exact test | This tests a different setup from a comparison of numeric means. |
| Test a single proportion or compare proportions | Proportion test or interval procedure | Statsmodels includes proportions_ztest and proportion_confint; check the conditions for the chosen method. |
SciPy documents chi-square and Fisher exact tests as well as tests for numeric data. These procedures are not interchangeable: select one according to the question, outcome, design, and model assumptions, rather than trying tests until one produces a small p-value.
3. Prepare and inspect the data
Each row or value should represent the unit described by your design. For example, if one person contributes repeated measurements, those measurements are not independent people. Check units, group membership, duplicates, missing values, and implausible values before testing.
For a two-group mean comparison, summarize group sizes, means, and standard deviations. Inspect the data for extreme values and distribution features that could make the chosen model inappropriate. A test call cannot repair selection bias, dependent observations, measurement errors, or an unsuitable study design.
4. Run a Welch t-test in Python
Install the libraries in an environment where you can run Python:
python -m pip install scipy numpy
Save this as test_means.py and run python test_means.py. The example data are illustrative; replace them with independent observations from your study.
import numpy as np
from scipy import stats
# Illustrative numeric observations from two independent groups.
group_a = np.array([12.1, 11.4, 13.0, 10.9, 12.6, 11.8])
group_b = np.array([10.5, 9.8, 11.2, 10.1, 9.6, 10.7])
# Remove missing observations explicitly for this example. In real work,
# first investigate why values are missing and whether omission is justified.
a = group_a[np.isfinite(group_a)]
b = group_b[np.isfinite(group_b)]
if len(a) < 2 or len(b) < 2:
raise ValueError("Each group needs at least two valid observations for this example")
# Welch's independent-samples t-test; the two-sided alternative is specified.
result = stats.ttest_ind(
a,
b,
equal_var=False,
alternative="two-sided",
nan_policy="raise",
)
print(f"Group A: n={len(a)}, mean={a.mean():.3f}, SD={a.std(ddof=1):.3f}")
print(f"Group B: n={len(b)}, mean={b.mean():.3f}, SD={b.std(ddof=1):.3f}")
print(f"Mean difference (A - B): {a.mean() - b.mean():.3f}")
print(f"t={result.statistic:.3f}, df={result.df:.2f}, p={result.pvalue:.4g}")
print("95% CI for mean difference:", result.confidence_interval(confidence_level=0.95))
equal_var=False requests Welch’s t-test. The SciPy function defaults to equal_var=True, which requests the conventional pooled-variance test, so set the option deliberately. alternative accepts "two-sided", "less", or "greater". The returned result includes the statistic, p-value, and degrees of freedom, and its confidence_interval() method returns an interval for the difference in population means. See the official ttest_ind reference for version-specific behavior and full parameter details.
Important options and data choices
equal_var:Falseselects Welch’s test;Trueselects the equal-variance pooled test.alternative: choose two-sided or directional inference from the hypothesis specified before analysis.nan_policy: SciPy supports policies such as"propagate","omit", and"raise". Omitting missing values changes the analyzed sample. Understand the missingness and report the resulting group sizes.- Paired observations: use a paired procedure such as
scipy.stats.ttest_relwhen values are matched or repeated on the same units; an independent-samples test answers the wrong question for paired data. - Different outcome: numeric mean tests do not test categorical counts or proportions. Select an appropriate method for those data.
5. Interpret the statistic, p-value, and interval
A p-value is calculated assuming the null model and describes how surprising data at least as extreme as those observed would be under that model. It is not the probability that the null hypothesis is true. SciPy describes the t-test p-value as quantifying the probability of observing results “as or more extreme” under the stated null model (SciPy documentation).
If the p-value is below the prespecified significance threshold, report evidence against the null under the selected model; the test does not prove the alternative. If it is above the threshold, report that the analysis did not provide sufficient evidence to reject the null. That result does not establish equality or prove that an effect is absent.
Report the estimated effect and its uncertainty, too. In the example, the mean difference is group A’s mean minus group B’s mean, and SciPy’s confidence interval gives a range of values compatible with the data and model at the selected confidence level. Statistical significance alone does not show whether an effect is large enough to matter in practice.
A useful reporting template
We compared [outcome] between [group A] (n=...) and [group B] (n=...) using a two-sided Welch independent-samples t-test. The estimated mean difference (A - B) was [estimate] [units], 95% CI [lower, upper]; t(df) = [statistic], p = [value].
Adapt the template to your design. Name the test, state group sizes and the effect estimate, and include an interval and test statistic where available. Describe exclusions and missing-data handling so readers can tell which observations informed the result.
6. Other common Python tests
Categorical association: chi-square and Fisher exact
For categorical observations, construct a contingency table of counts. SciPy provides chi2_contingency for a chi-square test of independence and fisher_exact for Fisher’s exact test in suitable tables. Follow the official SciPy tutorial examples and check the procedure’s assumptions and return values for your installed version.
from scipy import stats
# Rows and columns are categories; values are observed counts.
observed = [[12, 7], [5, 16]]
chi2 = stats.chi2_contingency(observed)
print("chi-square statistic:", chi2.statistic)
print("p-value:", chi2.pvalue)
print("degrees of freedom:", chi2.dof)
# For a 2 x 2 table, an exact alternative may be appropriate.
fisher = stats.fisher_exact(observed, alternative="two-sided")
print("Fisher odds ratio:", fisher.statistic)
print("Fisher p-value:", fisher.pvalue)
Proportions with Statsmodels
For proportion inference, Statsmodels documents proportions_ztest and proportion_confint. Install Statsmodels with python -m pip install statsmodels. The example compares the observed success proportions in two groups; verify that the method’s conditions fit your data.
from statsmodels.stats.proportion import proportions_ztest, proportion_confint
successes = [42, 35]
observations = [100, 90]
z_stat, p_value = proportions_ztest(successes, observations, alternative="two-sided")
print(f"z={z_stat:.3f}, p={p_value:.4g}")
for count, n in zip(successes, observations):
low, high = proportion_confint(count, n, alpha=0.05)
print(f"{count}/{n}: proportion={count/n:.3f}, 95% CI=({low:.3f}, {high:.3f})")
The proportion test and the contingency-table tests answer questions about categorical data, while a t-test concerns means of numeric outcomes. Choose the target quantity first. See the Statsmodels reference for the documented functions and related procedures.
7. Assumptions, edge cases, and reliability
- Independence: the sampling design determines whether observations are independent. Repeated measures, clusters, households, or matched pairs may require a different model.
- Small samples and unusual distributions: inspect the data and assess whether the test’s model is reasonable. Do not assume a test is valid solely because Python returns a p-value.
- Unequal spread: for two independent means, Welch’s test avoids the equal-variance assumption made by the pooled version; it does not fix dependence, confounding, or other design problems.
- Missing values: investigate why data are absent. Omitting missing rows is not automatically unbiased; report how many observations were analyzed.
- Outliers: verify whether unusual values are errors or valid observations. Do not remove them only because they change the result; explain any exclusion rule.
- Multiple tests: running many tests increases the chance of finding a small p-value by chance. Plan comparisons and use an appropriate multiplicity strategy where needed.
- Numerical and version details: record the Python and library versions for reproducibility, especially when relying on result fields or options documented for a particular SciPy release.
- Reproducibility: retain the analysis code, data-cleaning decisions, test direction, threshold, and summaries. A p-value without the design and analysis choices is difficult to evaluate.
8. Performance and cost
For ordinary in-memory samples, these SciPy and Statsmodels test calls are lightweight compared with data collection, cleaning, and model selection. If a dataset is large, the main work is often loading, validating, and transforming the data; avoid repeatedly rebuilding the same summaries inside loops. Benchmark your own workload if runtime matters, because no single runtime applies to every dataset and environment.
The core libraries are available as Python packages; the basic analysis does not require a paid statistical API. Your practical costs are typically compute, storage, and the time needed to validate the study design and assumptions. A fast calculation is not evidence that the analysis is appropriate.
9. Troubleshooting common errors
| Symptom | Likely cause | What to do |
|---|---|---|
ModuleNotFoundError: No module named 'scipy' |
SciPy is not installed in the Python environment running the script. | Run python -m pip install scipy numpy with the same interpreter used to run the script; install Statsmodels separately if needed. |
| NaN statistic or p-value | Input contains missing or non-finite values, has insufficient observations, or the data produce an undefined calculation. | Inspect values and group sizes. Make an explicit, justified missing-data choice; do not conceal the issue by blindly omitting values. |
| Unexpectedly different result from another tool | The tools may use different variance assumptions, one- or two-sided alternatives, missing-data rules, or data filters. | Compare the exact observations, test variant, alternative, and handling options. Set SciPy’s equal_var and alternative explicitly. |
| Very small p-value but tiny effect | A p-value measures compatibility with a null model, not practical importance; large samples can detect small differences. | Report the effect estimate and interval, then assess practical importance in the subject area. |
| Non-significant result described as “no difference” | Failure to reject the null is being treated as proof of equality. | Describe the result as insufficient evidence to reject the null, and report the estimate and interval. If the goal is to establish equivalence, use a design and method intended for that question. |
| Code runs, but the conclusion seems implausible | Observations may be paired, clustered, duplicated, mislabeled, or measured on the wrong scale. | Recheck the unit of analysis, data construction, and test choice before interpreting the output. |
10. Or skip the browser setup
If your analysis also needs a clean screenshot of a results page, dashboard, or report, ScreenshotNeo is a website screenshot API and MCP server. It does not run hypothesis tests; it captures webpages. A single GET request can return PNG, JPEG, WebP, or PDF. The DIY Python example below captures a webpage to WebP:
See the ScreenshotNeo API documentation for parameters and options.
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
ScreenshotNeo accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000, and every feature is available on every plan.
Use the same one-call endpoint from other clients:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));
Sign up for ScreenshotNeo to get 1,000 screenshots a month free with no card.
FAQ
What does a p-value tell me?
It describes how likely data at least as extreme would be under the null model and test assumptions. It does not give the probability that the null is true.
Does p > 0.05 mean the groups are equal?
No. It means the test did not provide sufficient evidence to reject its null at that threshold. Equality needs a question and method designed to establish equivalence.
Should I use a one-sided test?
Only when a directional alternative was justified and chosen before examining the result. Otherwise, a two-sided test is generally the appropriate choice for detecting differences in either direction.
Why report a confidence interval?
It shows uncertainty around an effect estimate, making the direction and plausible magnitude more informative than a p-value alone.


