Hypothesis Testing
What is Hypothesis Testing?
Hypothesis testing is a statistical method for determining whether an observed difference is likely to be real (statistically significant) or could simply be due to random chance. It is the statistical backbone of A/B testing, product experiments, quality control, and scientific research.
Without hypothesis testing, you cannot reliably answer questions like:
- Is the new checkout flow actually converting better, or did we just get lucky in our sample?
- Is the difference in revenue between two regions meaningful, or is it within normal variation?
- Does this marketing campaign actually increase customer lifetime value?
The Hypothesis Testing Framework
Step 1: Define the Hypotheses
Null Hypothesis (H0): The default assumption -- no effect, no difference. "The new button colour has no effect on conversion rate."
Alternative Hypothesis (H1): What you are trying to demonstrate. "The new button colour increases conversion rate."
You never "prove" the alternative hypothesis. You either:
- Reject the null hypothesis (evidence suggests H1 is true)
- Fail to reject the null hypothesis (insufficient evidence to support H1)
Step 2: Choose a Significance Level (alpha)
The significance level (alpha) is the threshold for deciding whether a result is statistically significant. Common values:
- alpha = 0.05 (5%): Most common in business. Accept a 5% probability of false positives.
- alpha = 0.01 (1%): More conservative. Used when false positives are costly.
Step 3: Calculate the Test Statistic and P-value
The p-value is the probability of observing a result at least as extreme as the one you observed, assuming the null hypothesis is true.
- p-value < alpha: Reject H0 (result is statistically significant)
- p-value >= alpha: Fail to reject H0 (result is not statistically significant)
Step 4: Make a Decision
Based on the p-value and significance level, make a clear decision and communicate the practical implications.
The T-Test
The t-test compares the means of one or two groups to determine if they are statistically different.
One-Sample T-Test
Compare a sample mean to a known or hypothesised population value.
from scipy.stats import ttest_1samp
import numpy as np
# Is the average order value significantly different from $100?
order_values = [112, 89, 134, 98, 145, 87, 123, 109, 91, 118]
stat, p_value = ttest_1samp(order_values, popmean=100)
print(f"T-statistic: {stat:.3f}")
print(f"P-value: {p_value:.4f}")
print(f"Significant at 0.05: {p_value < 0.05}")
Two-Sample T-Test (Independent Groups)
Compare the means of two independent groups.
from scipy.stats import ttest_ind
# Compare conversion rates for two different landing pages
page_a = [45, 52, 48, 41, 55, 49, 43, 51, 47, 53] # conversion amounts
page_b = [58, 61, 55, 63, 57, 59, 62, 60, 58, 64]
stat, p_value = ttest_ind(page_a, page_b)
print(f"T-statistic: {stat:.3f}")
print(f"P-value: {p_value:.4f}")
print(f"Significant difference: {p_value < 0.05}")
Paired T-Test
Compare the same group measured at two different times.
from scipy.stats import ttest_rel
# Customer satisfaction before and after a service improvement
before = [3.2, 3.8, 3.5, 4.1, 3.0, 3.7, 4.2, 3.6, 3.9, 3.4]
after = [3.8, 4.1, 3.9, 4.5, 3.4, 4.0, 4.6, 4.0, 4.3, 3.8]
stat, p_value = ttest_rel(before, after)
print(f"P-value: {p_value:.4f}")
print(f"Improvement is significant: {p_value < 0.05}")
Chi-Square Test
The chi-square test is used when comparing proportions across categories (not means).
from scipy.stats import chi2_contingency
import numpy as np
# Do conversion rates differ by device type?
# Rows: Converted / Not Converted, Columns: Mobile / Desktop / Tablet
observed = np.array([
[120, 350, 80], # Converted
[880, 650, 420] # Not converted
])
stat, p_value, dof, expected = chi2_contingency(observed)
print(f"Chi-square statistic: {stat:.3f}")
print(f"P-value: {p_value:.4f}")
print(f"Degrees of freedom: {dof}")
print(f"Conversion differs by device: {p_value < 0.05}")
A/B Testing in Practice
A/B testing is the most common application of hypothesis testing in business. The standard process:
- Define the metric: What are you trying to move? (Conversion rate, revenue per visitor, retention rate)
- Define H0 and H1: H0: No difference. H1: Version B performs better.
- Calculate required sample size: Use a sample size calculator before running the test. Running a test without adequate sample size leads to underpowered results.
- Run the experiment: Randomly assign visitors to A or B.
- Wait for statistical significance: Do not stop the test early just because you see a direction.
- Analyse and decide: Apply the t-test or chi-square test, check p-value, make a decision.
Sample Size Calculation
# How many users per group do we need?
# Power analysis using statsmodels
from statsmodels.stats.power import tt_ind_solve_power
required_n = tt_ind_solve_power(
effect_size=0.3, # expected standardised effect size
alpha=0.05, # significance level
power=0.8, # desired statistical power (80%)
alternative='larger' # one-tailed test
)
print(f"Required sample size per group: {int(required_n) + 1}")
Common Mistakes in Hypothesis Testing
Stopping tests early (peeking): Running a test, checking results daily, and stopping when p < 0.05 inflates the false positive rate. Pre-register the sample size and run the full test.
P-hacking: Running multiple tests and only reporting significant results. With 20 tests at alpha=0.05, you expect one false positive by chance.
Confusing statistical significance with practical significance: A result can be statistically significant (p < 0.05) but practically meaningless (the effect is too small to matter). Always report effect size alongside p-values.
Wrong test for the data: Using a t-test to compare proportions (use chi-square) or using a parametric test on very small non-normal samples (use non-parametric alternatives like Mann-Whitney U).
Practical Interpretation
def interpret_test(stat, p_value, alpha=0.05, effect_size=None):
print(f"Test statistic: {stat:.3f}")
print(f"P-value: {p_value:.4f}")
if p_value < alpha:
print(f"Result: STATISTICALLY SIGNIFICANT (p < {alpha})")
print("We reject the null hypothesis.")
else:
print(f"Result: NOT SIGNIFICANT (p >= {alpha})")
print("We fail to reject the null hypothesis.")
if effect_size:
print(f"Effect size: {effect_size:.3f} (small: 0.2, medium: 0.5, large: 0.8)")
Key Takeaways
- Hypothesis testing determines if an observed difference is likely real or due to random chance using the null hypothesis, significance level, and p-value.
- P-value < 0.05 (alpha): reject H0 (statistically significant). P-value >= 0.05: fail to reject H0 (insufficient evidence).
- T-tests compare means of one or two groups; chi-square tests compare proportions across categories.
- Calculate required sample size before running an A/B test -- running without adequate power leads to underpowered, unreliable results.
- Never confuse statistical significance with practical significance -- a tiny, meaningless difference can be statistically significant with a large enough sample.
Practice Exercise
from scipy.stats import ttest_ind, chi2_contingency
import numpy as np
# Scenario 1: A/B test for email subject lines
# Campaign A: 20 out of 200 recipients clicked
# Campaign B: 35 out of 200 recipients clicked
# Is Campaign B significantly better?
a_clicked, a_total = 20, 200
b_clicked, b_total = 35, 200
observed = np.array([
[a_clicked, b_clicked],
[a_total - a_clicked, b_total - b_clicked]
])
stat, p, dof, expected = chi2_contingency(observed)
print(f"P-value: {p:.4f}")
print(f"Campaign B significantly better: {p < 0.05}")
# Scenario 2: Order values after a redesign
before_redesign = np.random.normal(85, 25, 150)
after_redesign = np.random.normal(92, 28, 150)
stat, p = ttest_ind(before_redesign, after_redesign)
print(f"\nOrder value test P-value: {p:.4f}")
print(f"Significant improvement: {p < 0.05}")
# Calculate and report the effect size (Cohen's d)
pooled_std = np.sqrt((np.std(before_redesign)**2 + np.std(after_redesign)**2) / 2)
cohens_d = (np.mean(after_redesign) - np.mean(before_redesign)) / pooled_std
print(f"Effect size (Cohen's d): {cohens_d:.3f}")
Try it yourself
Key Takeaways
- Hypothesis testing determines if observed differences are statistically significant or could be due to random chance.
- P-value < alpha (typically 0.05): reject the null hypothesis (statistically significant); p >= alpha: fail to reject.
- T-tests compare means; chi-square tests compare proportions -- choose the right test for your data type.
- Calculate required sample size before running A/B tests; never stop a test early because you see a favorable direction (peeking inflates false positives).
- Statistical significance does not equal practical significance -- a tiny, meaningless effect can be statistically significant with a large enough sample.
Quick Quiz
1.What does a p-value of 0.03 mean in hypothesis testing?
2.What is the correct test when comparing conversion rates (a proportion) across multiple device types (mobile, desktop, tablet)?
3.What is the most serious problem with 'peeking' at A/B test results and stopping when p < 0.05?
4.An A/B test shows that the new headline generates 0.1% more conversions with p=0.001. What is the most important question to ask before concluding the test was a success?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx