Distributions and Probability
What is a Probability Distribution?
A probability distribution describes how the values of a variable are distributed -- what values are possible and how likely each is. Understanding distributions helps you choose the right statistical methods, make valid inferences from samples, and set realistic expectations.
The Normal Distribution
The normal distribution (also called the Gaussian distribution or "bell curve") is the most important distribution in statistics. It is symmetric around the mean and fully described by two parameters: mean (mu) and standard deviation (sigma).
Properties
- Symmetric: Mean = Median = Mode
- 68-95-99.7 Rule:
- 68% of data falls within 1 standard deviation of the mean
- 95% falls within 2 standard deviations
- 99.7% falls within 3 standard deviations
import numpy as np
import matplotlib.pyplot as plt
from scipy import stats
# Generate a normal distribution
mean, std = 170, 10 # e.g., heights in cm
x = np.linspace(mean - 4*std, mean + 4*std, 1000)
y = stats.norm.pdf(x, mean, std)
plt.figure(figsize=(10, 5))
plt.plot(x, y, linewidth=2, color='#059669')
plt.fill_between(x, y, where=(x >= mean - std) & (x <= mean + std),
alpha=0.2, color='#059669', label='68% (1 SD)')
plt.fill_between(x, y, where=((x >= mean - 2*std) & (x < mean - std)) |
((x > mean + std) & (x <= mean + 2*std)),
alpha=0.15, color='#059669', label='95% (2 SD)')
plt.title(f'Normal Distribution (mean={mean}, std={std})')
plt.legend()
plt.show()
What Follows a Normal Distribution?
- Human physical measurements (height, weight, blood pressure)
- Test scores in large groups
- Measurement errors in scientific instruments
- Aggregated daily returns in financial markets (approximately)
Many business metrics do NOT follow a normal distribution (revenue, customer counts, order values are typically right-skewed).
The Central Limit Theorem
The Central Limit Theorem (CLT) is one of the most important theorems in statistics. It states:
The distribution of sample means approaches a normal distribution as sample size increases, regardless of the shape of the population distribution.
In practice: even if the underlying data is skewed, the average of many samples from that data will be approximately normally distributed for sample sizes of n >= 30.
This is why many statistical tests (t-tests, ANOVA, regression) can be applied to non-normally distributed data as long as sample sizes are adequate -- the test statistics based on sample means will be approximately normally distributed.
# Demonstrating CLT: sample means from a skewed distribution
np.random.seed(42)
# Exponential distribution (very right-skewed)
population = np.random.exponential(scale=100, size=10000)
# Take 1000 samples of size 30 and calculate each sample's mean
sample_means = [np.mean(np.random.choice(population, size=30)) for _ in range(1000)]
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(12, 4))
ax1.hist(population, bins=50); ax1.set_title('Original Population (right-skewed)')
ax2.hist(sample_means, bins=50); ax2.set_title('Distribution of Sample Means (approx. normal)')
plt.tight_layout(); plt.show()
Other Important Distributions
Binomial Distribution
Used when counting the number of successes in a fixed number of binary trials (success/failure).
from scipy.stats import binom
# Email campaign: 10,000 emails, 2% click rate
n = 10000 # number of trials
p = 0.02 # probability of success
# Expected number of clicks
expected = n * p # 200
# Probability of getting fewer than 150 clicks
prob_less_than_150 = binom.cdf(149, n, p)
Business applications: Conversion rates, defect rates, success/failure events.
Poisson Distribution
Used to model the number of events occurring in a fixed time interval when events occur independently at a constant average rate.
from scipy.stats import poisson
# Average 8 customer service calls per hour
lambda_rate = 8
# Probability of receiving exactly 10 calls in an hour
prob_10_calls = poisson.pmf(10, lambda_rate)
# Probability of receiving 12 or more calls (need extra staffing)
prob_12_or_more = 1 - poisson.cdf(11, lambda_rate)
Business applications: Website traffic per minute, customer arrivals, equipment failures per day.
Probability Basics
Probability Rules
- P(event) ranges from 0 (impossible) to 1 (certain)
- P(not A) = 1 - P(A)
- P(A or B) = P(A) + P(B) - P(A and B) (if mutually exclusive: P(A) + P(B))
- P(A and B) = P(A) * P(B) (only if A and B are independent)
Conditional Probability
P(A | B) means "the probability of A given that B has occurred":
P(A | B) = P(A and B) / P(B)
Business example: What is the probability a customer makes a second purchase, given that they used a discount code on their first purchase?
Using Distributions for Business Analysis
Setting realistic targets
If monthly revenue follows a normal distribution with mean $500K and std $50K, you can calculate:
from scipy.stats import norm
# Probability of exceeding $600K in a month
prob_exceed_600k = 1 - norm.cdf(600000, loc=500000, scale=50000)
# About 2.3% probability
# Revenue value that will be exceeded 90% of the time (conservative target)
conservative_target = norm.ppf(0.10, loc=500000, scale=50000)
# About $436K -- will be exceeded 90% of months
Anomaly detection
Flag observations that are statistically unlikely under the expected distribution:
# Z-score: how many standard deviations from the mean?
z_scores = (df['revenue'] - df['revenue'].mean()) / df['revenue'].std()
anomalies = df[abs(z_scores) > 3] # flag values > 3 std devs from mean
Key Takeaways
- The normal distribution is symmetric with the 68-95-99.7 rule: 68% within 1 SD, 95% within 2 SD, 99.7% within 3 SD.
- The Central Limit Theorem means that for large enough samples (n >= 30), sample means are approximately normally distributed even when underlying data is skewed.
- Binomial distribution models binary outcomes (success/failure) across fixed trials; Poisson models event counts per time period.
- Z-scores measure how many standard deviations a value is from the mean -- values beyond 3 SD are statistically unusual.
- Use known distributions to calculate probabilities, set realistic targets, and detect anomalies in business data.
Practice Exercise
import numpy as np
from scipy import stats
import matplotlib.pyplot as plt
# 1. You run an e-commerce site with a 3% conversion rate and 5000 daily visitors.
# Using the binomial distribution, calculate:
# a. Expected number of conversions per day
# b. Probability of fewer than 120 conversions today
# c. Probability of more than 200 conversions today
from scipy.stats import binom
n, p = 5000, 0.03
print(f"Expected conversions: {n*p}")
print(f"P(less than 120): {binom.cdf(119, n, p):.4f}")
print(f"P(more than 200): {1-binom.cdf(200, n, p):.4f}")
# 2. Your support team receives an average of 15 tickets per hour (Poisson).
# Calculate the probability of receiving more than 25 tickets in one hour.
from scipy.stats import poisson
print(f"P(more than 25 tickets): {1-poisson.cdf(25, 15):.4f}")
# 3. Demonstrate the Central Limit Theorem with a right-skewed dataset:
# - Generate 10000 values from an exponential distribution
# - Take 1000 samples of size 50 and calculate each sample mean
# - Plot both distributions and comment on the result
Try it yourself
Key Takeaways
- The normal distribution is fully described by mean and standard deviation, with 68-95-99.7% of data within 1-2-3 standard deviations.
- The Central Limit Theorem: sample means are approximately normally distributed for n >= 30, enabling parametric tests on non-normal data.
- Binomial distribution models binary outcomes across fixed trials (conversion rates); Poisson models event counts per time period (tickets, arrivals).
- Z-scores standardise any value to units of standard deviations -- values beyond 3 SD are statistically unusual and worth investigating.
- Use known distributions to calculate probabilities, set realistic targets, and perform statistically valid anomaly detection.
Quick Quiz
1.What does the Central Limit Theorem state?
2.Which probability distribution is most appropriate for modelling the number of customer support tickets received per hour?
3.In a normal distribution, approximately what percentage of data falls within 2 standard deviations of the mean?
4.What is a Z-score and how is it used in anomaly detection?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx