Descriptive Statistics
Why Statistics for Data Analysts?
Statistics is the mathematical foundation of data analysis. Without it, you can describe data but you cannot draw reliable conclusions, quantify uncertainty, or make valid comparisons. Business decisions made on "the data looks higher" rather than statistically sound analysis are often wrong.
You do not need to be a statistician to be an effective data analyst. But understanding the core statistical concepts -- measures of central tendency, spread, distributions, and correlation -- will make your analysis significantly more accurate and credible.
Measures of Central Tendency
Central tendency describes the "typical" or "centre" of a dataset.
Mean (Average)
The sum of all values divided by the count.
import numpy as np
import pandas as pd
values = [120, 450, 89, 730, 200, 340, 1500, 95]
mean = np.mean(values) # 440.5
When to use: Best when the distribution is roughly symmetrical and has no extreme outliers.
Limitation: Highly sensitive to outliers. One extreme value pulls the mean significantly. In a team of 10 people earning $50,000 and one person earning $1,000,000, the mean salary is $136,364 -- not representative of anyone's actual salary.
Median
The middle value when data is sorted. If there is an even number of values, the median is the average of the two middle values.
median = np.median(values) # 270.0 (average of 200 and 340)
When to use: Better than mean when data is skewed or contains outliers. Income, house prices, and response times are typically median-reported for this reason.
Mode
The most frequently occurring value.
from scipy import stats
mode = stats.mode(values)
When to use: The only measure of central tendency applicable to categorical data (e.g., most common product category, most frequent complaint type).
Mean vs Median: Which to Use?
| Situation | Use |
|---|---|
| Symmetric distribution, no outliers | Either -- they will be similar |
| Skewed distribution | Median |
| Data with significant outliers | Median |
| Categorical data | Mode |
| Reporting income, home prices, order value | Median |
| Reporting exam scores, physical measurements | Mean |
Measures of Spread
Spread describes how much variation exists in the data.
Range
The difference between maximum and minimum values.
data_range = max(values) - min(values) # 1500 - 89 = 1411
Limitation: Highly sensitive to a single extreme value.
Variance and Standard Deviation
Variance measures the average squared deviation from the mean. Standard deviation is the square root of variance (expressed in the same units as the original data).
variance = np.var(values, ddof=1) # sample variance (ddof=1)
std_dev = np.std(values, ddof=1) # sample standard deviation
# About 68% of data falls within 1 standard deviation of the mean
# About 95% falls within 2 standard deviations (for normal distributions)
Higher standard deviation = more spread out. Lower = more concentrated around the mean.
Interquartile Range (IQR)
The range of the middle 50% of data (Q3 - Q1). Resistant to outliers.
Q1 = np.percentile(values, 25)
Q3 = np.percentile(values, 75)
IQR = Q3 - Q1
# Outlier bounds (Tukey's rule)
lower_bound = Q1 - 1.5 * IQR
upper_bound = Q3 + 1.5 * IQR
Percentiles and Quartiles
A percentile indicates the value below which a given percentage of observations fall.
# Common percentiles
p25 = np.percentile(values, 25) # Q1: 25th percentile
p50 = np.percentile(values, 50) # Median: 50th percentile
p75 = np.percentile(values, 75) # Q3: 75th percentile
p90 = np.percentile(values, 90) # 90th percentile
p99 = np.percentile(values, 99) # 99th percentile
# The describe() function shows all of these
pd.Series(values).describe()
Percentiles are widely used in analytics:
- P95 and P99 response times in web performance monitoring
- Salary percentiles in compensation benchmarking
- Revenue percentiles for customer segmentation
Distribution Shape
Understanding the shape of a distribution is important for choosing the right statistical methods and for understanding patterns in data.
Symmetry and Skewness
Symmetric distribution: Mean approximately equals median. Bell-shaped.
Right-skewed (positive skew): A long tail on the right. Mean > Median. Examples: income, order values, company sizes. Most values are low, but a few are very high.
Left-skewed (negative skew): A long tail on the left. Mean < Median. Less common in business data.
from scipy.stats import skew
print(f"Skewness: {skew(values):.2f}")
# Positive value = right skewed, negative = left skewed, near 0 = symmetric
Kurtosis
Measures how heavy the tails of the distribution are relative to a normal distribution. High kurtosis means more extreme outliers than a normal distribution would predict.
Summary Statistics in Practice
# Complete summary for a business dataset
df = pd.read_csv('orders.csv')
summary = df['order_amount'].agg([
('count', 'count'),
('mean', 'mean'),
('median', 'median'),
('std', 'std'),
('min', 'min'),
('q25', lambda x: x.quantile(0.25)),
('q75', lambda x: x.quantile(0.75)),
('max', 'max')
]).round(2)
print(summary)
# Check for skewness
print(f"Skewness: {df['order_amount'].skew():.2f}")
print(f"Mean vs Median: {df['order_amount'].mean():.2f} vs {df['order_amount'].median():.2f}")
Key Takeaways
- Mean is sensitive to outliers; median is robust. For skewed data (income, prices, response times), always report the median.
- Standard deviation quantifies spread around the mean; IQR quantifies the spread of the middle 50% and is robust to outliers.
- Percentiles (P25, P50, P75, P90, P99) provide more detailed distribution understanding than min/max alone.
- Positive (right) skew means most values are low with a few very high values -- the pattern for revenue, income, and order value in most businesses.
- Always visualise your data with a histogram or box plot alongside numeric statistics -- the shape tells you which statistics are most appropriate to use.
Practice Exercise
import pandas as pd
import numpy as np
from scipy.stats import skew
import matplotlib.pyplot as plt
# Generate sample e-commerce order data
np.random.seed(42)
orders = np.concatenate([
np.random.exponential(scale=150, size=900), # most orders
np.random.uniform(1000, 5000, size=100) # high-value orders
])
# 1. Calculate mean, median, mode, std dev, IQR
print(f"Mean: {np.mean(orders):.2f}")
print(f"Median: {np.median(orders):.2f}")
print(f"Std Dev: {np.std(orders, ddof=1):.2f}")
q1, q3 = np.percentile(orders, [25, 75])
print(f"IQR: {q3-q1:.2f}")
print(f"Skewness: {skew(orders):.2f}")
# 2. Plot histogram and box plot side by side
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(12, 4))
ax1.hist(orders, bins=50, color='steelblue', edgecolor='white')
ax1.axvline(np.mean(orders), color='red', linestyle='--', label=f'Mean: {np.mean(orders):.0f}')
ax1.axvline(np.median(orders), color='green', linestyle='-', label=f'Median: {np.median(orders):.0f}')
ax1.legend(); ax1.set_title('Order Value Distribution')
ax2.boxplot(orders, vert=True); ax2.set_title('Box Plot')
plt.tight_layout(); plt.show()
# 3. Explain why mean and median differ so much in this dataset
Try it yourself
Key Takeaways
- Mean is sensitive to outliers; use median for skewed data like income, house prices, and order values.
- Standard deviation measures spread around the mean; IQR measures the spread of the middle 50% and is robust to outliers.
- Percentiles (P25, P50, P75, P90, P99) describe distribution in more detail than min/max alone.
- Right (positive) skew means most values are low with a few very high values -- the typical pattern for business revenue data.
- Always visualise data alongside numeric statistics -- summary statistics alone can describe very different distributions with identical numbers.
Quick Quiz
1.A company reports that the average employee salary is $95,000, but most employees earn between $45,000 and $65,000. What is the most likely explanation?
2.What does the interquartile range (IQR) measure?
3.What does positive (right) skewness in a distribution indicate?
4.Why should you always visualise data (histogram or box plot) alongside numeric summary statistics?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx