Descriptive Statistics
Summarising Data With a Few Numbers
Exploratory Data Analysis (EDA) is the stage where you get to know your data before modelling it. It starts with descriptive statistics: a handful of numbers that summarise thousands of rows. They answer two questions:
- Where is the centre? What is a typical value?
- How spread out is it? How much do values vary?
import pandas as pd
import numpy as np
np.random.seed(7)
salaries = pd.Series(np.random.lognormal(mean=12.6, sigma=0.5, size=1000)).round(-3)
salaries.describe()
describe() gives you count, mean, standard deviation, minimum, quartiles and maximum in one call. Learn to read its output fluently.
Measures of Centre
Mean
The arithmetic average: the sum divided by the count. salaries.mean()
Median
The middle value when the data is sorted. Half the values are below it and half above. salaries.median()
Mode
The most frequent value. It is most useful for categorical data: "What is the most common payment channel?" df["channel"].mode()
Mean vs median: the most important comparison
Imagine ten employees at a Lagos startup earning N300,000 a month, and the founder earning N10,000,000.
- Mean: (10 x 300,000 + 10,000,000) / 11 = about N1.18 million
- Median: N300,000
The mean suggests a typical employee earns over a million Naira. The median tells the truth. When the mean is much larger than the median, the data is right-skewed: a long tail of high values is pulling the mean up. Income, house prices, transaction sizes and social media followers are almost always right-skewed.
Measures of Spread
Range
Maximum minus minimum. It is simple, but one outlier can distort it.
Variance and standard deviation
Variance is the average squared distance from the mean. The standard deviation (std) is its square root, measured in the same units as the data. A small std means values cluster tightly around the mean; a large std means they are spread out.
salaries.std()
salaries.var()
Quartiles and the IQR
- Q1 (25th percentile): 25% of values are below it
- Q2 (50th percentile): the median
- Q3 (75th percentile): 75% of values are below it
- IQR = Q3 - Q1, the range of the middle 50% of the data
q1, q3 = salaries.quantile([0.25, 0.75])
iqr = q3 - q1
The IQR is robust: extreme values do not affect it, which makes it ideal for skewed data and for detecting outliers.
Distributions
A distribution describes how often each value occurs. Always plot it:
import seaborn as sns
import matplotlib.pyplot as plt
sns.histplot(salaries, bins=40, kde=True)
plt.axvline(salaries.mean(), color="red", label="Mean")
plt.axvline(salaries.median(), color="green", label="Median")
plt.legend()
plt.show()
Common shapes:
| Shape | Description | Example |
|---|---|---|
| Normal (bell curve) | Symmetric; mean is about equal to median | Heights, measurement errors |
| Right-skewed | Long tail to the right; mean > median | Income, transaction amounts |
| Left-skewed | Long tail to the left; mean < median | Age at retirement, exam scores on an easy test |
| Bimodal | Two peaks | Restaurant orders at lunch and dinner |
| Uniform | All values roughly equally likely | Rolling a fair die |
The normal distribution and the 68-95-99.7 rule
For normally distributed data:
- About 68% of values fall within 1 std of the mean
- About 95% fall within 2 std
- About 99.7% fall within 3 std
This is why a value more than 3 standard deviations from the mean is often flagged as unusual. Card fraud systems at banks apply similar ideas to spot transactions far outside a customer's normal spending.
Skewness in numbers
salaries.skew() # > 0 right-skewed, < 0 left-skewed, about 0 symmetric
A common fix for strong right skew before modelling is a log transform: np.log1p(salaries).
Describing Categorical Data
channel = pd.Series(["App", "USSD", "App", "POS", "App", "USSD", "Web", "App"])
channel.value_counts() # counts
channel.value_counts(normalize=True) # proportions
channel.mode()[0] # most common: 'App'
Putting It Together
When you first open a dataset, run this routine:
df.describe()for numeric columnsdf.describe(include="object")for text columns- Compare the mean and median of every key numeric column
- Plot a histogram of each key numeric column
- Run
value_counts()on each key categorical column
Use the lab below to build intuition: drag the values and watch how the mean, median and standard deviation react.
Try it yourself
Key Takeaways
- Descriptive statistics summarise the centre (mean, median, mode) and the spread (range, std, IQR) of your data.
- Compare the mean and the median: a large gap signals skew, and for skewed data such as income the median is the honest typical value.
- The standard deviation measures typical distance from the mean, while the IQR measures the spread of the middle 50% and is robust to outliers.
- Always plot distributions; normal, skewed, bimodal and uniform shapes each suggest different analysis choices.
- For normal data, about 68%, 95% and 99.7% of values fall within 1, 2 and 3 standard deviations of the mean.
Quick Quiz
1.In a dataset of customer transaction amounts, the mean is N85,000 and the median is N22,000. What does this tell you?
2.Why is the IQR described as a robust measure of spread?
3.For normally distributed data, roughly what percentage of values lie within 2 standard deviations of the mean?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx