Data Visualisation for EDA
Charts as a Thinking Tool
In module 2 you learned how to draw charts. In EDA, charts are not decoration: they are how you ask questions of the data. The statistician Francis Anscombe made the point in 1973 with four small datasets that have almost identical means, variances and correlations but look completely different when plotted. One is a clean line, one is a curve, and two are driven by a single outlier. Summary statistics alone would have hidden all of that.
The rule: never trust a statistic you have not plotted.
This lesson focuses on the four workhorse EDA charts, what each one reveals, and the mistakes to avoid.
import seaborn as sns
import matplotlib.pyplot as plt
sns.set_theme(style="whitegrid")
# df is the fintech dataset from the previous lesson
1. Histograms: What Does One Variable Look Like?
sns.histplot(data=df, x="income", bins=50, kde=True)
plt.title("Income distribution of app users")
plt.show()
Look for:
- Skew: a long right tail suggests a log transform before modelling
- Multiple peaks: two peaks often mean two different groups mixed together (for example, salary earners and business owners)
- Spikes at round numbers: a spike at exactly 0, 100,000 or 999 usually signals a default value or data entry habit
- Impossible values: negative ages, dates in the future
Pitfall: the number of bins changes the story. Try several values (10, 30, 100) before drawing conclusions.
Compare a distribution across groups with hue:
sns.histplot(data=df, x="monthly_txns", hue="churned", stat="density", common_norm=False, element="step")
2. Scatter Plots: How Do Two Numeric Variables Relate?
sns.scatterplot(data=df, x="tenure_months", y="monthly_volume", hue="channel", alpha=0.4)
plt.yscale("log")
plt.show()
Look for: direction (up or down), shape (straight or curved), strength (tight or loose), clusters, and outliers.
Pitfalls:
- Overplotting: with thousands of points, everything becomes a blob. Use
alpha=0.2, take a sample withdf.sample(1000), or switch tosns.histplot(x=..., y=...)for a 2D density view. - Skewed axes: a log scale often reveals structure hidden in a squashed corner.
sns.pairplot(df[["age", "income", "monthly_txns", "churned"]], hue="churned") draws a scatter plot for every pair of columns, which makes it a fast first look at a new dataset.
3. Box Plots: How Does a Number Vary Across Groups?
A box plot shows five numbers at once: the median (middle line), Q1 and Q3 (the box), the whiskers (up to 1.5 x IQR), and outliers (dots beyond the whiskers).
sns.boxplot(data=df, x="state", y="income", order=df.groupby("state")["income"].median().sort_values().index)
plt.yscale("log")
plt.xticks(rotation=30)
plt.show()
Look for: differences in median between groups, differences in spread, and which groups have many outliers.
Tip: a violin plot (sns.violinplot) shows the full distribution shape per group. It is useful when a box plot might hide two peaks.
4. Heatmaps: How Do Many Variables Relate at Once?
corr = df.select_dtypes("number").corr()
mask = np.triu(np.ones_like(corr, dtype=bool)) # hide the duplicate upper triangle
sns.heatmap(corr, mask=mask, annot=True, fmt=".2f", cmap="RdBu_r", vmin=-1, vmax=1)
Heatmaps also work for any two-dimensional summary, such as a pivot table of churn rate by state and channel, or transactions by day of week and hour of day. That second one is how payment companies like Paystack and Stripe spot peak load times.
Pitfall: always fix vmin and vmax for correlations and use a diverging colour map, so that 0 is neutral and the colours mean the same thing on every chart.
Choosing the Right Chart for the Question
| Question | Chart |
|---|---|
| What does the distribution of X look like? | Histogram / KDE |
| Is X related to Y (both numeric)? | Scatter plot |
| How does numeric X differ across categories? | Box plot / violin plot |
| How do many numeric variables correlate? | Correlation heatmap |
| How does a metric vary over two categories? | Pivot table heatmap |
| How does X change over time? | Line chart |
| How do category counts compare? | Bar chart (sorted) |
An EDA Visual Checklist
For every new dataset, produce at least:
- A histogram of every key numeric column
- A bar chart of every key categorical column
- A correlation heatmap
- Box plots of the target variable (for example, churn or spend) across the main categories
- Scatter plots of the target against its most correlated features
Write one sentence under each chart describing what you learned. Those sentences become the backbone of your EDA report.
Practise: The lab below gives you real analyst questions. Pick the right chart for each one before you move on.
Try it yourself
Key Takeaways
- In EDA, charts are a thinking tool; as Anscombe's quartet shows, never trust a statistic you have not plotted.
- Histograms reveal a variable's shape, skew, multiple peaks and suspicious spikes; try several bin counts.
- Scatter plots reveal relationships between two numeric variables; fix overplotting with alpha, sampling or density plots.
- Box plots compare the median, spread and outliers across groups, and heatmaps summarise many correlations or two-way pivots at once.
- Follow a visual checklist for every dataset and write a one-sentence takeaway under each chart.
Quick Quiz
1.What did Anscombe's quartet demonstrate?
2.Your scatter plot of 200,000 transactions is a solid blob of colour. What is the best fix?
3.In a box plot, what do the dots beyond the whiskers usually represent?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx