Exploratory Data Analysis
What is Exploratory Data Analysis?
Exploratory Data Analysis (EDA) is the process of investigating a dataset to understand its structure, distributions, relationships, and patterns before formal analysis or modelling. It is the analytical equivalent of reading a brief before writing a report -- you need to understand what you are working with before drawing conclusions.
EDA answers questions like:
- What are the distributions of key variables?
- Are there outliers or unexpected values?
- How are variables correlated?
- What patterns exist across different segments?
- What questions does this data raise or answer?
Libraries for EDA
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
# Set a clean plotting style
plt.style.use('seaborn-v0_8-whitegrid')
sns.set_palette('Set2')
Step 1: Understand the Data
df = pd.read_csv('sales_data.csv', parse_dates=['order_date'])
# Shape and structure
print(df.shape)
print(df.dtypes)
print(df.head())
# Statistical summary
df.describe()
# For non-numeric columns
df.describe(include='object')
Step 2: Understand Distributions
Numeric distributions
# Histogram: distribution of order amounts
plt.figure(figsize=(10, 4))
df['amount'].hist(bins=50, edgecolor='white')
plt.title('Distribution of Order Amounts')
plt.xlabel('Amount')
plt.ylabel('Count')
plt.show()
# Box plot: spot outliers
plt.figure(figsize=(8, 4))
df['amount'].plot(kind='box')
plt.title('Order Amount Distribution')
plt.show()
# With seaborn (more options)
fig, axes = plt.subplots(1, 2, figsize=(12, 4))
sns.histplot(df['amount'], bins=50, ax=axes[0])
sns.boxplot(y=df['amount'], ax=axes[1])
plt.tight_layout()
plt.show()
Categorical distributions
# Count of each category
df['status'].value_counts()
# Bar chart
df['country'].value_counts().head(10).plot(kind='bar', figsize=(10, 4))
plt.title('Orders by Country (Top 10)')
plt.xticks(rotation=45, ha='right')
plt.tight_layout()
plt.show()
Step 3: Explore Trends Over Time
# Monthly revenue trend
monthly = df.groupby(df['order_date'].dt.to_period('M'))['amount'].sum().reset_index()
monthly['order_date'] = monthly['order_date'].astype(str)
plt.figure(figsize=(12, 4))
plt.plot(monthly['order_date'], monthly['amount'], marker='o', linewidth=2)
plt.title('Monthly Revenue Trend')
plt.xlabel('Month')
plt.ylabel('Revenue')
plt.xticks(rotation=45, ha='right')
plt.tight_layout()
plt.show()
Step 4: Explore Segment Performance
# Revenue by country
country_revenue = df.groupby('country').agg(
total_revenue=('amount', 'sum'),
order_count=('order_id', 'count'),
avg_order=('amount', 'mean')
).sort_values('total_revenue', ascending=False).head(10)
print(country_revenue)
# Grouped bar chart
fig, ax = plt.subplots(figsize=(12, 5))
country_revenue['total_revenue'].plot(kind='bar', ax=ax)
ax.set_title('Revenue by Country (Top 10)')
ax.set_xlabel('Country')
ax.set_ylabel('Total Revenue')
plt.xticks(rotation=45, ha='right')
plt.tight_layout()
plt.show()
Step 5: Explore Correlations
# Correlation matrix for numeric columns
corr = df[['amount', 'orders', 'discount', 'customer_age']].corr()
print(corr)
# Heatmap visualisation
plt.figure(figsize=(8, 6))
sns.heatmap(corr, annot=True, fmt='.2f', cmap='RdYlGn', center=0)
plt.title('Correlation Matrix')
plt.tight_layout()
plt.show()
# Scatter plot: relationship between two variables
plt.figure(figsize=(8, 5))
plt.scatter(df['orders'], df['amount'], alpha=0.3)
plt.xlabel('Number of Orders')
plt.ylabel('Total Amount')
plt.title('Orders vs Amount')
plt.show()
# With regression line (seaborn)
sns.lmplot(data=df, x='orders', y='amount', scatter_kws={'alpha': 0.2})
plt.title('Orders vs Amount with Trend Line')
plt.show()
Step 6: Detect Outliers
# IQR method for outlier detection
Q1 = df['amount'].quantile(0.25)
Q3 = df['amount'].quantile(0.75)
IQR = Q3 - Q1
lower_bound = Q1 - 1.5 * IQR
upper_bound = Q3 + 1.5 * IQR
outliers = df[(df['amount'] < lower_bound) | (df['amount'] > upper_bound)]
print(f"Outliers: {len(outliers)} rows ({len(outliers)/len(df)*100:.1f}%)")
print(outliers['amount'].describe())
EDA Documentation: The Analytical Report
As you explore, document your findings. A good EDA report includes:
- Data overview: size, time period, source, key columns
- Data quality: missing values, duplicates, type issues found and how resolved
- Key distributions: for each important numeric column, describe the distribution (mean, median, skew, outliers)
- Segment analysis: which segments are largest, best-performing, most interesting
- Trend analysis: what patterns exist over time
- Correlations and relationships: what variables relate to each other
- Hypotheses generated: what the data suggests that should be investigated further
Key Takeaways
- EDA investigates distributions, trends, segment performance, correlations, and outliers before formal analysis or modelling.
- Always start with univariate analysis (individual columns) before moving to bivariate (relationships between two columns) and multivariate analysis.
- Combine numeric summaries (.describe()) with visualisations (histograms, box plots, bar charts, scatter plots) -- each reveals different aspects.
- Use the IQR method (Q1 - 1.5IQR, Q3 + 1.5IQR) to systematically identify outliers in numeric columns.
- Document findings as you explore -- the insights you generate during EDA are often as valuable as the final analysis.
Practice Exercise
Use the following public dataset to perform a complete EDA:
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
# Complete these EDA steps:
# 1. Assess data quality (missing values, types, duplicates)
# 2. Distribution of 'Fare' and 'Age' -- histogram and box plot
# 3. Survival rate by passenger class (Pclass)
# 4. Correlation between numeric columns -- heatmap
# 5. Outlier detection in 'Fare' using IQR method
# 6. Time breakdown is not available -- analyse by embarked port (Embarked)
# Document 3 insights from your exploration
Try it yourself
Key Takeaways
- EDA investigates distributions, trends, segment performance, correlations, and outliers before formal analysis.
- Follow a systematic workflow: data overview, distributions, time trends, segment analysis, correlations, outlier detection.
- Use both numeric summaries (.describe()) and visualisations (histograms, box plots, heatmaps) -- they reveal different aspects of the data.
- The IQR method (flag values below Q1-1.5*IQR or above Q3+1.5*IQR) is a robust approach to systematic outlier detection.
- Document insights as you explore -- the findings generated during EDA are often as valuable as any subsequent formal analysis.
Quick Quiz
1.What is the primary purpose of Exploratory Data Analysis (EDA)?
2.What is the IQR method for outlier detection?
3.What does a correlation coefficient of -0.8 between two variables indicate?
4.Why is it important to document your findings during EDA, not just at the end?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx