Data Visualisation with Python
The Python Visualisation Ecosystem
Python offers multiple visualisation libraries, each suited to different use cases:
Matplotlib: The foundational library. Full control over every element. Verbose but powerful. Seaborn: Statistical visualisation built on matplotlib. Simpler API for common analytical charts. Plotly: Interactive charts (hover, zoom, click). Excellent for dashboards and web embedding. Pandas .plot(): Quick charts directly from DataFrames, uses matplotlib underneath.
For most analytical work, use seaborn for static publication-quality charts and Plotly for interactive dashboards.
import matplotlib.pyplot as plt
import seaborn as sns
import plotly.express as px
import pandas as pd
# Global style settings
plt.style.use('seaborn-v0_8-whitegrid')
sns.set_palette('Set2')
Line Charts: Trends Over Time
# Monthly revenue trend (matplotlib)
fig, ax = plt.subplots(figsize=(12, 5))
ax.plot(monthly['month'], monthly['revenue'], marker='o', linewidth=2, color='#059669')
ax.fill_between(monthly['month'], monthly['revenue'], alpha=0.1, color='#059669')
ax.set_title('Monthly Revenue Trend', fontsize=14, fontweight='bold')
ax.set_xlabel('Month')
ax.set_ylabel('Revenue ($)')
plt.xticks(rotation=45, ha='right')
plt.tight_layout()
plt.savefig('monthly_trend.png', dpi=150, bbox_inches='tight')
plt.show()
# Interactive with Plotly (much simpler)
fig = px.line(monthly, x='month', y='revenue',
title='Monthly Revenue Trend',
markers=True)
fig.show()
Bar Charts: Comparing Categories
# Horizontal bar chart (better for many categories)
top_countries = df.groupby('country')['revenue'].sum().nlargest(10).reset_index()
fig, ax = plt.subplots(figsize=(10, 6))
bars = ax.barh(top_countries['country'], top_countries['revenue'], color='#059669')
ax.set_title('Revenue by Country (Top 10)', fontsize=14, fontweight='bold')
ax.set_xlabel('Total Revenue ($)')
# Add value labels
for bar in bars:
ax.text(bar.get_width(), bar.get_y() + bar.get_height()/2,
f'${bar.get_width():,.0f}', va='center', ha='left', fontsize=9)
plt.tight_layout()
plt.show()
# Grouped bar chart (compare two metrics across categories)
fig = px.bar(country_data, x='country', y=['revenue', 'orders'],
barmode='group', title='Revenue and Orders by Country')
fig.show()
Distributions: Histograms and Box Plots
# Distribution of order amounts
fig, axes = plt.subplots(1, 2, figsize=(14, 5))
# Histogram
sns.histplot(df['amount'], bins=50, kde=True, ax=axes[0], color='#059669')
axes[0].set_title('Distribution of Order Amounts')
axes[0].set_xlabel('Amount ($)')
# Box plot by segment
sns.boxplot(data=df, x='segment', y='amount', ax=axes[1])
axes[1].set_title('Amount Distribution by Segment')
axes[1].set_xlabel('Segment')
axes[1].set_ylabel('Amount ($)')
plt.tight_layout()
plt.show()
# Violin plot (shows distribution shape AND box plot stats)
sns.violinplot(data=df, x='segment', y='amount')
plt.title('Order Amount Distribution by Segment')
plt.show()
Scatter Plots: Relationships
# Scatter plot with colour coding
fig = px.scatter(df,
x='orders',
y='revenue',
color='country',
size='discount',
hover_data=['customer_name'],
title='Orders vs Revenue by Country')
fig.show()
# Seaborn scatter with regression line
sns.lmplot(data=df, x='orders', y='revenue',
hue='segment',
scatter_kws={'alpha': 0.4},
line_kws={'linewidth': 2})
plt.title('Orders vs Revenue with Trend Lines by Segment')
plt.show()
Heatmaps: Correlation and Grid Data
# Correlation heatmap
numeric_cols = df.select_dtypes(include='number').columns
corr_matrix = df[numeric_cols].corr()
fig, ax = plt.subplots(figsize=(10, 8))
sns.heatmap(corr_matrix,
annot=True,
fmt='.2f',
cmap='RdYlGn',
center=0,
ax=ax,
linewidths=0.5)
ax.set_title('Correlation Matrix')
plt.tight_layout()
plt.show()
# Pivot heatmap (e.g., revenue by month and category)
pivot = df.pivot_table(values='revenue', index='category', columns='month_name', aggfunc='sum')
sns.heatmap(pivot, annot=True, fmt=',.0f', cmap='YlOrRd')
plt.title('Revenue by Category and Month')
plt.show()
Chart Formatting Best Practices
# Professional chart template
fig, ax = plt.subplots(figsize=(12, 6))
# Plot data
ax.plot(x, y, color='#059669', linewidth=2, marker='o', markersize=4)
# Title and labels
ax.set_title('Clear, Descriptive Title', fontsize=14, fontweight='bold', pad=15)
ax.set_xlabel('X Axis Label (unit)', fontsize=11)
ax.set_ylabel('Y Axis Label (unit)', fontsize=11)
# Clean the spines (borders)
ax.spines['top'].set_visible(False)
ax.spines['right'].set_visible(False)
# Format y-axis as currency
import matplotlib.ticker as mticker
ax.yaxis.set_major_formatter(mticker.FuncFormatter(lambda x, p: f'${x:,.0f}'))
# Annotate key points
ax.annotate('Key Event', xy=(peak_x, peak_y),
xytext=(peak_x + 1, peak_y * 1.1),
arrowprops=dict(arrowstyle='->', color='red'),
color='red')
plt.tight_layout()
plt.savefig('chart.png', dpi=150, bbox_inches='tight')
plt.show()
Key Takeaways
- Use seaborn for clean static analytical charts and Plotly for interactive charts suitable for dashboards and sharing.
- Match chart type to data type: line for trends, bar for comparisons, histogram/box plot for distributions, scatter for relationships, heatmap for correlation.
- Format charts professionally: descriptive titles, clear axis labels, remove unnecessary borders (spines), use consistent colour.
- plt.savefig('filename.png', dpi=150, bbox_inches='tight') saves publication-quality images for reports and presentations.
- Plotly's px.scatter(), px.line(), px.bar() provide interactive charts with hover details in just a few lines of code.
Practice Exercise
Using a public dataset, create a four-panel visualisation figure:
import pandas as pd, matplotlib.pyplot as plt, seaborn as sns
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
fig, axes = plt.subplots(2, 2, figsize=(14, 10))
fig.suptitle('Titanic Dataset Analysis', fontsize=16, fontweight='bold')
# 1. Distribution of passenger ages (histogram)
sns.histplot(df['Age'].dropna(), bins=30, ax=axes[0, 0], color='steelblue')
axes[0, 0].set_title('Age Distribution')
# 2. Survival rate by passenger class (bar chart)
survival_by_class = df.groupby('Pclass')['Survived'].mean() * 100
survival_by_class.plot(kind='bar', ax=axes[0, 1], color=['#EF4444','#F59E0B','#059669'])
axes[0, 1].set_title('Survival Rate by Class')
axes[0, 1].set_ylabel('Survival Rate (%)')
# 3. Fare distribution by class (box plot)
sns.boxplot(data=df, x='Pclass', y='Fare', ax=axes[1, 0])
axes[1, 0].set_title('Fare by Passenger Class')
# 4. Correlation heatmap
sns.heatmap(df[['Age','Fare','Survived','Pclass']].corr(),
annot=True, fmt='.2f', ax=axes[1, 1], cmap='RdYlGn', center=0)
axes[1, 1].set_title('Correlation Matrix')
plt.tight_layout()
plt.savefig('titanic_eda.png', dpi=150)
plt.show()
Try it yourself
Key Takeaways
- Use seaborn for static publication-quality charts and Plotly for interactive charts -- both are simpler than raw matplotlib for most analytical charts.
- Match chart type to data type: line for trends, bar for comparisons, histogram/box plot for distributions, scatter for relationships, heatmap for correlation.
- Format charts professionally: descriptive titles, clear axis labels, remove top and right spines, use consistent colours.
- plt.savefig('file.png', dpi=150, bbox_inches='tight') saves high-resolution images for reports; always call it before plt.show().
- Plotly express (px.line, px.bar, px.scatter) creates interactive charts in 2-3 lines of code -- use it for shared and embedded outputs.
Quick Quiz
1.Which Python library is best for creating interactive charts with hover details?
2.What chart type is most appropriate for showing the distribution of order values across different customer segments?
3.What does plt.savefig('chart.png', dpi=150, bbox_inches='tight') do?
4.You want to show the relationship between advertising spend and revenue, coloured by marketing channel. Which chart type and library combination is best?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx