Correlation and Regression
Understanding Relationships Between Variables
One of the most common analytical questions is: "Is there a relationship between these two variables?" Correlation and regression provide the tools to measure and quantify those relationships.
Correlation
Correlation measures the strength and direction of the linear relationship between two continuous variables.
Pearson Correlation Coefficient (r)
The most common measure of correlation. Ranges from -1 to +1:
- r = +1: Perfect positive linear relationship
- r = 0.7 to 0.9: Strong positive relationship
- r = 0.4 to 0.6: Moderate positive relationship
- r = 0 to 0.3: Weak or no relationship
- r = -0.7 to -0.9: Strong negative relationship
- r = -1: Perfect negative linear relationship
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
from scipy.stats import pearsonr, spearmanr
# Calculate correlation between two variables
r, p_value = pearsonr(df['marketing_spend'], df['revenue'])
print(f"Pearson r: {r:.3f}")
print(f"P-value: {p_value:.4f}")
print(f"Significant: {p_value < 0.05}")
# Correlation matrix for multiple variables
corr_matrix = df[['revenue', 'marketing_spend', 'customer_count', 'avg_order_value']].corr()
print(corr_matrix)
# Visualise as heatmap
sns.heatmap(corr_matrix, annot=True, fmt='.2f', cmap='RdYlGn', center=0)
plt.title('Correlation Matrix')
plt.show()
Spearman Correlation
Non-parametric alternative to Pearson. Use when:
- Data is not normally distributed
- Relationships are monotonic but not necessarily linear
- Data contains significant outliers
r_spearman, p_value = spearmanr(df['marketing_spend'], df['revenue'])
Correlation vs Causation
This is the most important concept in data analysis: correlation does not imply causation.
Just because two variables move together does not mean one causes the other. Alternative explanations:
- Reverse causation: Higher revenue might cause higher marketing spend (not the other way around).
- Confounding variable: A third variable causes both. Ice cream sales and drowning rates are correlated because both increase in summer (confounded by temperature).
- Spurious correlation: Random chance. With enough variables, some will correlate by coincidence.
Always think critically about mechanism: WHY would A cause B? Is the relationship plausible? Can you design an experiment to test it?
Simple Linear Regression
Linear regression models the linear relationship between a dependent variable (y) and one or more independent variables (x). It provides:
- The equation of the best-fit line
- The coefficient (how much y changes for each unit change in x)
- R-squared (how much of y's variation is explained by x)
Simple Linear Regression (one predictor)
from scipy import stats
# Linear regression: marketing spend vs revenue
slope, intercept, r_value, p_value, std_err = stats.linregress(
df['marketing_spend'],
df['revenue']
)
print(f"Slope: {slope:.3f}") # revenue change per $1 of marketing spend
print(f"Intercept: {intercept:.2f}") # baseline revenue when spend = 0
print(f"R-squared: {r_value**2:.3f}")# proportion of variance explained
print(f"P-value: {p_value:.4f}")
# Prediction
predicted_revenue = slope * 50000 + intercept
print(f"Predicted revenue for $50K spend: ${predicted_revenue:,.0f}")
# Visualise with regression line
plt.figure(figsize=(10, 6))
plt.scatter(df['marketing_spend'], df['revenue'], alpha=0.4, color='#059669')
x_line = np.array([df['marketing_spend'].min(), df['marketing_spend'].max()])
plt.plot(x_line, slope * x_line + intercept, 'r-', linewidth=2, label=f'y = {slope:.2f}x + {intercept:.0f}')
plt.xlabel('Marketing Spend ($)')
plt.ylabel('Revenue ($)')
plt.title('Revenue vs Marketing Spend')
plt.legend()
plt.show()
Interpreting Regression Output
Slope: For every $1 increase in marketing spend, revenue increases by [slope] dollars on average.
Intercept: The expected revenue when marketing spend is zero (may be meaningless outside the range of your data).
R-squared (R2): Ranges from 0 to 1. An R2 of 0.72 means 72% of the variation in revenue is explained by marketing spend. The remaining 28% is explained by other factors.
Multiple Linear Regression
Multiple regression includes multiple predictor variables:
from sklearn.linear_model import LinearRegression
from sklearn.preprocessing import StandardScaler
import pandas as pd
# Predict revenue from multiple factors
X = df[['marketing_spend', 'customer_count', 'avg_order_value', 'season_index']]
y = df['revenue']
# Fit the model
model = LinearRegression()
model.fit(X, y)
print(f"R-squared: {model.score(X, y):.3f}")
print("Coefficients:")
for feature, coef in zip(X.columns, model.coef_):
print(f" {feature}: {coef:.3f}")
print(f"Intercept: {model.intercept_:.2f}")
# Make predictions
new_data = pd.DataFrame({
'marketing_spend': [50000],
'customer_count': [1200],
'avg_order_value': [95],
'season_index': [1.2]
})
predicted = model.predict(new_data)[0]
print(f"Predicted revenue: ${predicted:,.0f}")
Practical Applications
Forecasting
Use regression to forecast future values based on known relationships:
# Given next quarter's planned marketing spend, forecast revenue
next_quarter_spend = 150000
forecast = slope * next_quarter_spend + intercept
confidence_interval = 1.96 * std_err * np.sqrt(1/len(df) + ...)
Attribution Analysis
Use regression to estimate how much each marketing channel contributes to revenue:
channels = ['social_spend', 'search_spend', 'email_spend', 'tv_spend']
X = df[channels]
model = LinearRegression().fit(X, df['revenue'])
contribution = pd.Series(model.coef_, index=channels).sort_values(ascending=False)
Customer Lifetime Value Prediction
Predict a customer's future lifetime value based on early purchase behaviour.
Key Takeaways
- Pearson correlation (r) measures linear relationship strength from -1 to +1; use Spearman for non-normal data or outliers.
- Correlation does not imply causation -- always consider reverse causation, confounding variables, and spurious correlations.
- Simple linear regression models y as a function of one predictor; multiple regression uses several predictors.
- R-squared measures how much of the dependent variable's variation is explained by the model (0-1, higher is better).
- The regression coefficient (slope) tells you how much y is expected to change for each unit increase in x, holding other variables constant.
Practice Exercise
import pandas as pd, numpy as np, matplotlib.pyplot as plt
from scipy import stats
# Generate sample data: marketing spend and revenue
np.random.seed(42)
n = 100
marketing_spend = np.random.uniform(10000, 200000, n)
revenue = 3.5 * marketing_spend + np.random.normal(0, 100000, n) + 500000
df = pd.DataFrame({'marketing_spend': marketing_spend, 'revenue': revenue})
# 1. Calculate and interpret Pearson correlation
r, p = stats.pearsonr(df['marketing_spend'], df['revenue'])
print(f"Correlation: r={r:.3f}, p={p:.4f}")
# 2. Run simple linear regression
slope, intercept, r2, p_val, se = stats.linregress(df['marketing_spend'], df['revenue'])
print(f"Slope: {slope:.2f} (revenue per $1 spend)")
print(f"R-squared: {r2**2:.3f}")
# 3. Plot scatter plot with regression line
# 4. Interpret: What does the slope tell you about ROI?
# 5. Predict revenue for a $100,000 marketing budget
prediction = slope * 100000 + intercept
print(f"Predicted revenue: ${prediction:,.0f}")
Try it yourself
Key Takeaways
- Pearson r measures linear relationship strength (-1 to +1); use Spearman r for non-normal data or when the relationship is monotonic but non-linear.
- Correlation does not imply causation -- always consider reverse causation, confounding variables, and coincidental patterns.
- Simple linear regression gives a slope (change in y per unit of x), intercept, and R-squared (proportion of variance explained).
- R-squared measures model fit: an R2 of 0.70 means 70% of the variation in y is explained by the predictor(s).
- Multiple regression includes several predictors; each coefficient represents the effect of that variable while holding others constant.
Quick Quiz
1.What does a Pearson correlation coefficient of 0.85 indicate?
2.What is R-squared (R2) in regression analysis?
3.You find that cities with more ice cream shops have higher rates of crime. What is the most likely explanation?
4.The regression equation is Revenue = 3.5 * Marketing_Spend + 200,000. If marketing spend is $100,000, what is the predicted revenue?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx