Your First ML Model
The Goal: Predict Lagos Apartment Rents
You will build a linear regression model that predicts the annual rent of an apartment in Lagos from its size, number of bedrooms, area and distance to the nearest major road. Along the way you will learn the pattern that every scikit-learn model follows, which works the same way for a simple regression or a gradient boosting model at a big tech company.
Open in Google Colab: Lagos Rent Predictor
scikit-learn comes preinstalled in Colab. Open a new notebook and run each block below in order. By the end you will have trained, evaluated and used a real model.
Step 1: Create the Dataset
import numpy as np
import pandas as pd
rng = np.random.default_rng(10)
n = 800
area = rng.choice(["Lekki", "Yaba", "Ikeja", "Surulere", "Ajah"], n)
bedrooms = rng.integers(1, 5, n)
size_sqm = np.round(bedrooms * 35 + rng.normal(20, 12, n))
dist_road_km = np.round(rng.uniform(0.1, 5, n), 1)
area_premium = pd.Series(area).map({"Lekki": 2.2, "Ikeja": 1.6, "Yaba": 1.3, "Surulere": 1.1, "Ajah": 0.9}).values
rent = (size_sqm * 28000 + bedrooms * 400000) * area_premium - dist_road_km * 150000 + rng.normal(0, 600000, n)
homes = pd.DataFrame({
"area": area, "bedrooms": bedrooms, "size_sqm": size_sqm,
"dist_road_km": dist_road_km, "annual_rent": np.round(rent, -4),
})
homes.head()
Step 2: Separate Features (X) and Target (y)
X = homes[["bedrooms", "size_sqm", "dist_road_km", "area"]]
y = homes["annual_rent"]
Models need numbers, but area is text. One-hot encoding turns it into a 0/1 column per area:
X = pd.get_dummies(X, columns=["area"], drop_first=True)
X.head()
drop_first=True drops one area (Ajah) as the baseline, because its value can be inferred when all the other area columns are 0.
Step 3: Train/Test Split
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42
)
print(X_train.shape, X_test.shape) # (640, 7) (160, 7)
The model only sees the training set. The test set is locked away to simulate brand-new apartments, which gives you an honest estimate of real-world performance. random_state makes the split reproducible.
Step 4: Train (Fit) the Model
from sklearn.linear_model import LinearRegression
model = LinearRegression()
model.fit(X_train, y_train)
That is it: fit finds the coefficients that minimise the squared error between predictions and actual rents on the training data.
Every scikit-learn model follows the same three calls:
| Step | Code |
|---|---|
| Create | model = SomeModel() |
| Train | model.fit(X_train, y_train) |
| Predict | model.predict(X_new) |
Step 5: Inspect What It Learned
Linear regression is interpretable: each feature gets a coefficient.
coefs = pd.Series(model.coef_, index=X.columns).sort_values()
print(coefs.round(0))
print("Intercept:", round(model.intercept_))
Read a coefficient as: "holding everything else constant, a one-unit increase in this feature changes the predicted rent by this many naira." You should see a negative coefficient for dist_road_km (further from a main road means cheaper) and a large positive one for area_Lekki.
Step 6: Evaluate on the Test Set
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
pred = model.predict(X_test)
mae = mean_absolute_error(y_test, pred)
rmse = mean_squared_error(y_test, pred) ** 0.5
r2 = r2_score(y_test, pred)
print(f"MAE: N{mae:,.0f}")
print(f"RMSE: N{rmse:,.0f}")
print(f"R2: {r2:.3f}")
| Metric | Meaning |
|---|---|
| MAE | On average, predictions are off by this many naira. Easy to explain to stakeholders |
| RMSE | Like MAE, but penalises big errors more heavily |
| R2 | Share of the variation in rent that the model explains. 1.0 is perfect; 0 is no better than always predicting the average |
Always compare against a baseline. What if you just predicted the average rent for every apartment?
baseline = np.full(len(y_test), y_train.mean())
print(f"Baseline MAE: N{mean_absolute_error(y_test, baseline):,.0f}")
If your model does not clearly beat this, it is not adding value.
Step 7: Visualise the Errors
import matplotlib.pyplot as plt
plt.scatter(y_test, pred, alpha=0.5)
lims = [y_test.min(), y_test.max()]
plt.plot(lims, lims, "r--") # perfect-prediction line
plt.xlabel("Actual rent"); plt.ylabel("Predicted rent")
plt.title("Predicted vs actual")
plt.show()
Points close to the red line are good predictions. A pattern in the errors, such as a consistent underestimate for expensive homes, suggests the model is missing something.
Step 8: Predict for a New Apartment
new_flat = pd.DataFrame([{
"bedrooms": 3, "size_sqm": 125, "dist_road_km": 0.8,
"area_Ikeja": 0, "area_Lekki": 1, "area_Surulere": 0, "area_Yaba": 0,
}])
print(f"Predicted annual rent: N{model.predict(new_flat)[0]:,.0f}")
The new row must have exactly the same columns, in the same order, as the training data. That is a common source of bugs in production.
What You Just Did
You defined X and y, encoded a category, split the data, trained a model, read its coefficients, evaluated it against a baseline, and made a prediction. That is the full supervised learning loop. Every model from here on is a variation of it.
Try it yourself
Key Takeaways
- Every scikit-learn model follows the same pattern: create the model, fit it on the training data, and predict on new data.
- Separate features (X) from the target (y), and one-hot encode categorical columns with pd.get_dummies.
- train_test_split holds out unseen data so that evaluation reflects real-world performance.
- Evaluate regression with MAE, RMSE and R squared, and always compare against a baseline such as predicting the mean.
- Linear regression coefficients are interpretable, and new data must have exactly the same columns as the training data.
Quick Quiz
1.Why do we split data into training and test sets before fitting a model?
2.A linear regression coefficient for dist_road_km is -150,000. How do you interpret it?
3.Your model has an MAE of N1.9 million. Predicting the average rent for every apartment gives an MAE of N2.0 million. What should you conclude?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx