Building Your Data Science Portfolio
Your Portfolio Is Your Proof
For entry-level data science roles, recruiters at Kuda, Moniepoint, Flutterwave, Andela and global remote-first companies all face the same problem: hundreds of applicants with similar certificates. A portfolio of real projects is how you prove you can do the work. For many hiring managers it matters more than where you studied.
A strong portfolio needs 3 to 5 quality projects, not 20 tutorial copies.
What Makes a Portfolio Project Stand Out
| Weak project | Strong project |
|---|---|
| Titanic survival prediction | Churn prediction for a telecom with a cost-benefit analysis |
| Iris classification | Naira exchange rate volatility analysis from CBN data |
| Copied Kaggle notebook | Original question, own data collection and cleaning |
| Code with no explanation | Clear README, markdown narrative and business recommendations |
| "Accuracy: 97%" | "Targeting the top 1,000 finds 2.6x more churners than random" |
The formula: a real question, messy real data, sound methods, and a clear business conclusion.
A balanced set of projects
- An end-to-end EDA project using local data, such as the ShopNaija case study or an NBS inflation analysis
- A supervised ML project with honest evaluation, such as churn or loan default prediction
- A data collection project using an API or scraping, such as a World Bank or exchange rate tracker
- A deployed or interactive project: a Streamlit app or dashboard that anyone can click
- (Optional) A domain-specific project for the industry you want to join: fintech, health, agriculture or energy
GitHub: Your Professional Home
Every project should live in its own GitHub repository with a structure like this:
nigeria-telecom-churn/
├── README.md <- the most important file
├── data/
│ └── README.md <- where the data came from (do not commit huge or private files)
├── notebooks/
│ ├── 01_eda.ipynb
│ └── 02_modelling.ipynb
├── src/
│ └── features.py <- reusable functions
├── reports/
│ └── figures/
└── requirements.txt
The README template recruiters love
# Predicting Telecom Churn in Nigeria
## Problem
A retention team can target 1,000 subscribers a week. Who should they target?
## Key Results
- Gradient boosting model: ROC AUC 0.81
- Top-1,000 list contains 2.6x more churners than random targeting
- Top drivers: falling recharges, second SIM, complaints
## Approach
Data generation/collection -> cleaning -> EDA -> feature engineering -> 3 models -> business evaluation
## Recommendations
1. Prioritise subscribers with falling recharges for bundle offers
2. Route high-risk subscribers who have complained to customer care
## Tech
Python, pandas, scikit-learn, seaborn
## How to run
pip install -r requirements.txt, then open notebooks/01_eda.ipynb
Put the results first. Many recruiters read only the top of the README.
Also: pin your best repositories on your GitHub profile, write a profile README introducing yourself, and commit regularly. A consistent activity history shows sustained practice.
Kaggle and Zindi
- Kaggle: publish well-documented notebooks, enter competitions, and earn medals. Notebook upvotes show your ability to communicate.
- Zindi: Africa's data science competition platform, with challenges from African companies and organisations. A top finish on an African problem is a strong signal for local employers.
You do not need to win. A clean, well-explained notebook that ranks mid-table is still great portfolio material.
Make One Project Interactive
A live app makes your work tangible for non-technical interviewers. Streamlit turns a Python script into a web app:
# app.py
import streamlit as st
import joblib
import pandas as pd
model = joblib.load("churn_model.joblib")
st.title("Telecom Churn Risk Checker")
tenure = st.slider("Tenure (months)", 1, 72, 12)
complaints = st.number_input("Complaints in the last 90 days", 0, 10, 0)
# ... more inputs ...
if st.button("Predict"):
risk = model.predict_proba(pd.DataFrame([{"tenure_months": tenure, "complaints_90d": complaints}]))[0, 1]
st.metric("Churn risk", f"{risk:.0%}")
Deploy it for free on Streamlit Community Cloud or Hugging Face Spaces, and put the link at the top of your README.
Presenting Findings to Stakeholders
Your portfolio should show that you can communicate, not only code. Use this structure for every write-up and presentation:
- Headline: the one-sentence answer ("Pay on Delivery drives 60% of cancellations")
- Why it matters: the business impact in naira or customers
- Evidence: two or three clear charts, no more
- Recommendation: what to do next, and how to test it
- Caveats: limitations of the data and methods
Avoid jargon with non-technical audiences. Say "the model finds 2.6 times as many at-risk customers as random targeting" rather than "the model has an AUC of 0.81". Record a 3-minute Loom video walking through one project; it lets recruiters see your communication skills before an interview.
Portfolio Checklist
- 3 to 5 projects, at least two using real or local data
- Each repository has a results-first README
- At least one deployed, interactive project
- A Kaggle or Zindi profile with at least one public notebook
- LinkedIn posts summarising each project with one chart
- Your CV links directly to your GitHub and live app
Try it yourself
Key Takeaways
- A portfolio of 3 to 5 original, well-documented projects is your strongest proof of skill for entry-level data science roles.
- Strong projects combine a real question, messy real or local data, sound methods and a clear business conclusion.
- Give each project its own GitHub repository with a results-first README, a clean structure and instructions to run it.
- Kaggle and Zindi notebooks, plus at least one deployed Streamlit app, make your skills visible and interactive.
- Present findings as headline, impact, evidence, recommendation and caveats, and translate metrics into business language.
Quick Quiz
1.Which project is most likely to impress a hiring manager at a Nigerian fintech?
2.What should appear at the very top of a project README?
3.When presenting to a non-technical manager, which phrasing is best?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx