Project Planning for Data Scientists
Why Most Data Science Projects Fail
Industry surveys have repeatedly found that a large share of data science and ML projects never make it into production. The usual reasons are rarely technical. They are:
- The problem was vague, so nobody could agree whether the result was useful
- The data needed did not exist, was not accessible, or was far messier than expected
- The model answered a question nobody was going to act on
- The project took so long that the business moved on
Good planning prevents all four. Before you open a notebook, spend time scoping the project. This lesson gives you a framework you can use in your portfolio projects and at work, whether at a Lagos fintech or a global company.
Step 1: Start With the Decision, Not the Data
Ask the stakeholder: "What decision will you make differently with this result?"
| Vague request | Decision-focused framing |
|---|---|
| "Analyse our customer data" | "Which 5,000 customers should get our retention offer this month?" |
| "Build an AI model for loans" | "Should we approve or decline applications under N200,000 automatically?" |
| "Look at our sales" | "Which three states should we prioritise for new agents in Q3?" |
If nobody can name a decision, the project is not ready.
Step 2: Define Success in Numbers
Agree on two kinds of metrics before you start:
- Business metric: what the organisation cares about, such as reducing churn from 4.2% to 3.5% a month, or cutting fraud losses by N50 million a quarter
- Model metric: how you will judge the model, such as recall of at least 70% at a precision of at least 40%
Also define the baseline: what happens today without your model? A model is only valuable if it beats the current process, which might be a simple rule or a manager's judgement.
Step 3: Audit the Data
For every data source, answer:
| Question | Why it matters |
|---|---|
| Does it exist? | Many projects assume data that was never recorded |
| Can I access it, and who approves? | Access requests at banks can take weeks |
| How far back does it go? | Seasonality needs at least a year or two |
| Is the target labelled? | Supervised learning needs historical outcomes |
| Is it personal data? | The Nigeria Data Protection Act 2023 and GDPR apply |
| Will it be available at prediction time? | Otherwise you have data leakage |
Data leakage is one of the most common mistakes: using information that would not be known at the moment of prediction. Using "number of calls to the cancellation line" to predict churn looks brilliant in testing and is useless in production, because by then the customer has already decided to leave.
Step 4: Scope the Deliverables
Be explicit about what you will hand over:
- An exploratory analysis with findings (days to weeks)
- A dashboard that updates regularly
- A model plus documentation
- A deployed API or batch scoring job integrated with a product
- A presentation with recommendations
Start with the smallest version that proves value, a minimum viable analysis. A simple model with a CSV list of high-risk customers that the retention team can use next week beats a perfect deployed system six months from now.
Step 5: Plan the Timeline and Risks
A typical four-week portfolio or first-iteration project:
| Week | Focus |
|---|---|
| 1 | Scoping, data access, data audit, baseline |
| 2 | Cleaning and EDA; share early findings with the stakeholder |
| 3 | Feature engineering, modelling and evaluation |
| 4 | Final model, write-up, presentation and handover |
List the top risks (data arrives late, labels are unreliable, stakeholder availability) and a mitigation for each.
Step 6: Write a One-Page Project Brief
Put it all on one page and get the stakeholder to agree before you build:
PROJECT: Reduce prepaid churn in Lagos
DECISION: Which subscribers receive a retention bundle each week
SUCCESS: Monthly churn in the target group falls from 4.2% to 3.5%
MODEL METRIC: Recall >= 70% at precision >= 40% on a held-out month
BASELINE: Current rule - offer to anyone inactive for 14+ days
DATA: 12 months of usage, recharge and complaint logs (access: approved by the CRM lead)
DELIVERABLE (v1): Weekly CSV of the top 5,000 at-risk subscribers + a slide deck
TIMELINE: 4 weeks
RISKS: Complaint data only goes back 6 months -> test the model with and without it
This brief is also a great portfolio artefact. Recruiters love seeing that you think about the business, not only the algorithm.
Try it: Use the planner below to scope a project of your own. It flags common gaps before you start.
Try it yourself
Key Takeaways
- Most data science projects fail because of vague problems, missing data or unused outputs, not because of algorithms.
- Start with the decision the stakeholder will make differently, then define business and model success metrics with numbers.
- Always measure a baseline, meaning the current process or a simple rule, that your model must beat.
- Audit data for existence, access, history, labels, privacy and availability at prediction time to avoid data leakage.
- Scope a minimum viable deliverable, plan a short timeline with named risks, and agree a one-page brief before building.
Quick Quiz
1.A stakeholder asks you to 'analyse our customer data'. What is the best first question to ask?
2.Which of these is an example of data leakage in a churn model?
3.Why start with a 'minimum viable analysis' instead of a fully deployed system?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx