Reading Data and A/B Testing
Data Without Interpretation Is Just Numbers
A dashboard full of charts means nothing until a PM asks the right questions of it: Is this change real, or noise? Is it caused by what we shipped, or something else entirely? This lesson covers how to read data critically and how to run A/B tests — the most reliable way product teams answer "did this actually work?"
Correlation Is Not Causation
Imagine Flutterwave notices that transaction volume rose 20% the same week a new dashboard redesign shipped. It is tempting to credit the redesign. But that same week might also have been salary week, a public holiday shopping period, or a marketing campaign running independently. Without a controlled comparison, a PM cannot separate the effect of the redesign from everything else happening at the same time.
This is exactly the problem A/B testing solves.
What Is an A/B Test?
An A/B test (also called a split test or randomized controlled experiment) randomly divides users into two or more groups:
- Control (A): sees the current experience
- Variant (B): sees the new experience
Because users are assigned randomly, the two groups should be statistically similar in every way except the one thing being tested. Any meaningful difference in outcomes can then be attributed to the change itself, not to external factors like salary week.
A Simple Example
Konga wants to know if adding a "Only 3 left in stock" urgency label increases purchase conversion.
- Group A (50% of visitors): sees the product page as normal
- Group B (50% of visitors): sees the same page with the urgency label
After two weeks: Group A converts at 4.1%, Group B converts at 4.6%. Is that difference real, or could it have happened by chance?
Statistical Significance, in Plain English
Statistical significance tells you how confident you can be that an observed difference is real and not due to random chance.
Key ingredients:
| Concept | What it means |
|---|---|
| Sample size | How many users were in each group. Bigger samples give more reliable results. |
| Baseline conversion rate | The starting rate you're trying to improve (4.1% above). |
| Minimum detectable effect | The smallest improvement worth detecting (e.g. is a 0.5pp lift even meaningful to the business?). |
| p-value | The probability the observed difference could occur by pure chance if there were actually no real difference. A p-value below 0.05 is the common (though somewhat arbitrary) threshold for calling a result "statistically significant." |
| Confidence interval | A range around your result showing how much uncertainty remains, e.g. "we are 95% confident the true lift is between 0.1pp and 0.9pp." |
The most common mistake new PMs make: peeking at results after two days and declaring a "winner" before enough users have gone through the test. Small samples produce noisy, unreliable percentages that often reverse a week later. Always calculate the required sample size before launching a test, and let it run to completion.
Designing a Good Experiment
- Write a clear hypothesis. "We believe adding urgency labels will increase conversion by at least 0.3 percentage points, because scarcity signals reduce hesitation at checkout."
- Pick one primary metric. Testing ten metrics at once increases the odds one looks significant purely by chance (a problem called p-hacking).
- Randomize properly. Users should be assigned to a group in a way that avoids bias, and each user should consistently see the same variant throughout the test.
- Calculate sample size up front. Tools like Optimizely, Statsig, or even a simple online A/B test calculator can tell you how many users you need for a given baseline rate and expected lift.
- Run for full business cycles. A test that runs Monday to Wednesday misses weekend behavior. Most teams run tests for at least one to two full weeks.
- Watch for interaction effects. If Access Bank runs a pricing test and an onboarding test on the same users at the same time, the two experiments might interfere with each other's results.
Beyond Simple A/B Tests
- Multivariate tests test several changes at once (e.g. headline + button color + image) to see which combination performs best, but need much larger sample sizes.
- Phased or gradual rollouts release a new feature to 5%, then 25%, then 100% of users, catching major problems early without a full controlled experiment.
- Holdout groups keep a small percentage of users on the old experience permanently, so a team can measure the long-run cumulative impact of many small changes over months, not just days.
The PM's Job in All of This
A PM does not need to personally compute p-values by hand, but must understand enough statistics to ask good questions of a data scientist or analyst: "Was this sample size big enough?" "Could this result be a fluke?" "What did we NOT measure that might explain it?" That healthy skepticism, paired with a bias toward testing rather than debating, is what separates evidence-driven product teams from ones running on opinions.
Try it yourself
Key Takeaways
- Correlation is not causation: two things happening at the same time does not prove one caused the other, which is why controlled A/B tests exist.
- An A/B test randomly splits users into control and variant groups so any performance difference can be attributed to the change itself.
- Statistical significance (commonly p < 0.05) measures confidence that a result is real, not proof of how large or valuable the impact is.
- Sample size should be calculated before launching a test; peeking early on small samples produces unreliable, often reversible results.
- Beyond simple A/B tests, teams also use multivariate tests, phased rollouts, and holdout groups depending on the question being answered.
Quick Quiz
1.Flutterwave sees transaction volume rise 20% the same week a redesign ships. Why can't the PM immediately conclude the redesign caused the rise?
2.What does a p-value below 0.05 commonly indicate in an A/B test?
3.Why is it risky to declare a 'winner' after only two days of an A/B test?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx