AI Bias and Fairness
What is AI Bias?
AI systems learn patterns from data. When that data reflects historical inequalities, social biases, or unrepresentative samples, the AI system learns and often amplifies those biases. This is called AI bias: systematic errors in AI outputs that disadvantage certain groups of people.
AI bias is not a fringe concern -- it has caused real harm in high-stakes domains including hiring, lending, healthcare diagnosis, and criminal justice. As an AI Automation Engineer, understanding bias and building systems that mitigate it is a professional and ethical responsibility.
Types of AI Bias
Historical Bias
The training data reflects historical discrimination. A hiring model trained on historical hiring decisions will learn that certain demographic groups were less often hired -- and may replicate that pattern, even if the historical discrimination was wrong.
Example: A resume screening AI trained on historical hiring data from a male-dominated industry may learn to deprioritise resumes from women, because historically fewer women were hired.
Representation Bias
Certain groups are underrepresented in the training data. The model performs poorly for those groups because it has seen fewer examples of them.
Example: A facial recognition system trained primarily on lighter-skinned faces performs significantly worse on darker-skinned faces -- a bias documented in multiple commercial systems.
Measurement Bias
The data used to train or evaluate the model does not measure what we actually care about. Proxy variables (things we can measure) imperfectly represent the target concept (what we actually want to predict).
Example: A loan default prediction model uses zip code as a feature. Zip code correlates with race due to historical housing discrimination, so the model effectively discriminates by race while appearing to only use geographic data.
Aggregation Bias
The model treats different groups as homogeneous when meaningful differences exist between them.
Example: A medical AI trained primarily on data from men may give inaccurate recommendations for women, whose symptoms for certain conditions present differently.
Feedback Loop Bias
AI systems can create self-reinforcing cycles. A predictive policing algorithm predicts higher crime in certain areas, leading to more policing in those areas, generating more arrests, which reinforces the algorithm's prediction.
Fairness: Multiple Definitions, Real Tensions
There is no single universally agreed definition of fairness in AI. Different definitions can directly conflict with each other:
Demographic parity: The model's positive outcome rate should be the same across demographic groups. (Equal outcomes)
Equal opportunity: The model's true positive rate should be the same across groups. (Equal accuracy for those who qualify)
Predictive parity: When the model predicts a positive outcome, the probability should be the same regardless of group membership. (Equal precision)
Individual fairness: Similar individuals should receive similar predictions.
Example of the tension: In criminal justice risk assessment, demographic parity (equal release rates across groups) and predictive parity (equal accuracy in predicting reoffending) are mathematically impossible to achieve simultaneously when base rates differ between groups.
This means that "fair" AI requires value judgements about which definition of fairness matters most for a given context -- and that decision should involve the communities affected, not just the engineers.
Detecting Bias in AI Systems
Audit your data first
- What are the demographic characteristics of your training data?
- Are certain groups underrepresented?
- Does the labelling process reflect the biases of the labellers?
Evaluate disaggregated metrics
Do not just measure overall accuracy. Measure performance separately for each demographic group:
function evaluateFairness(predictions, groundTruth, groups) {
const results = {};
for (const group of [...new Set(groups)]) {
const groupIndices = groups.map((g, i) => g === group ? i : -1).filter(i => i >= 0);
const groupPredictions = groupIndices.map(i => predictions[i]);
const groupTruth = groupIndices.map(i => groundTruth[i]);
const tp = groupPredictions.filter((p, i) => p === 1 && groupTruth[i] === 1).length;
const tn = groupPredictions.filter((p, i) => p === 0 && groupTruth[i] === 0).length;
const fp = groupPredictions.filter((p, i) => p === 1 && groupTruth[i] === 0).length;
const fn = groupPredictions.filter((p, i) => p === 0 && groupTruth[i] === 1).length;
results[group] = {
accuracy: (tp + tn) / groupIndices.length,
precision: tp / (tp + fp) || 0,
recall: tp / (tp + fn) || 0,
falsePositiveRate: fp / (fp + tn) || 0,
};
}
return results;
}
Mitigating Bias: What Engineers Can Do
Before training
- Audit training data for representation gaps
- Collect more data for underrepresented groups
- Review and audit labels for systematic bias
- Remove or transform features that are proxies for protected characteristics
During development
- Test your system on disaggregated subgroups from the beginning, not just overall
- Use fairness-aware algorithms and constraints
- Have diverse perspectives in your evaluation process
After deployment
- Monitor outcomes across demographic groups over time
- Create mechanisms for affected communities to report concerns
- Conduct regular fairness audits
- Be willing to pause or retract a system if bias is discovered post-deployment
The limits of technical solutions
Bias in AI ultimately reflects bias in society and in the choices made during development. Technical debiasing can reduce but rarely eliminate bias, and some fairness definitions cannot be achieved simultaneously. Recognise that your technical choices embed value judgements -- make those judgements explicitly and transparently.
Bias in Language Models
LLMs trained on internet text absorb the biases present in that text:
- Associating certain professions with certain genders ("nurse" more often associated with women, "engineer" with men)
- Differential language generation for different demographic groups
- Perpetuating stereotypes present in training data
When building AI systems with LLMs, test for these associations. Use techniques like prompt diversity testing and output auditing for systematic patterns.
Key Takeaways
- AI bias is the systematic tendency of AI systems to produce outputs that disadvantage certain groups, often learned from biased training data.
- The five main bias types are historical, representation, measurement, aggregation, and feedback loop bias.
- Multiple definitions of fairness exist and can mathematically conflict -- choosing between them requires value judgements, not just technical decisions.
- Always evaluate model performance disaggregated by demographic group, not just as an aggregate.
- Technical debiasing can reduce but not eliminate bias -- building fair AI also requires diverse teams, community involvement, and ongoing monitoring.
Try it yourself
Key Takeaways
- AI bias is the systematic tendency of AI systems to disadvantage certain groups, typically learned from biased training data or flawed evaluation.
- The five main bias types are historical, representation, measurement, aggregation, and feedback loop -- each requiring different mitigation strategies.
- Multiple mathematically incompatible fairness definitions exist -- choosing between them requires value judgements, not just technical expertise.
- Always evaluate model performance disaggregated by demographic group; aggregate metrics mask differential performance.
- Technical debiasing reduces but cannot eliminate bias -- diverse teams, community involvement, and ongoing monitoring are equally essential.
Quick Quiz
1.What is representation bias in AI?
2.What is a feedback loop bias?
3.Why can multiple fairness definitions conflict with each other?
4.What is the most important technical step for detecting AI bias?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx