Bias and Discrimination Risks in Data-Driven Emerging Technologies
Keywords:
algorithmic bias, discrimination risk, fairness AI, representation bias, feedback loop, predictive policing, facial recognition, LLM biasAbstract
Bias and discrimination risks in data-driven emerging technologies arise from multiple sources -- historical data encoding past discrimination, measurement instruments that are less accurate for marginalised groups, model architectures that amplify statistical patterns of inequality, and deployment contexts that direct high-risk AI to high-disparity populations. This paper proposes the Bias and Discrimination Risk Assessment Framework (BDRAF), a systematic taxonomy and measurement approach covering six bias sources -- historical bias, representation bias, measurement bias, aggregation bias, deployment bias, and feedback loop bias -- applied to five emerging technology domains: large language models, facial recognition, predictive policing, credit scoring AI, and medical diagnosis AI. BDRAF introduces the Bias Risk Score (BRS) integrating bias source prevalence, demographic disparity magnitude, harm severity, and mitigation feasibility. Key results: predictive policing achieves the highest BRS (0.924) driven by historical and feedback loop bias; LLMs show the broadest bias source distribution (all six sources active); representation bias is the most prevalent source (present in all five domains); feedback loop bias is the most ethically severe single source (self-reinforcing discrimination). The framework provides bias risk prioritisation guidance and source-matched mitigation strategies.
