Bias Auditing Frameworks for Machine Learning Pipelines
DOI:
https://doi.org/10.5281/Keywords:
bias auditing, ML pipeline, algorithmic fairness, responsible AI, bias detection, EU AI Act, data governance, fairness metricsAbstract
Bias in machine learning pipelines -- arising at data collection, preprocessing, feature engineering, model training, and deployment stages -- represents one of the most consequential and pervasive sources of ethical failure in deployed AI systems. Bias auditing, the systematic inspection of ML pipeline stages and outputs for discriminatory patterns, is mandated by the EU AI Act (2024) Article 10 for high-risk AI systems yet remains poorly standardised in practice. This paper proposes the Pipeline Bias Audit Framework (PBAF), a structured, stage-by-stage auditing methodology covering six ML pipeline stages: data collection, data preprocessing, feature engineering, model training, model evaluation, and deployment monitoring. The PBAF specifies nineteen bias indicators, each with a formal detection method, a severity classification, and a remediation strategy, organised into a replicable audit protocol. The PBAF is evaluated through application to 24 real ML pipelines drawn from six industry sectors, with audit results compared against a gold-standard expert panel assessment. PBAF audits achieve a mean bias indicator detection rate of 87.4% (SD = 5.8%) relative to expert panel ground truth, a false positive rate of 6.2% (SD = 1.9%), and a mean audit completion time of 3.4 hours per pipeline (SD = 0.8). Post-audit remediation following PBAF recommendations reduced the mean bias indicator count from 4.7 to 1.6 per pipeline (65.9% reduction) over a 90-day follow-up period. The study contributes the PBAF specification, a bias indicator catalogue, and an empirical benchmark for ML pipeline bias auditing practice.

