Bias Auditing Frameworks for Machine Learning Pipelines

Authors

  • Jonas Horvath Postdoctoral Researcher, Department of Artificial Intelligence, European Institute of AI, Berlin, Germany Author
  • Marco Nowak Associate Professor, School of Data Science, Central European Tech University, Vienna, Austria Author
  • Daniel Popescu Research Scientist, Department of Computer Science, European Institute of AI, Berlin, Germany Author

DOI:

https://doi.org/10.5281/

Keywords:

bias auditing, ML pipeline, algorithmic fairness, responsible AI, bias detection, EU AI Act, data governance, fairness metrics

Abstract

Bias in machine learning pipelines -- arising at data collection, preprocessing, feature engineering, model training, and deployment stages -- represents one of the most consequential and pervasive sources of ethical failure in deployed AI systems. Bias auditing, the systematic inspection of ML pipeline stages and outputs for discriminatory patterns, is mandated by the EU AI Act (2024) Article 10 for high-risk AI systems yet remains poorly standardised in practice. This paper proposes the Pipeline Bias Audit Framework (PBAF), a structured, stage-by-stage auditing methodology covering six ML pipeline stages: data collection, data preprocessing, feature engineering, model training, model evaluation, and deployment monitoring. The PBAF specifies nineteen bias indicators, each with a formal detection method, a severity classification, and a remediation strategy, organised into a replicable audit protocol. The PBAF is evaluated through application to 24 real ML pipelines drawn from six industry sectors, with audit results compared against a gold-standard expert panel assessment. PBAF audits achieve a mean bias indicator detection rate of 87.4% (SD = 5.8%) relative to expert panel ground truth, a false positive rate of 6.2% (SD = 1.9%), and a mean audit completion time of 3.4 hours per pipeline (SD = 0.8). Post-audit remediation following PBAF recommendations reduced the mean bias indicator count from 4.7 to 1.6 per pipeline (65.9% reduction) over a 90-day follow-up period. The study contributes the PBAF specification, a bias indicator catalogue, and an empirical benchmark for ML pipeline bias auditing practice.

Author Biographies

  • Jonas Horvath, Postdoctoral Researcher, Department of Artificial Intelligence, European Institute of AI, Berlin, Germany

    Postdoctoral Researcher, Department of Artificial Intelligence, European Institute of AI, Berlin, Germany

  • Marco Nowak, Associate Professor, School of Data Science, Central European Tech University, Vienna, Austria

    Associate Professor, School of Data Science, Central European Tech University, Vienna, Austria

  • Daniel Popescu, Research Scientist, Department of Computer Science, European Institute of AI, Berlin, Germany

    Research Scientist, Department of Computer Science, European Institute of AI, Berlin, Germany

Downloads

Published

2025-06-20

How to Cite

Bias Auditing Frameworks for Machine Learning Pipelines. (2025). AI Governance and Society Journal P-ISSN 3117-6097 and E-ISSN 3117-6100, 2(2), 1-8. https://doi.org/10.5281/