Explainability-Driven Debugging of Ethical Failures in AI Systems

Authors

DOI:

https://doi.org/10.5281/

Keywords:

explainability-driven debugging, ethical AI failures, responsible AI, root cause analysis, SHAP, fairness debugging, AI transparency, EU AI Act

Abstract

Ethical failures in deployed AI systems -- encompassing discriminatory outputs, unjustified adverse decisions, privacy violations, and unsafe behaviour -- are frequently detected through post-deployment monitoring, user complaints, or external audits, yet the root causes of these failures are rarely systematically diagnosed and remediated. Explainability methods, originally developed to provide interpretable accounts of individual predictions, offer an underexplored capability as diagnostic tools for identifying the architectural, data, and training-process causes of ethical failures. This paper proposes the Explainability-Driven Ethical Debugging (XDED) methodology, a structured six-step process for diagnosing and remediating ethical failures in AI systems using explainability techniques as primary diagnostic instruments. The XDED methodology is evaluated through a case study analysis of 22 documented ethical AI failures drawn from healthcare, financial services, criminal justice, and recruitment domains, and through a controlled experiment involving 14 AI development teams applying XDED versus conventional debugging approaches to seeded ethical failures in three classification models. Controlled experiment results demonstrate that XDED-applying teams achieve a 52.4% higher rate of correct root-cause identification (XDED: 79.6%, SD = 7.3%; conventional: 52.2%, SD = 9.1%; p < 0.001) and a 38.7% reduction in time-to-remediation (XDED mean: 4.2 hours, SD = 0.9; conventional: 6.9 hours, SD = 1.4; p < 0.001). The study contributes the XDED methodology specification, a taxonomy of ethical failure root causes diagnosable through explainability, and empirical evidence for explainability as a practical ethical debugging tool.

Author Biography

  • Hugo Dubois, Assistant Professor, School of Data Science, Advanced Computing University, Paris, France

    Assistant Professor, School of Data Science, Advanced Computing University, Paris, France

Downloads

Published

2025-03-25

How to Cite

Explainability-Driven Debugging of Ethical Failures in AI Systems. (2025). AI Governance and Society Journal P-ISSN 3117-6097 and E-ISSN 3117-6100, 2(1), 19-26. https://doi.org/10.5281/