Benchmarking Frameworks for Trustworthy and Responsible AI Systems

Authors

  • Elena Klein Associate Professor, Department of Artificial Intelligence, Central European Tech University, Vienna, Austria Author

DOI:

https://doi.org/10.5281/

Keywords:

AI benchmarking, trustworthy AI, responsible AI evaluation, trustworthiness-accuracy trade-off, EU AI Act, NIST AI RMF, standardised evaluation, AI performance

Abstract

Benchmarking -- the systematic comparison of AI systems against standardised performance criteria using controlled evaluation datasets and protocols -- is a cornerstone of empirical progress in machine learning but has developed almost exclusively around technical accuracy metrics. The responsible AI community has produced a parallel proliferation of ethical evaluation instruments, yet a comprehensive benchmarking framework that integrates technical performance and trustworthiness criteria into a unified, standardised, and regulatorily-aligned evaluation protocol remains absent. This paper proposes the Trustworthy AI Benchmarking Framework (TABF), a structured benchmarking methodology integrating eight trustworthiness dimensions -- accuracy, fairness, transparency, explainability, robustness, privacy, safety, and accountability -- into a unified benchmark protocol with standardised evaluation datasets, scoring procedures, and compliance reporting aligned with the EU AI Act (2024) and NIST AI RMF (2023). TABF is validated through benchmark evaluation of 40 AI systems across five high-stakes domains and through an expert validation study with 38 responsible AI practitioners. TABF evaluation achieves strong inter-evaluator consistency (mean ICC = 0.84) and reveals significant trustworthiness-accuracy trade-offs: systems in the top accuracy quartile achieve a mean trustworthiness score of only 58.3/100 (SD = 11.4), while systems with the highest trustworthiness scores (mean = 79.6, SD = 8.2) accept a mean accuracy reduction of 6.4% relative to unconstrained baselines. The study identifies a Trustworthiness-Accuracy Frontier for each domain and demonstrates that frontier-optimal systems -- those achieving the best trustworthiness at a given accuracy level -- can be systematically identified and promoted through the TABF protocol. The paper contributes the TABF specification, a Trustworthy AI Scorecard (TASC) reporting template, and empirical evidence of the trustworthiness- accuracy trade-off landscape across five high-stakes AI domains.

Author Biography

  • Elena Klein, Associate Professor, Department of Artificial Intelligence, Central European Tech University, Vienna, Austria

    Associate Professor, Department of Artificial Intelligence, Central European Tech University, Vienna, Austria

Downloads

Published

2025-12-15

How to Cite

Benchmarking Frameworks for Trustworthy and Responsible AI Systems. (2025). AI Governance and Society Journal P-ISSN 3117-6097 and E-ISSN 3117-6100, 2(4), 17-24. https://doi.org/10.5281/