Responsible Data Collection Strategies for Machine Learning Applications
DOI:
https://doi.org/10.5281/Keywords:
responsible data collection, machine learning, data quality, representativeness, collection ethics, EU AI Act, GDPR, data governanceAbstract
Responsible data collection -- the application of ethical, legal, and technical principles to the design, execution, and documentation of data collection activities for machine learning -- is a foundational but frequently underspecified requirement of responsible AI. Data collection decisions made at the earliest stage of the ML pipeline propagate through model training and deployment, making upstream responsible collection the most cost-effective intervention point for preventing downstream ethical failures. Despite this, the responsible AI literature has devoted substantially more attention to post-collection bias mitigation and model fairness than to the principled design of data collection processes. This paper proposes the Responsible Data Collection Framework (RDCF), a structured methodology for planning, executing, and documenting ML data collection that embeds ethical, legal, and quality requirements from the earliest design stage. The RDCF comprises six collection design principles and fifteen collection practice specifications, organised into four phases: collection planning, source and sampling design, collection execution, and collection documentation. The framework is evaluated through application to twenty-two data collection projects across five ML domains and through a survey of 112 ML practitioners assessing current responsible collection practice. Survey results document substantial gaps: only 34.8% of practitioners systematically assess representativeness before collection, 21.4% conduct pre-collection ethics review, and 18.3% maintain collection decision logs. RDCF-aligned projects achieve a 48.6% lower post-collection bias indicator count (p < 0.001) and 91.3% EU AI Act Article 10 documentation completeness compared to 43.7% for non-RDCF projects. The study contributes the RDCF specification, a Collection Responsibility Assessment (CRA) instrument, and the first large-scale empirical benchmark of responsible data collection practice in the ML community.

