AI Evaluation Practices Consensus
Le résumé fourni par la source
This protocol outlines a preregistered, cross-sector Delphi process to produce (1) a summary of AI evaluation practices recommended by a broad spectrum of participants, and (2) a comprehensive transparent record of expert agreement and disagreement. The process sets explicit standards for both consensus and disagreements. This can support expectation-setting for evaluations across the AI ecosystem, while remaining explicitly non-binding and non-prescriptive. The resulting consensus hopes to promote better performance by evaluators, and informed use of evaluations by different actors in the AI ecosystem, while allowing future revision as technologies, norms, and contexts evolve. The project uses a modified Delphi-conference method to elicit participant views. (Linstone & Turoff, 2002) It specifies—in advance—the study’s design, governance, stratification, rating scales, consensus thresholds, aggregation logic, stopping rules, and deviation handling. As such, it is a methodological commitment and reference artifact, similar to a registered report or preregistration, and does not discuss evaluation practices themselves. This structure, and the outputs, are designed to improve coordination and transparency around AI evaluation practices by reporting areas of agreement, conditional agreement, and persistent disagreement across stakeholder groups. Disagreements about evaluation practices are highlighted, not suppressed, especially since they reflect not just technical debates, but differences in institutional role, incentives, downstream use, and normative assumptions. Note that this preregistration was largely written by GPT5.2 on the basis of project notes and documents, but all text was reviewed by the submitter.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.