Result data for: Augmenting Canary Analysis with Artificial Intelligence: Large Language Models and Vision-Language Models as the Automated Judge
Rattachement africain : ro, md. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Every result artefact behind the paper, in the directory layout the analysis code expects, together with the scripts that recompute the paper's tables from them. The testbed that produced the measurements is the kayenta-ai-canary-judge repository, archived separately at 10.5281/zenodo.22553878; the scripts here import their metric implementations from it, so a regenerated table is produced by the same code as the published one. Which directories to use. The paper's numbers come from results/corrected/ (the 180-scenario benchmark), results/pilot-45-corrected/ (the 45-scenario pilot) and results/n5-corrected/ (the five-seed sweep). The runs under reference-results/ are the earlier, uncorrected ones; they are retained as evidence, not for use. README.md opens with a table stating which directory is which. Why the uncorrected runs are here. Section 11.2 of the paper reports a defect in the evaluation harness: a tolerant parser turned a failed model call into a well-formed failing verdict at score zero with an empty error field, indistinguishable from a genuine judgement. The affected configurations were re-run in full under the fixed harness, and every published number comes from that re-measurement. The earlier runs are kept beside the corrected ones so that the reported defect can be checked independently rather than taken on trust. PROVENANCE.md documents what the defect was, how it was found, and exactly which rows it touched. Contents: the 180-scenario benchmark before and after correction; the 45-scenario pilot in both forms; the five-seed sweep re-run under the fixed harness; the hosted-model and rubric-ablation runs; the derived tables (confidence intervals, McNemar, ROC/PR, ensemble, false-positive guards) behind the manuscript's tables and figures; and tools/, which maps every manuscript table to the script and artefact that regenerate it. The data is CC-BY-4.0; the scripts under tools/ are Apache-2.0, as in the repository they were written for. See README.md for the layout, PROVENANCE.md for which artefacts are pre- and post-correction, and tools/README.md for the table-by-table regeneration map, including what cannot be regenerated.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.