StatSheets: dataset and reproducibility kit for spreadsheet table understanding
Rattachement africain : fr. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Code and dataset accompanying the CIKM 2026 paper "Structured Prediction for Scalable Spreadsheet Table Understanding: From Cell Types to Table Ranges". Includes the StatSheets benchmark – 737 manually annotated sheets from 14 public statistical organizations, in six languages and five file formats, with complete cell-level cell-type annotations and 818 annotated table ranges – together with the cell-type classification and table detection pipelines and the scripts needed to reproduce all experiments. This archive bundles three differently licensed components. The source code (cell-type-classification/, table-detection/) is released under the MIT License. The StatSheets annotations (dataset/annotations/) are released under CC-BY 4.0. The source spreadsheets (dataset/spreadsheets/) remain the property of their original publishers and are subject to their respective terms of use; dataset/manifest.csv records the origin URL of each file. See DATA-LICENSE.md for details.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.