BABYCRY-UJM-AXA: A sound library of cries from 1,000 full-term babies aged between 1 and 3 days
Rattachement africain : fr. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Abstract The BABYCRY-UJM-AXA sound library is a comprehensive dataset containing 107,312 cry sequences from 1,000 full-term newborns aged 15.81 to 141.60 hours (104.72 +/- 108.59 sequences per newborn, range = 1-727; mean sequence duration = 4.60 +/- 3.41 seconds, range = 2.40-94.08). Data acquisition relied on continuous, high-fidelity audio recordings conducted in a maternity ward, thereby capturing the natural acoustic environment. To process this large corpus and extract cry segments, we developed an automated pipeline combining active learning and ensemble methods. This pipeline was rigorously validated against human annotations to ensure the reliable exclusion of speech and background noise. The resulting dataset consists of .wav files containing automatically detected cry sequences, each paired with detailed metadata, including anthropometric measurements, pregnancy history, and delivery context. The dataset, the extraction methodology, and a technical validation demonstrating both high audio quality and the ability to reproduce established findings are described in Bonafos et al. (2026, Scientific Data). This resource will support the development of robust machine learning models for cry analysis and facilitate research into the acoustic correlates of neonatal health. Data description BABYCRY-UJM-AXA.zip contains 107,312 individual .wav files, corresponding to cries from 1,000 babies. Audio files are sampled at 44,100 Hz (32-bit). File naming convention: #baby_#cry.wav (#baby = anonymized infant identifier; #cry =sequential cry instance number for that infant). In addition, it contains metadata.csv which resumes, for each wav file, the information associated to the crying baby plus a set of 176 acoustic descriptors. code.zip contains the python code of the pipeline used for the extraction of the cry sequences, as well as the R code used for the technical validation presented in Bonafos et al. (2026, Scientific Data). This dataset was produced as part of the BABYCRY project, a collaborative effort involving researchers from the ENES, Hubert Curien, and SAINBIOSE laboratories at the University of Saint-Étienne (France).
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.