Accelerating Machine Learning Research in Fusion Through a Pydantic‑Centred Human‑in‑the‑Loop Architecture
Rattachement africain : gb. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Training high-performing ML models requires large volumes of high quality, labelled data. While nuclear fusion experiments generate extensive datasets – JET produced 105,929 pulses, each with more than 10GB of raw data [1] – much of this data lacks the annotations necessary for effective ML training. Retrospective annotation is both challenging and time consuming, as domain experts lack the appropriate tools to effectively label historic pulses. To address this challenge, RSEs and domain experts have collaborated to develop an open-source human-in-the-loop annotation platform for labelling tokamak diagnostic data [2]. This combines a web-based user interface and a REST API built around data models defined using Pydantic [3], with validated annotations stored in a MongoDB [4] database for later reuse. The modular design supports multiple data access layers from different machines, with simple integration of new sources. Custom ML models can be defined by the user, with training and prediction tasks scheduled via the UI and executed using Ray [5]. These reduce the workload for scientists, who can focus on refining model predictions rather than labelling from scratch. This demonstrates how infrastructure co-developed between RSEs and domain experts, focussed on data validation, extensibility and reusability, plays a key role in the research journey. [1]: Vega, J., et al. "New developments at JET in diagnostics, real-time control, data acquisition and information retrieval with potential application to ITER." Fusion Engineering and Design 84.12 (2009): 2136-2144. [2]: TokTagger. https://github.com/ukaea/toktagger [3]: Pydantic. https://github.com/pydantic/pydantic [4]: The MongoDB Database. https://github.com/mongodb/mongo [5]: Ray. https://github.com/ray-project/ray
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.