Aller au contenu principal
Accès ouvert déclaré 2026 dataset

A Systematically Designed Open-Access Dataset for Cross-Microscope Machine Learning Benchmarking in Complex Steel Microstructure Classification

0Citations signalées, ce qui n’est pas une note de qualité
5Institutions déclarées
2Pays d’affiliation déclarés

Rattachement africain : de, ar. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Machine learning (ML) for microstructure analysis has seen rapid progress, yet a fundamental challenge remains largely unaddressed: models trained on one imaging configuration frequently fail to generalize to others, a problem known as domain shift. This issue is particularly acute for complex steel microstructures — specifically bainite and martensite — whose visual appearance is highly sensitive to microscopy modality, magnification, and acquisition settings, and for which no standardized nomenclature or ground truth assignment procedure exists. This dataset was designed to systematically capture and enable the study of these imaging-induced variances. Its defining characteristic is the correlative, pixel-wise registration of identical regions of interest (ROI) imaged across three distinct light optical microscopes — an Olympus LEXT confocal laser scanning microscope, a Leica DM6000, and a Zeiss Axio Imager M2 — at multiple magnifications (20x, 50x, and 100x for the LSM) and under deliberately varied acquisition conditions, including aperture, exposure, and optical filter settings. Since registration to a common reference is performed prior to annotation, ground truth labels apply consistently across all imaging variants of each ROI, ensuring label coherence and eliminating the annotation bottleneck typically associated with multi-condition datasets. The dataset consists of 5,327 annotated image patches extracted from 20 steel samples representing a broad range of quenched and quenched-and-tempered microstructures, covering typical bainitic and martensititic constituents - simplified to two classes. Patches are provided at standardized sizes (256 px at 50x, 96 px at 20x), ready for direct use in standard convolutional neural network (CNN) architectures without further resizing. All micrographs are accompanied by harmonized metadata encoding microscope identity, magnification, pixel size, aperture, exposure level, and filter configuration, assigned via unique identifiers that allow full traceability of each patch to its imaging origin. In the accompanying publication, the dataset was used to systematically evaluate the influence of microscopy modality and magnification on CNN-based bainite/martensite classification, revealing a resolution-dependent asymmetry in cross-domain transfer, systematic class bias inversions under cross-scale validation, and a hierarchy of imaging variance types in their contribution to model robustness. Beyond this initial analysis, the dataset is well suited for cross-microscope and cross-magnification domain adaptation studies, generative image-to-image translation and super-resolution (enabled by the spatially aligned multi-condition image pairs), unsupervised representation and embedding analyses, the development and validation of domain coverage metrics, the study of metadata-informed or physics-aware ML approaches and much more. The dataset complies with FAIR data principles and is intended as a community benchmark for reproducible and transferable ML workflows in microstructure-based materials science.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

La source scientifique ouverte est momentanément indisponible.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.