Aller au contenu principal
Accès ouvert déclaré 2022 article

DIVIS: a semantic DIstance to improve the VISualisation of heterogeneous phenotypic datasets

5Citations signalées, ce qui n’est pas une note de qualité
3Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : fr. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

BACKGROUND: Thanks to the wider spread of high-throughput experimental techniques, biologists are accumulating large amounts of datasets which often mix quantitative and qualitative variables and are not always complete, in particular when they regard phenotypic traits. In order to get a first insight into these datasets and reduce the data matrices size scientists often rely on multivariate analysis techniques. However such approaches are not always easily practicable in particular when faced with mixed datasets. Moreover displaying large numbers of individuals leads to cluttered visualisations which are difficult to interpret. RESULTS: We introduced a new methodology to overcome these limits. Its main feature is a new semantic distance tailored for both quantitative and qualitative variables which allows for a realistic representation of the relationships between individuals (phenotypic descriptions in our case). This semantic distance is based on ontologies which are engineered to represent real-life knowledge regarding the underlying variables. For easier handling by biologists, we incorporated its use into a complete tool, from raw data file to visualisation. Following the distance calculation, the next steps performed by the tool consist in (i) grouping similar individuals, (ii) representing each group by emblematic individuals we call archetypes and (iii) building sparse visualisations based on these archetypes. Our approach was implemented as a Python pipeline and applied to a rosebush dataset including passport and phenotypic data. CONCLUSIONS: The introduction of our new semantic distance and of the archetype concept allowed us to build a comprehensive representation of an incomplete dataset characterised by a large proportion of qualitative data. The methodology described here could have wider use beyond information characterizing organisms or species and beyond plant science. Indeed we could apply the same approach to any mixed dataset.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
DIVIS: a semantic DIstance to improve the VISualisation of heterogeneous phenotypic datasets
Date Crossref
04/04/2022
Éditeur
Springer Science and Business Media LLC
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Institut Agro Rennes-Angers pays non établi dans la notice
    Université ou école supérieure
  • Institut National de Recherche pour l'Agriculture pays non établi dans la notice
    Organisme public
  • L'Institut Agro pays non établi dans la notice
    Université ou école supérieure
  • Univ Angers Institut Agro pays non établi dans la notice
    Université ou école supérieure
  • IRHS - Institut de Recherche en Horticulture et Semences (Institut Agro Agrocampus Ouest pays non établi dans la notice
    Structure de recherche

Institut Agro Rennes-Angers, Institut National de Recherche pour l'Agriculture et L'Institut Agro, avec 2 autres affiliations.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Biomedical Text Mining and OntologiesData Visualization and AnalyticsSpecies Distribution and Climate Change

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.