Aller au contenu principal
2025 conference-abstract

Abstract 1064: Cancer data aggregator: a new cancer data discovery tool

1Citations signalées, ce qui n’est pas une note de qualité
4Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : us. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Cancer research data is generated and managed in silos. Data repositories frequently use different models and vocabularies to describe their data holdings, and many specialize in certain data types, necessitating that data from a large integrative project be spread across several repositories. For researchers who want to reuse existing cancer data, just navigating this ecosystem of repositories, formats and models can be overwhelming. The Cancer Data Aggregator (CDA) is a new free National Cancer Institute (NCI) service that aims to increase the discoverability of cancer datasets by making them searchable from a common interface, using a single model and vocabulary. We combine metadata for thousands of studies hosted at multiple data repositories across the NCI, and make that metadata available for search from a unified database, so researchers can more easily find and reuse existing cancer research data. The CDA ensures that researchers can find what they’re looking for using a single set of search terms by thoroughly cleaning, harmonizing, and cross-referencing the study metadata. This allows researchers to find subjects that have participated in multiple studies, discover data from a disease that was originally described in different ways at each repository, or compile all the data from a large consortium project like CPTAC - no matter how fragmented it ended up. With the CDA, one set of search terms can return results from data repositories including the Genomic Data Commons, the Proteomic Data Commons, the Imaging Data Commons, Cancer Data Services, and the Integrated Canine Data Commons. Results are displayed in a simple table format (tsv) that is compatible with all standard spreadsheet apps (Excel, Google sheets, Numbers, etc) or loaded directly for analysis in a Jupyter notebook. Results include relevant clinical, phenotypic, genotypic values and provenance details, as well as information about where data files can be downloaded from using the GA4GH DRS API. The Cancer Data Aggregator publishes a new release at the end of each month, and only publishes public data. This ensures that all researchers can use the CDA anonymously without any special privileges, and can always expect data to be up to date. Non-computational users can search the CDA using fill-in-the-blank style queries, while power users can create software for custom searches using our Python library. Visit https://cda.readthedocs.io/ to get started. Citation Format: Amanda Charbonneau, Arthur Brady, Tanner Coon, Rachel Kutner, Surya Saha, Katherine Thayer, Alexander Baumann, Todd Phil, Henry Schaefer, Heather Creasy, Jack DiGiovanna, David Pot, Bing-Xing Huo. Cancer data aggregator: a new cancer data discovery tool [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 1064.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Abstract 1064: Cancer data aggregator: a new cancer data discovery tool
Date Crossref
21/04/2025
Éditeur
American Association for Cancer Research (AACR)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Artificial Intelligence in HealthcareCardiovascular Health and Risk FactorsBioinformatics and Genomic Networks

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.