Aller au contenu principal
Accès ouvert déclaré 2025 dataset

Cell-free DNA methylome and fragmentome analysis for relapse monitoring of Ewing Sarcoma - data processing and classifier building pipelines

0Citations signalées — pas une note de qualité
4Institutions déclarées
2Pays d’affiliation déclarés

Résumé fourni par la source

This repository contains the input files and Nextflow workflows used to generate processed qseaSet objects from T7-MBD-seq data (an enrichment-based methylation capture method using a methyl-binding domain protein), and to construct the EwingSign classifier. OverviewT7-MBD-seq was applied to 86 plasma cell-free DNA (cfDNA) samples derived from Ewing sarcoma (EwS) and CIC-rearranged patients, as well as 107 non-cancer controls (NCC) (83 of which were used for classifier training). After pre-processing fastq files using the in-house T7-MBD-seq Nextflow pipeline, the resulting qseaSet objects were used within a second Nextflow pipeline for the training and application of the EwingSign classifier. This classifier is an ensemble of 25 XGBoost models trained on EwS/CIC methylation array data from Koelsche et al., 2021 (Data Reference: GSE140686). Contents of the repository: 1. Input FilesThese include all necessary input files for running both the T7-MBD-seq and EwingSign classifier Nextflow pipelines. T7-MBD-seq pipeline inputs (T7MBDseqInputFiles.zip):- starterSheet.csv — contains file paths to the FASTQ files.- sampleTable.csv — contains metadata for each cfDNA sample.- config files EwingSign classifier pipeline inputs (ClassifierInputFiles.zip):- CSV files specifying paths to array methylation data (IDAT files)- NCC qseaSet object- CpG_beta_Count_lookupTable.csv — used to convert methylation data into qseaSet objects- Contrasts.csv — defines sample groups for differentially methylated region (DMR) selection- mix.csv — defines group proportions for generating in silico mixture samples 2. Nextflow Pipelines- T7MBDseqNextflowPipeline.zip — Nextflow main workflow, modules, and R scripts for T7-MBD-seq data preprocessing- ClassifierBuildingNextflowPipeline.zip — Nextflow main workflow, modules, and R scripts for training and applying the EwingSign classifier 3. Scripts.zip- Bash scripts to run the T7-MBD-seq Nextflow pipeline for specific cfDNA sample cohorts (EwS, NCC, and NCChigherDepth)- Bash script to run the EwingSign classifier- R script to generate NCC split qseaSet objects from the original NCC qseaSet object- R script to generate CSV files containing the paths to EWS/CIC and NCC split qseaSet objects.- R and bash scripts to generate in silico dilution mixture sets (using array data from Patrizi et al., 2024; Data Reference: GSE276012) and to obtain EwingSign classifier predictions (related to Figure 3B) Code to reproduce all figures in the paper is available at https://doi.org/10.5281/zenodo.17527450 (unrestricted).

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

La source scientifique ouverte est momentanément indisponible.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, ROR et la Banque mondiale, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune donnée externe enregistrée en base. Sources et limites.