P340: IN SEARCH FOR GENOMIC AND TRANSCRIPTOMIC PREDICTORS OF RELAPSE IN PEDIATRIC T-ALL
Résumé fourni par la source
Background: Prognostic markers are urgently needed in T-ALL since risk stratification is mainly MRD-based. Integrative omics might aid the identification of prognostic markers. Aims: We used whole genome (WGS) and transcriptome sequencing (mRNA-seq) 1/ to identify genes affected by copy number alterations (CNA) associated with refractory/relapsed (RR) course of T-ALL; 2/ to propose a classifier of RR T-ALL based on gene expression at diagnosis. Methods: 63 samples from diagnosis (Dg) and 24 matched remission samples (panel of controls) were analyzed by WGS. Reads were aligned to GRCh38 reference genome and pre-processed using GATKv4.2.6.1. Somatic variants were identified with Mutect2 and annotated with VariantEffectPredictorv107. CNAs were identified using GATKv4.1.1.2 and matched with genes using RefSeq exon coordinates. Percentage of blasts in Dg samples was used to correct for tumor purity. We defined the boundaries of recurrent CNAs using GISTIC2.0. We compared distribution of recurrent CNAs in Dg samples of 17 patients with subsequently RR vs. 46 non-refractory/relapsed (non-RR) patients. We integrated WGS with mRNA-seq, available for Dg samples of 53/63 patients. RNA-seq reads were aligned to GRCh38. We used CNVRanger to identify genes which expression associated with CNAs. Using edgeR, we identified genes differentially expressed (DEGs) between RR and non-RR. Using caret R library, we created a SVM classifier, evaluated by 10-fold cross-validation, with down-sampling to account for imbalanced dataset. Cross-validation loop included feature selection, set to select 1 to max. 20 features independently. Results: We identified 1431 CNAs, showing high heterogeneity across samples. Unsupervised clustering of samples based on genomic profile of CNAs did not separate RR vs. non-RR cases. Recurrent CNAs with significantly different frequency (p<0.05) in RR vs. non-RR patients included only 4 regions of copy number gains (CN_gains) and 4 regions of losses (CN_losses). Recurrent CN_gains affected: cyclin‐dependent kinase 6 gene, CDK6 (chr7); 3 histone genes HIST1H4B, HIST1H3B, HIST1H2AB (chr6), and 2 larger regions encompassing 14 protein-coding genes (chr4), and 98 protein-coding genes (chr16); all regions gained in 3 RR vs. 0 non-RR patients. Recurrent CN_losses affected: gene of tubulin polymerization promoting protein, TPPP (chr5; loss in 3 RR vs. 0 non-RR); P53 regulated DNA replication inhibitor gene, KLLN (chr10; 4 RR vs. 0 non-RR), a region in chr10 encompassing 9 protein-coding genes and 7 pseudogenes (3 RR vs. 0 non-RR), and a region in chr9 including MLLT3 super elongation complex subunit, MLLT3, 3 pseudogenes, 2 miRNA genes, 1 lncRNA (0 RR vs. 12 non-RR). P-values for recurrent CNAs were >0.05, when adjusted for multiple testing. CNVRanger revealed 119 genes which expression associated with CNAs. These did not include genes affected by recurrent CNAs differing RR vs. non-RR. In search for transcriptomic predictors of relapse, we identified 357 DEGs differing RR vs. non-RR, 11 of which were used in a classifier (NPIPA8; EML4; PTPRC; APC; KIDINS220; VEZF1; LNPEP; PPP2R5E; SPPL2A; MFAP3; LOC105378667) with AUC: 0,928. Summary/Conclusion: CNAs reflect high genomic heterogeneity of T-ALL and seem suboptimal candidates for relapse predictors. Gene expression at diagnosis efficiently discriminates RR vs. non-RR and enabled to develop an 11-gene classifier which needs to be independently validated. We observed an advantage of transcriptomic over genomic features in prediction of T-ALL relapse. Funding: NCRD Poland: STARTEGMED3/30456/5/NCBR/2017, European Union’s Horizon 2020 Grant: 952304Keywords: relapsed/refractory, Gene expression profile, Gene dosage, T cell acute lymphoblastic leukemia