Host Genetic Regulation of NLRP3 Inflammasome Cytokines Reveals Immune and Vascular Pathways in HIV
Rattachement africain : us. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
This repository contains R and shell scripts implementing a full analysis pipeline for whole-genome sequencing (WGS)-based association studies of circulating NLRP3 inflammasome cytokines (IL-6, IL-1β, IL-18) and atherosclerotic vascular outcomes in people with HIV, stratified by ancestry (European, EUR5; African, AFR5). The variant calling pipeline processes raw paired-end whole-genome sequencing reads aligned to GRCh38 (full analysis set with decoy and HLA sequences) through adapter trimming (Trimmomatic v0.39), alignment (BWA-MEM), duplicate marking (GATK MarkDuplicatesSpark), base quality score recalibration (GATK BQSR), per-sample variant calling (GATK HaplotypeCaller in GVCF mode), joint genotyping (GATK GenomicsDBImport and GenotypeGVCFs), and variant quality score recalibration (VQSR at 99.9% truth sensitivity). Post-VCF filtering retained autosomal SNPs with minor allele frequency ≥ 1%, variant missingness ≤ 5%, Hardy–Weinberg equilibrium p > 1×10⁻⁶, and individual missingness ≤ 5%. Association analyses were performed using generalized linear mixed models (GLMM) implemented in GMMAT, with population structure and cryptic relatedness accounted for via a genomic relatedness matrix (GRM) computed using the GENESIS pipeline (KING → PC-AiR → PC-Relate). Five genotype principal components were included as fixed-effect covariates in all models. The following analyses are implemented: Genome-wide association study (GWAS): Genome-wide score tests across all autosomal SNPs passing quality filters (gmmat_gwas.R), with post-processing including QQ and Manhattan plots and LD clumping of genome-wide significant variants (process_gmmat_gwas_results.py) Targeted Wald tests: Effect size estimation for pre-specified SNPs of interest (gmmat_wald.R) Rare variant burden testing: Gene-based aggregation tests using SMMAT (O and E tests) (smmat.R) Transcriptome-wide association study (TWAS): Score tests against PrediXcan-predicted gene expression across four GTEx tissues — Whole Blood, Spleen, Coronary Artery, and Heart Left Ventricle — using tissue-specific elastic net model databases (run_grex.sh, gmmat_twas.R) Mendelian randomization (MR): Two-sample MR using LD-clumped GWAS instruments with IVW, MR Egger, Weighted Median, and MR-PRESSO methods, including heterogeneity and pleiotropy sensitivity analyses (MR_analysis.R) Descriptive and quality control analyses: Cohort descriptive figures and population stratification heatmaps to evaluate ancestry-related trends (descriptive_figures.Rmd, stratification_heatmap.Rmd) Gene set enrichment analysis: Gene set enrichment analysis of the nearest genes from GWAS, gene groups from rare variant analysis, and genes predicted using TWAS to identify biological pathways associated with cytokine levels (GSEA.R) CRISPR functional validation: Functional validation of the top hits using public CRISPR HIV infectivity and perturbation datasets (CRISPR_functional_validation.Rmd) External replication — All of Us IL-6 GWAS: Cross-cohort replication of IL-6 GWAS findings using the All of Us Research Program Controlled Tier Dataset v8 (European and African ancestry strata), run on the All of Us Researcher Workbench (Google Cloud). The pipeline covers data extraction and phenotype preparation from BigQuery with variant QC and PCA via Hail (01_data_extraction_aou.py), GMMAT-based GWAS with PC covariate sweep and genomic inflation factor diagnostics (02_gwas_gmmat.r), SNP overlap and sign-concordance processing against the primary cohort results (03_snp_overlap_processing.r), and allele-harmonized downstream analyses and publication-quality concordance figures (04_downstream_analyses.r, 04_downstream_analyses_notebook.ipynb) Dependencies: R (≥ 4.0) with GMMAT, GENESIS, TwoSampleMR, SeqArray, SeqVarTools, tidyverse, ggplot2, ggrepel, dplyr, forcats, and Matrix. TWAS expression prediction requires Python 3 with PrediXcan (MetaXcan). The All of Us pipeline additionally requires Hail (≥ 0.2), PLINK2, and Google Cloud SDK (gsutil), and must be run within the All of Us Researcher Workbench for Steps 1–2. The variant calling pipeline requires Snakemake, BWA-MEM, GATK, Trimmomatic v0.39, SAMtools, BCFtools, FastQC, and PLINK2. See the repository README for full installation instructions and usage. Recommended workflow order: Snakefile — variant calling from raw reads to filtered VCF Ancestry assignment and phenotype/covariate file preparation genesis_kinship.R — GRM computation gmmat_gwas.R / gmmat_wald.R / smmat.R — association testing process_gmmat_gwas_results.py — GWAS post-processing run_grex.sh → gmmat_twas.R — TWAS MR_analysis.R — Mendelian randomization descriptive_figures.Rmd, stratification_heatmap.Rmd — descriptive analyses GSEA.R — gene set enrichment analysis CRISPR_functional_validation.Rmd — CRISPR functional validation 01_data_extraction_aou.py → 02_gwas_gmmat.r → 03_snp_overlap_processing.r → 04_downstream_analyses.r / 04_downstream_analyses_notebook.ipynb — All of Us external replication (must run Steps 11a–b within the All of Us Researcher Workbench)
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.