Deposit for Extended datas and extended codes for the publication : PIGSTI: a modular, reproducible pipeline for detecting species identity, pathogens, and microbes from animal palaeogenomic data (Large files and codes)
Résumé fourni par la source
Repository for large dataset and modified code PIGSTI: a modular, reproducible pipeline for detecting species identity, pathogens, and microbes from animal palaeogenomic data Repository for large dataset and modified code in the paper above. This Zenodo record (doi:10.5281/zenodo.22135224) accompanies the PIGSTI F1000Research software article. It contains supplementary screening tables (952-sample empirical cohort and simulation benchmarking), and modified PathPhynder source code used for phylogenetic placement. PIGSTI pipeline: GitHub — LouisLhote/PIGSTI Contents of this archive File / folder Description Large_tables.xlsx Excel workbook with supplementary tables S1–S8 (see below). Pathphynder_v1.2.3_modif/ Modified PathPhynder (v1.2.3) with proportional alternative-allele tolerance (--maximumToleranceProp). Large_tables.xlsx — sheet guide S1 — Pathogen panel and thresholds Pathogen reference panel screened in the empirical analysis, with per-pathogen E-value and read-count thresholds. Corresponds to the Pathogen_spreadsheet.csv configuration described in the Methods. S2 — Newly generated library sequencing data Sequencing data for the datasets newly generated in this study. For each library: sample of origin, library identifier, UDG treatment, and number of raw read pairs. S3 — Simulated benchmarking results (60 datasets) Detection scores and authentication metrics for the 60 simulated benchmarking datasets (10 pathogens × 6 abundance tiers). Used to derive the high-confidence filter thresholds applied to empirical data (Figure 2). S4 — Sample metadata (952 datasets) Metadata for all 952 screened datasets, including sample identifiers, accessions, host taxon, material, age, and provenance. Comprises 799 previously published datasets from metAaRCive (doi:10.5281/zenodo.20085065) and 153 libraries newly generated for this study. S5 — Candidate pathogen hits (1,023 rows) All 1,023 candidate hits passing the initial Guellil E-value threshold (> 0.001), with screening, mapping, and authentication metrics. S6 — High-confidence pathogen hits (109 rows) The 109 hits passing the high-confidence filter set derived from simulation benchmarking: relative entropy ≥ 0.45, genus ranking = 1, edit-distance decay without damage ≥ 0.70, breadth ratio ≥ 0.50, ANI > 96.5%, and mapping ratio > 0.50. S7 — Samples with authenticated detections (89 rows) The 89 samples carrying at least one authenticated pathogen detection (89 of 952 screened samples). S8 — Newly generated sample metadata (153 datasets) Metadata for the 153 datasets newly generated in this study, including archaeological labels and internal laboratory codes. Archaeological context for newly reported sites is described in the paper (Description of archaeological sites for newly reported material). Pathphynder_v1.2.3_modif/ — modified PathPhynder Standard PathPhynder (Martiniano et al.) places ancient samples on a reference phylogeny using branch-defining SNPs. For low-coverage pathogen BAMs, a fixed maximum count of conflicting alternative alleles can stop traversal too early. This modified v1.2.3 adds: CLI option --maximumToleranceProp (0–1): maximum proportion of conflicting ALT alleles allowed per branch before stopping traversal. When set, this replaces the fixed-count rule for traversal (see pathphynder_maxToleranceProp.patch in this folder). Paper setting: --maximumToleranceProp 0.45 for Rickettsia, Leptospira, and Erysipelothrix placements (Methods — phylogenetic placement). Upstream reference: Martiniano R, et al. Placing Ancient DNA Sequences into Reference Phylogenies. Mol Biol Evol (2022). PathPhynder Licensing Component License Large_tables.xlsx (S1–S8) CC-BY 4.0 Pathphynder_v1.2.3_modif/ MIT — see Pathphynder_v1.2.3_modif/LICENSE Contact Louis L'Hôte — lhtel@tcd.ie