Arbovirus surveillance in emergency care units in Salvador, Brazil
Résumé fourni par la source
Introduction: Unrelenting progress in High-throughput sequencing technologies has facilitated infection surveillance.At the same time the initial analysis of these data currently is not routine procedure and requires substantial computational resources to process and interpret the vast amounts of data generated.The COVID-19 pandemic has established genome sequencing of the virus as a standard procedure, facilitated by readily accessible solutions.These solutions depend on mapping reads to a predefined reference.However, for the majority of viruses, this approach is impractical due to substantial genetic diversitythe reference may diverge significantly from the reads present in the sample.For example, the genomic sequence CY103892, which corresponds to the fourth segment of the A/little_yellow-shouldered_bat/Guatemala/060/2010(H17N10) influenza A virus (IAV) strain, and the genomic sequence HQ020375, from the fourth segment of the A/great_cormorant/Qinghai/1/2009(H5N1) IAV strain, share 425 out of 633 nucleotides (67.14%) in less than half of the fourth segment.However, the remaining portion of this segment shows no homology.In this case, the optimal solution would be to select the most suitable reference from a representative database.We demonstrated the effectiveness of this approach using IAV as an example, which confirms its applicability to any other viruses when an appropriate database is available.Methods: A representative database of IAV (n = 872,467) was created for each segment (n = 8) of the IAV based on data obtained from GenBank using the query (txid11320).The software CD-HIT was employed to eliminate sequences that were more than 99.8% identical in nucleotide composition.The selection of the reference was achieved through a direct search for portions of reads in the corresponding database, as demonstrated using the example of the influenza virus ( https://github.com/ana-way/IAVCP_workflow).Thus, a reference was obtained for the most prevalent virus in the population, and it is also possible to detect traces of less prevalent types.Results: To assess the accuracy of the assembly, three samples from experiment SRP483129 were used, where the IAV H1N1 virus was isolated from the nasal wash of ferrets: SRR27490307, SRR27490310, and SRR27490310.The consensus sequences obtained from these three samples were compared with the assembly of the IAV (A/California/07/2009(H1N1)) GCF_001343785.1.The nucleotide sequence's identity ranged between 99% and 100%.IGV visualization demonstrates that the identified substitutions perfectly match the reads, indicating that these discrepancies are unlikely to be errors in data processing.Discussion: Thus, the number of segments determines the number of operations required to compare raw reads with the database.Creating representative databases and applying a modified approach tailored to the number of segments will become a universal solution for the identification of viral genomes.Conclusion: Implementing the proposed approach will streamline the processing of sequencing results, thereby enhancing viral infection surveillance.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Arbovirus surveillance in emergency care units in Salvador, Brazil
- Date Crossref
- 01/03/2025
- Éditeur
- Elsevier BV
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.