Data: Genome assembly, gene prediction and associated proteins of L. flavolineata
Rattachement africain : gb, us. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Files correspond to genome assmbly (*.fna.gz), gene prediction (*.gff) and protein sequences (*.faa) of Liostenogaster flavolineata (NCBI bioproject: PRJNA1415745). Genome annotation V1 RNA-seq reads from mostly adult reproductive and non-reproductive brain tissue, and one whole body dataset were aligned to the genome using STAR v2.7.6a. Gene models were predicted using AUGUSTUS v3.4a trained on Hymenoptera BUSCO v5 models as hints for CDS, exons, and introns matched to RNA-seq data, with BRAKER v3.0.2 and GeneMark v4.69, supplemented by GenomeThreader protein alignments from NCBI Annotation Release 103. Proteins were functionally annotated against the NCBI nr (non-redundant) database using Diamond v2.0.25. Genome annotation V2 Repeats were annotated by HiTE v3.1.2 before being masked by Bedtools v2.31.1. The RNA-seq data was first trimmed by Trimmomatic v0.39 and a quality control was then made using MultiQC v1.25.2 and FastQC v0.12.1 before being mapped onto the genomes using HISAT v2.2.1. Mapped reads were then sorted using Samtools v1.21. Gene prediction was performed by AUGUSTUS v3.6.0, BRAKER v3.0.8 and GeneMark-ETP v1.02 using a protein database containing sequences extracted from the hymenopteran RefSeq database. TSEBRA was configured as follows: P 5; E 10; C 10 ;M 5; intron_support 0.2; stasto_support 1; e_1 0.3; e_2 1; e_3 0.2; e_4 0.2; e_5 0.18; e_6 0.18. The best_by_compleasm.py script (from BRAKER) was then used to keep non-spurious models filtered out by TSEBRA. Odorant receptors (ORs) have been annotated using happy-abcenth v1.0 and OR sequences of hymenopterans find in literature. Some gene families have been annotated using BITACORA v1.4.2 and pfam database v37.2. Gene models were filtered and merged using Bedtools and AGAT v1.4.1. Nucleotide and amino acid CDS sequences were then extracted using AGAT v1.4.1. files correspond to: - *.V2.gff: gene models predicted with alternative isoforms- *.V2.faa: protein sequences of the longest isoforms BUSCO protein mode V1 (evalue threshold = 1e-5, db = hymenoptera_odb10) Number of BUSCOs Category Percentage 4741 Complete and single-copy BUSCOs (S) 79.1% 561 Complete and duplicated BUSCOs (D) 9.4% 163 Fragmented BUSCOs (F) 2.7% 526 Missing BUSCOs (M) 8.8% BUSCO protein mode V2 (same parameters) Number of BUSCOs Category Percentage 5024 Complete and single-copy BUSCOs (S) 83.9% 641 Complete and duplicated BUSCOs (D) 10.7% 99 Fragmented BUSCOs (F) 1.7% 227 Missing BUSCOs (M) 3.7%
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.