RiboGRAM learns multi-scale non-coding RNA grammar for structural and regulatory inference
Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Protein-coding sequences are read through the triplet genetic code, whereas non-coding RNA sequences encode function through motifs, their positioning, and base-paired architectures. Here, we identify a non-coding RNA grammar defined by multi-scale motif usage and 5'-to-3' positional gradients, and develop RiboGRAM to learn this organization. Across nine organisms, non-coding RNAs showed scale-dependent divergence from coding sequences, stop-codon-like enrichment and directional U-rich and AU-rich gradients. Integrating multi-scale sequence encoding with contextual modeling, RiboGRAM organized RNA families, revealing motif-associated and pair-like patterns across layers. Pretraining on non-coding RNA yielded stronger zero-shot base-pair recovery than coding or messenger RNA pretraining, with these embeddings supporting alignment-free diffusion-based 3D modeling of short RNAs. In human RNA–protein binding profiles, RiboGRAM identified position-resolved motif organization across protein targets and cell types, with sequence-preference changes associated with reduced cross-cell transfer. Together, these results establish a learnable grammar of non-coding RNA sequence organization for structural and regulatory inference.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
Où se fait cette recherche
-
Beijing Academy of Artificial Intelligence pays non établi dans la noticeInstitution
-
Institute of Pharmacy and Molecular Biotechnology pays non établi dans la noticeStructure de recherche
Beijing Academy of Artificial Intelligence et Institute of Pharmacy and Molecular Biotechnology.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.