Aller au contenu principal
Accès ouvert déclaré 2026 preprint

FlashSign: Pose-Free Guidance for Efficient Sign Language Video Generation

0Citations signalées, ce qui n’est pas une note de qualité
2Institutions déclarées
2Pays d’affiliation déclarés

Rattachement africain : ps, us. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Sign language plays a crucial role in bridging communication gaps between the deaf and hard-of-hearing communities. However, existing sign language video generation models often rely on complex intermediate representations, which limits their flexibility and efficiency. In this work, we propose a novel pose-free framework for real-time sign language video generation. Our method eliminates the need for intermediate pose representations by directly mapping natural language text to sign language videos using a diffusion-based approach. We introduce two key innovations: (1) a pose-free generative model based on the a state-of-the-art diffusion backbone, which learns implicit text-to-gesture alignments without pose estimation, and (2) a Trainable Sliding Tile Attention (T-STA) mechanism that accelerates inference by exploiting spatio-temporal locality patterns. Unlike previous training-free sparsity approaches, T-STA integrates trainable sparsity into both training and inference, ensuring consistency and eliminating the train-test gap. This approach significantly reduces computational overhead while maintaining high generation quality, making real-time deployment feasible. Our method increases video generation speed by 3.07x without compromising video quality. Our contributions open new avenues for real-time, high-quality, pose-free sign language synthesis, with potential applications in inclusive communication tools for diverse communities. Code: https://github.com/AIGeeksGroup/FlashSign.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

Aucun DOI disponible pour le contrôle Crossref.

Où se fait cette recherche

  • University College of Applied Science pays non établi dans la notice
    Université ou école supérieure
  • The University of Texas Health Science Center at Houston pays non établi dans la notice
    Université ou école supérieure
  • UCAS pays non établi dans la notice
    Institution
  • UTHealth Houston pays non établi dans la notice
    Institution

University College of Applied Science, The University of Texas Health Science Center at Houston et UCAS, avec 1 autre affiliation.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Hand Gesture Recognition SystemsFace recognition and analysisHuman Pose and Action Recognition

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.