Recognizing Human Postures Using Two Versions of Dedicated Transformer Network
Rattachement africain : pl, jp. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
This paper proposes a method for human posture classification using two specialized Transformer architectures. The broader goal is to leverage recognized posture sequences to infer complex human actions, like distinguishing ‘sitting down’ from ‘standing up’, for applications like human–robot interaction and intelligent interface systems. While Transformers excel in NLP and image tasks, their potential for 3D skeleton-based posture classification remains underexplored. State-of-the-art action recognition methods often rely on deep, resource-intensive networks struggling with viewpoint variability and temporal inconsistency, and requiring large annotated datasets. To address these limitations, we introduce two lightweight, effective transformer variants: the Central Posture Transformer (CPT) and Isolated Attention Transformer (IAT). CPT uses acentral posture mechanismto identify and emphasize the most representative frame within a sequence via contextual attention, enhancing interpretability and efficiency. IATdecouples spatial and sequential attention, first modeling joint relationships before integrating temporal dynamics, improving modularity and robustness. These design choices distinguish our models from standard transformers, which compute attention uniformly across all input tokens. Evaluated on the customWUT-20 dataset and publicPKU-MMD benchmark, IAT achieves top accuracy on WUT-20 (0.9981 for 12-frame sequences and 0.9963 for 8-frame sequences), while CPT excels on PKU-MMD (0.9573 and 0.9349). Both outperform baselines such as Ts-CNN, IndRNN, and PoSeqTNet by over 3% in most settings. Qualitative results show CPT’s discriminative posture extraction and IAT’s robustness to temporal resolution changes. Overall, our models achieve high accuracy, computational efficiency, and generalizability across datasets, making them well-suited for real-time human-centered applications in constrained or diverse environments.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Recognizing Human Postures Using Two Versions of Dedicated Transformer Network
- Date Crossref
- 01/01/2025
- Éditeur
- Institute of Electrical and Electronics Engineers (IEEE)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.