Aller au contenu principal
Accès ouvert déclaré 2025 conference-paper

Optimizing Human Pose Estimation Through Focused Human and Joint Regions

3Citations signalées, ce qui n’est pas une note de qualité
4Institutions déclarées
2Pays d’affiliation déclarés

Rattachement africain : cn, sg. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Human pose estimation has given rise to a broad spectrum of novel and compelling applications, including action recognition, sports analysis, as well as surveillance. However, accurate video pose estimation remains an open challenge. One aspect that has been overlooked so far is that existing methods learn motion clues from all pixels rather than focusing on the target human body, making them easily misled and disrupted by unimportant information such as background changes or movements of other people. Additionally, while the current Transformer-based pose estimation methods has demonstrated impressive performance with global modeling, they struggle with local context perception and precise positional identification. In this paper, we try to tackle these challenges from three aspects: (1) We propose a bilayer Human-Keypoint Mask module that performs coarse-to-fine visual token refinement, which gradually zooms in on the target human body and keypoints while masking out unimportant figure regions. (2) We further introduce a novel deformable cross attention mechanism and a bidirectional separation strategy to adaptively aggregate spatial and temporal motion clues from constrained surrounding contexts. (3) We mathematically formulate the deformable cross attention, constraining that the model focuses solely on the regions centered at the target person body. Empirically, our method achieves state-of-the-art performance on three large-scale benchmark datasets. A remarkable highlight is that our method achieves an 84.8 mean Average Precision (mAP) on the challenging wrist joint, which significantly outperforms the 81.5 mAP achieved by the current state-of-the-art method on the PoseTrack2017 dataset.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Optimizing Human Pose Estimation Through Focused Human and Joint Regions
Date Crossref
11/04/2025
Éditeur
Association for the Advancement of Artificial Intelligence (AAAI)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Jilin University pays non établi dans la notice
    Université ou école supérieure
  • Zhejiang Gongshang University pays non établi dans la notice
    Université ou école supérieure
  • National University of Singapore pays non établi dans la notice
    Université ou école supérieure
  • College of Computer Science and Technology pays non établi dans la notice
    Université ou école supérieure
  • Zhejiang University Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security The State Key Laboratory of Blockchain and Data Security pays non établi dans la notice
    Université ou école supérieure
  • School of Computing pays non établi dans la notice
    Université ou école supérieure

Jilin University, Zhejiang Gongshang University et National University of Singapore, avec 3 autres affiliations.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Human Pose and Action RecognitionGait Recognition and AnalysisHand Gesture Recognition Systems

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.