PFHAR: Practically Adopting Multi-Modal Foundation Model for Human Activity Recognition Through Edge-Cloud Collaborative Learning
Résumé fourni par la source
Multi-modal human activity recognition (HAR) is a key technology for a wide range of applications and has received widespread attention in recent years. However, the difficulty of achieving generalizability in multi-modal sensing models, combined with heterogeneous and unlabeled downstream data, significantly hinders their broader adoption. In this work, we proposePFHAR, a unified framework for practically adopting multi-modal foundation HAR model to target user groups.PFHARuses a novel dynamic masked contrastive learning method to pre-train a foundation model on various heterogeneous public HAR datasets, ensuring strong generalizability across different modal combinations. It then adopts semi-supervised edge-cloud collaborative learning to fine-tune the pre-trained model with heterogeneous and unlabeled local data, adapting it for the target user group. Our evaluations on public and self-collected datasets demonstrate thatPFHARsignificantly outperforms SOTA baselines in both the pre-training and edge-cloud collaborative fine-tuning stages.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- PFHAR: Practically Adopting Multi-Modal Foundation Model for Human Activity Recognition Through Edge-Cloud Collaborative Learning
- Date Crossref
- 01/08/2026
- Éditeur
- Institute of Electrical and Electronics Engineers (IEEE)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.