Aller au contenu principal
Accès ouvert déclaré 2025 article

CHEF-VL: Detecting Cognitive Sequencing Errors in Cooking with Vision-language Models

0Citations signalées, ce qui n’est pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : us. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Minimally obtrusive support for individuals with subjective cognitive decline (SCD) is important for fostering independence in completing daily tasks. In overseeing these tasks, occupational therapists may choose to help as errors arise and provide corrective courses of action. To accomplish this, therapists must be able to recognize task-specific actions, as well as the appropriate sequence for them to occur. However, manual monitoring by therapists is not always feasible in real-world environments, motivating the need for automated systems capable of recognizing actions and detecting sequencing errors. To address this, we present CHEF-VL, an online C ognitive H uman E rror Detection F ramework with V ision -L anguage Models in smart kitchen environments. CHEF-VL combines two novel vision-language models, with one fine-tuned for online human action recognition and the other specially engineered to track key environmental states. An Action-State Merger integrates these two streams of information to reduce prediction noise and correct misrecognized actions. A two-year occupational therapy project of over 100 participants with and without SCD was organized to collect video data for task evaluation. Empirical results demonstrate that CHEF-VL improves both action recognition and sequencing error detection performance, offering a promising solution for real-world assistive technologies in smart home settings.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
CHEF-VL: Detecting Cognitive Sequencing Errors in Cooking with Vision-language Models
Date Crossref
02/12/2025
Éditeur
Association for Computing Machinery (ACM)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Washington University in St. Louis Engineering pays non établi dans la notice
    Université ou école supérieure
  • Washington University School of Medicine in St. Louis Occupational Therapy pays non établi dans la notice
    Université ou école supérieure

Engineering — Washington University in St. Louis et Occupational Therapy — Washington University School of Medicine in St. Louis.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Multimodal Machine Learning ApplicationsHuman Pose and Action RecognitionContext-Aware Activity Recognition Systems

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.