AI Performance on Image-based Medical Case Scenarios: A Cross-Sectional Comparative Study
Jia-Wei Liu, Yue‐Tong Qian, Xiao Ma, Junping Fan et autres
cn (code pays fourni par la source)
Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.
Jia-Wei Liu, Yue‐Tong Qian, Xiao Ma, Junping Fan et autres
cn (code pays fourni par la source)
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani et autres
We present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g., sports, music, dance, bike repair). 740 participants from 13 cities worldwide performed these activities in 123 …
us, gb, in, sg, co, ca, it, jp, sa (code pays fourni par la source)
Haiming Wang, Mert Unsal, Xiaohan Lin, Mantas Baksys et autres
We introduce Kimina-Prover Preview, a large language model that pioneers a novel reasoning-driven exploration paradigm for formal theorem proving, as showcased in this preview release. Trained with a large-scale reinforcement learning pipeline from Qwen2.5-72B, Kimina-Prover demonstrates strong performance in Lean 4 proof …
Gui-Chen Ling, Chang Su, Yongjie Guo, Xia Qiu et autres
Dermatomyositis (DM) positive for anti-melanoma differentiation-associated gene 5 (MDA5) antibodies, mainly when linked with rapidly progressive interstitial lung disease (RP-ILD), is considered a refractory disease. Our report describes a critical case of clinically amyopathic dermatomyositis (CADM) with RP-ILD that tested positive for …
cn, hk (code pays fourni par la source)
Joya Chen, Zhaoyang Lv, Shiwei Wu, Kevin Qinghong Lin et autres
Recent Large Language Models have been enhanced with vision capabilities, enabling them to comprehend images, videos, and interleaved vision-language content. However, the learning methods of these large multimodal models typically treat videos as predetermined clips, making them less effective and efficient at …
Joya Chen, Zhaoyang Lv, Shiwei Wu, Kevin Qinghong Lin et autres
Recent Large Language Models (LLMs) have been en-hanced with vision capabilities, enabling them to compre-hend images, videos, and interleaved vision-language con-tent. However, the learning methods of these large multi-modal models (LMMs) typically treat videos as predeter-mined clips, rendering them less effective and …
sg, us (code pays fourni par la source)
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani et autres
We present Ego-Exo4D, a diverse, large-scale multi-modal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured ego-centric and exocentric video of skilled human activities (e.g., sports, music, dance, bike repair). 740 participants from 13 cities worldwide performed these activities in 123 …
gb, us, sg, in, bo, ca, it, jp, sa (code pays fourni par la source)
Weifeng Liu, Tianyi She, Jia-Wei Liu, Dongyu Yao et autres
In recent years, DeepFake technology has achieved unprecedented success in high-quality video synthesis, but these methods also pose potential and severe security threats to humanity. DeepFake can be bifurcated into entertainment applications like face swapping and illicit uses such as lip-syncing fraud. …
Jay Zhangjie Wu, Guian Fang, Haoning Wu, Xintao Wang et autres
Generative models have demonstrated remarkable capability in synthesizing high-quality text, images, and videos. For video generation, contemporary text-to-video models exhibit impressive capabilities, crafting visually stunning videos. Nonetheless, evaluating such videos poses significant challenges. Current research predominantly employs automated metrics such as FVD, …
Weifeng Liu, Tianyi She, Jia-Wei Liu, Dongyu Yao
Jia-Wei Liu, Weining Wang, Sihan Chen, Xinxin Zhu et autres
As a combination of visual and audio signals, video is inherently multi-modal. However, existing video generation methods are primarily intended for the synthesis of visual frames, whereas audio signals in realistic videos are disregarded. In this work, we concentrate on a rarely …
cn (code pays fourni par la source)
Sihan Chen, Xinxin Zhu, Dongze Hao, Wei Liu et autres
The quality of video representation directly decides the performance of video related tasks, for both understanding and generation. In this paper, we propose single-modality pretrained feature fusion technique which is composed of reasonable multi-view feature extraction method and designed multi-modality feature fusion …
cn (code pays fourni par la source)
BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.
L'essentiel de l'actu tech du Burkina & d'Afrique, chaque semaine dans votre boîte mail.
Gratuit · sans spam · désinscription en un clic