Aller au contenu principal
Accès ouvert déclaré 2026 conference-paper

VideoExplorer: Advancing Long-Horizon Video Understanding via Hierarchical Orchestration

0Citations signalées, ce qui n’est pas une note de qualité
5Institutions déclarées
2Pays d’affiliation déclarés

Rattachement africain : cn, it. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Current agentic frameworks for Long-Video Understanding (LVU) remain limited by two critical problems: ineffective control, where traditional monolithic agents struggle with high-branching, multi-granularity decision processes; and inefficient supervision, where sparse, outcome-based feedback fails to guide long-horizon reasoning. To resolve these challenges, we propose VideoExplorer, a novel agentic system designed to advance long-video reasoning on top of structured control and trajectory-level optimization. First, VideoExplorer innovates a hierarchically orchestrated framework: it employs a planning agent to focus on creating high-level reasoning strategies and specialized sub-agents (including a temporal grounder and a visual perceiver) to accomplish fine-grained reasoning executions, thereby substantially reducing the complexity of reasoning process. Second, VideoExplorer introduces a novel optimization approach, Trajectory level Direct Preference Optimization (TDPO), to mitigate inefficient supervision. Unlike standard methods that optimize single turn responses, TDPO aligns the entire planning trajectory, including evidence routing and termination decisions, with end task success, which effectively mitigates premature commitment and compounding errors. To better support the conduct of TDPO, we further create fine-grained supervision data via a teacher-guided, difficulty-adaptive sampling process. Extensive experiments on MLVU, LVBench, and MH-NIAH demonstrate that VideoExplorer consistently outperforms monolithic baselines in both accuracy and efficiency, validating the effectiveness of structured control and trajectory-level optimization in long-video reasoning. Our code is available in this repository.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
VideoExplorer: Advancing Long-Horizon Video Understanding via Hierarchical Orchestration
Date Crossref
08/08/2026
Éditeur
ACM
Type
proceedings-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Renmin University of China pays non établi dans la notice
    Université ou école supérieure
  • Beijing Academy of Artificial Intelligence pays non établi dans la notice
    Institution
  • Beijing University of Posts and Telecommunications pays non établi dans la notice
    Université ou école supérieure
  • Peking University pays non établi dans la notice
    Université ou école supérieure
  • University of Trento pays non établi dans la notice
    Université ou école supérieure
  • Gaoling School of Artificial Intelligence pays non établi dans la notice
    Université ou école supérieure
  • China and Hong Kong Polytechnic University Beijing pays non établi dans la notice
    Université ou école supérieure

Renmin University of China, Beijing Academy of Artificial Intelligence et Beijing University of Posts and Telecommunications, avec 4 autres affiliations.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.