混合专家协同的多步无监督视频动作定位方法
Multi‑stage unsupervised temporal action localization with mixture‑of‑experts collaboration
  
DOI:
中文关键词:  视频动作定位;混合专家机制;特征解耦;视频理解
英文关键词:temporal action localization; mixture-of-experts(MoE); feature decoupling; video understanding
基金项目:温州图盛控股集团有限公司科技项目(CF058809002023001)资助项目
作者单位
徐钰栋 乐清市电力实业有限公司,浙江 温州 325600 
彭潮海 乐清市电力实业有限公司,浙江 温州 325600 
龙春福 乐清市电力实业有限公司,浙江 温州 325600 
徐文斌 乐清市电力实业有限公司,浙江 温州 325600 
洪雷 乐清市电力实业有限公司,浙江 温州 325600 
唐昊煜 山东大学 软件学院,山东 济南 250101 
摘要点击次数: 157
全文下载次数: 87
中文摘要:
      针对时序动作定位中存在的伪标签噪声累积、跨模态特征融合僵化及前景背景特征纠缠三 大技术瓶颈,提出一种基于混合专家协同的迭代优化框架。该框架包含三大创新模块:其一,设计 了标签优化模块,通过改进的相互邻域机制提纯聚类伪标签,抑制噪声传播;其二,构建了双分支 混合专家网络(DB-MoE),利用门控机制动态融合 RGB与光流特征;其三,提出了动态分离混合专 家网络(DS-MoE),通过显式解耦实现前景与背景表征的分离。在 THUMOS’14 和 ActivityNet v1.2 基准数据集上的大量实验证明,文中方法的性能超越了现有主流方案,达到了领先水平。
英文摘要:
      To address the three major technical bottlenecks in temporal action localization—namely, pseudo-label noise accumulation, static cross-modal feature fusion, and foreground-background feature entanglement—this paper proposes an iterative optimization framework with a mixture-of-experts collaborative optimization mechanism. This framework comprises three innovative core modules: first, a label optimization module that leverages an improved reciprocal-neighbor mechanism to refine clustered pseudo-labels and suppress noise propagation; second, a dual-branch mixture-of-experts (DB-MoE) that utilizes a gating mechanism to dynamically fuse RGB and optical flow features; and third, a dynamic separation MoE(DS-MoE) designed to explicitly disentangle foreground and background representations. Extensive experiments conducted on the THUMOS’14 and ActivityNet v1.2 benchmark datasets demonstrate that the proposed method comprehensively outperforms existing mainstream solutions, achieving state-of-the-art performance.
查看全文  查看/发表评论   附件

你是第5855912访问者
版权所有《南京邮电大学学报(自然科学版)》编辑部
Tel:86-25-85866913 E-mail:xb@njupt.edu.cn
技术支持:本系统由北京勤云科技发展有限公司设计