| 混合专家协同的多步无监督视频动作定位方法 |
| Multi‑stage unsupervised temporal action localization with mixture‑of‑experts collaboration |
| |
| DOI: |
| 中文关键词: 视频动作定位;混合专家机制;特征解耦;视频理解 |
| 英文关键词:temporal action localization; mixture-of-experts(MoE); feature decoupling; video understanding |
| 基金项目:温州图盛控股集团有限公司科技项目(CF058809002023001)资助项目 |
| 作者 | 单位 | | 徐钰栋 | 乐清市电力实业有限公司,浙江 温州 325600 | | 彭潮海 | 乐清市电力实业有限公司,浙江 温州 325600 | | 龙春福 | 乐清市电力实业有限公司,浙江 温州 325600 | | 徐文斌 | 乐清市电力实业有限公司,浙江 温州 325600 | | 洪雷 | 乐清市电力实业有限公司,浙江 温州 325600 | | 唐昊煜 | 山东大学 软件学院,山东 济南 250101 |
|
| 摘要点击次数: 157 |
| 全文下载次数: 87 |
| 中文摘要: |
| 针对时序动作定位中存在的伪标签噪声累积、跨模态特征融合僵化及前景背景特征纠缠三
大技术瓶颈,提出一种基于混合专家协同的迭代优化框架。该框架包含三大创新模块:其一,设计
了标签优化模块,通过改进的相互邻域机制提纯聚类伪标签,抑制噪声传播;其二,构建了双分支
混合专家网络(DB-MoE),利用门控机制动态融合 RGB与光流特征;其三,提出了动态分离混合专
家网络(DS-MoE),通过显式解耦实现前景与背景表征的分离。在 THUMOS’14 和 ActivityNet v1.2
基准数据集上的大量实验证明,文中方法的性能超越了现有主流方案,达到了领先水平。 |
| 英文摘要: |
| To address the three major technical bottlenecks in temporal action localization—namely,
pseudo-label noise accumulation, static cross-modal feature fusion, and foreground-background feature
entanglement—this paper proposes an iterative optimization framework with a mixture-of-experts collaborative optimization mechanism. This framework comprises three innovative core modules: first, a label
optimization module that leverages an improved reciprocal-neighbor mechanism to refine clustered
pseudo-labels and suppress noise propagation; second, a dual-branch mixture-of-experts (DB-MoE) that
utilizes a gating mechanism to dynamically fuse RGB and optical flow features; and third, a dynamic
separation MoE(DS-MoE) designed to explicitly disentangle foreground and background representations.
Extensive experiments conducted on the THUMOS’14 and ActivityNet v1.2 benchmark datasets demonstrate that the proposed method comprehensively outperforms existing mainstream solutions, achieving
state-of-the-art performance. |
| 查看全文 查看/发表评论 附件 |