面向边缘智能的可见光-红外图像融合权重在线学习方法综述

A Review of Online Learning Methods for Weighted Fusion of Visible Light-Infrared Images for Edge Intelligence

  • 摘要: 面向边缘智能场景下可见光-红外跨模态融合任务中融合权重实时自适应调整的核心需求,从跨模态特征融合、在线凸优化、轻量化神经网络与神经架构搜索四个维度系统综述该领域的研究进展。首先,梳理了基于注意力机制的动态融合权重分配方法,分析了交叉注意力、通道注意力与自注意力机制在模态互补性挖掘中的应用,并通过对比实验揭示了不同融合策略在信息熵、标准差、空间频率等关键指标上的性能差异;其次,阐述了在线凸优化理论框架及其在实时权重更新中的数学基础,探讨了无梯度优化与Bandit反馈场景下的遗憾界分析,建立了理论保证与工程实践之间的桥梁;再次,总结了轻量化神经网络架构设计与模型压缩技术,包括知识蒸馏、网络剪枝与量化方法,对比分析了SqueezeNet、MobileNet系列及EfficientNet等主流架构的计算效率与精度权衡;最后,评述了硬件感知神经架构搜索在超微算力约束下的自动化网络设计策略。在此基础上,深入剖析了当前研究面临的三大核心挑战:异构模态特征对齐与语义一致性建模、“硬件-模型-实时性-准确性”的耦合优化、以及超微算力约束下的在线学习机理,并展望了神经符号融合、持续学习与硬件协同设计等未来研究方向。

     

    Abstract: Addressing the core requirement of real-time adaptive adjustment of fusion weights in visible–infrared cross-modal fusion tasks for edge intelligence scenarios, this paper systematically reviews the research progress in this field from four dimensions: cross-modal feature fusion, online convex optimization, lightweight neural networks, and neural architecture search. First, the dynamic fusion weight allocation methods based on attention mechanisms are summarized, and the applications of cross-attention, channel attention, and self-attention mechanisms in mining modal complementarity are analyzed; comparative experiments are conducted to reveal the performance differences among various fusion strategies in terms of key metrics such as information entropy, standard deviation, and spatial frequency. Second, the theoretical framework of online convex optimization and its mathematical foundation for real-time weight updating are elaborated, and the regret bound analysis under gradient-free optimization and bandit feedback scenarios is discussed, establishing a bridge between theoretical guarantees and engineering practice. Third, lightweight neural network architecture design and model compression techniques are summarized, including knowledge distillation, network pruning, and quantization methods; the computational efficiency and accuracy trade-offs of mainstream architectures such as SqueezeNet, MobileNet series, and EfficientNet are comparatively analyzed. Finally, hardware-aware neural architecture search for automated network design under ultra-low computational power constraints is reviewed. On this basis, three core challenges facing current research are thoroughly analyzed: heterogeneous modal feature alignment and semantic consistency modeling, the coupled optimization of “hardware–model–real-time–accuracy”, and the online learning mechanism under ultra-low computational power constraints; future research directions such as neuro-symbolic fusion, continual learning, and hardware–software co-design are also prospected.

     

/

返回文章
返回