高级检索

基于轻量级视听融合适配器的煤岩截割事件定位

Lightweight audio-visual fusion adapter for coal-rock cutting event localization

  • 摘要: 煤岩截割事件定位的目标是通过视频传感器确定既可听又可见截割事件发生的起止时间,并识别其类别。研发具有煤岩截割事件定位功能的视频传感器,对煤矿安全高效开采起着至关重要的作用。然而,煤尘和音频噪声的干扰影响截割事件定位的准确性,且视听多模态数据处理造成的高时延,严重影响视频传感器定位功能的实时性。为解决上述问题,提出一种基于轻量级视听融合适配器(LAVFA)的Swin-LAVFA方法,将LAVFA适配于预训练的移位窗口Transformer(Swin-T)模型,使其具有实时高效定位煤岩截割事件的能力。首先,引入特定模态的压缩模块,利用非对称的交叉注意力机制将高维输入的视觉和音频特征在低维的潜在标记作用下,分别压缩为表征煤岩截割事件线索的视觉和音频特征。然后,利用跨模态交叉注意力模块,将具有事件线索的视觉(或音频)特征与其相对模态的特征融合,获得增强的视觉和音频特征。最后,通过轻量的适配模块,将增强的视听特征适配于每层冻结的Swin-T模型,利用非对称注意力机制迭代地将输入数据提取到一个压缩的潜在瓶颈中,在不损失模型性能的前提下降低模型的计算复杂度。同时,消除Swin-T模型中自注意力计算的二次复杂性问题,提高模型的推理速度。结果表明:该方法在自建的采煤机截割状态(MSCS)数据集中具有优异的煤岩截割事件定位性能(准确率为80.3%,推理速度为33.5 帧/s,可训练参数量为4.6 M),从而辅助视频传感器实时准确地定位煤岩截割事件。

     

    Abstract: Coal-rock cutting event localization is performed to determine the start and end times of cutting events that are both audible and visible and to identify their categories using video sensors. The development of video sensors capable of localizing coal-rock cutting events is crucial for safe and efficient coal mining. However, the accuracy of audio-visual event localization is affected by the interference from coal dust and audio noise. Furthermore, the real-time localization performance of video sensors is significantly affected by the high latency caused by multimodal audio-visual data processing. To solve these problems, a Swin-LAVFA method based on a lightweight audio-visual fusion adapter (LAVFA) is proposed. LAVFA is adapted to the pretrained Shifted Window Transformer (Swin-T) model to enable the real-time and efficient localization of coal-rock cutting events. Specifically, a modality-specific compression module is introduced.Using an asymmetric cross-attention mechanism, the high-dimensional input visual and audio features are compressed into visual and audio features that representing coal-rock cutting event cues, respectively, under the guidance of low-dimensional latent tokens.Then,using a cross-modal cross-attention module, the visual (or audio) features containing event cues are fused with the corresponding-modality features, thereby obtaining enhanced visual and audio features. Finally, the enhanced audio-visual features are adapted to each frozen layer of the Swin-T model through a lightweight adaptation module. During this process, the asymmetric attention mechanism is used to iteratively extract the input data into a compressed latent bottleneck, thereby reducing the computational complexity of the model while perserving its performance. The quadratic complexity of the self-attention computation in the Swin-T model is also eliminated, and the inference speed of the model is consequently improved. The results show that excellent coal-rock cutting event localization performance is achieved by the proposed method on the self-built mine shearer cutting states (MSCS) dataset, with the accuracy of 80.3%, the influence speed of 33.5 fps, and 4.6 M trainable parameters, thereby assisting video sensors to accurately localize coal-rock cutting events in real time.

     

/

返回文章
返回