Advance Search
Xu Qian,Wu Yunxia,Zou Zhengyang. Lightweight audio-visual fusion adapter for coal-rock cutting event localizationJ. Coal Science and Technology,2026,54(8):371−382. DOI: 10.12438/cst.2025-1264
Citation: Xu Qian,Wu Yunxia,Zou Zhengyang. Lightweight audio-visual fusion adapter for coal-rock cutting event localizationJ. Coal Science and Technology,2026,54(8):371−382. DOI: 10.12438/cst.2025-1264

Lightweight audio-visual fusion adapter for coal-rock cutting event localization

  • Coal-rock cutting event localization is performed to determine the start and end times of cutting events that are both audible and visible and to identify their categories using video sensors. The development of video sensors capable of localizing coal-rock cutting events is crucial for safe and efficient coal mining. However, the accuracy of audio-visual event localization is affected by the interference from coal dust and audio noise. Furthermore, the real-time localization performance of video sensors is significantly affected by the high latency caused by multimodal audio-visual data processing. To solve these problems, a Swin-LAVFA method based on a lightweight audio-visual fusion adapter (LAVFA) is proposed. LAVFA is adapted to the pretrained Shifted Window Transformer (Swin-T) model to enable the real-time and efficient localization of coal-rock cutting events. Specifically, a modality-specific compression module is introduced.Using an asymmetric cross-attention mechanism, the high-dimensional input visual and audio features are compressed into visual and audio features that representing coal-rock cutting event cues, respectively, under the guidance of low-dimensional latent tokens.Then,using a cross-modal cross-attention module, the visual (or audio) features containing event cues are fused with the corresponding-modality features, thereby obtaining enhanced visual and audio features. Finally, the enhanced audio-visual features are adapted to each frozen layer of the Swin-T model through a lightweight adaptation module. During this process, the asymmetric attention mechanism is used to iteratively extract the input data into a compressed latent bottleneck, thereby reducing the computational complexity of the model while perserving its performance. The quadratic complexity of the self-attention computation in the Swin-T model is also eliminated, and the inference speed of the model is consequently improved. The results show that excellent coal-rock cutting event localization performance is achieved by the proposed method on the self-built mine shearer cutting states (MSCS) dataset, with the accuracy of 80.3%, the influence speed of 33.5 fps, and 4.6 M trainable parameters, thereby assisting video sensors to accurately localize coal-rock cutting events in real time.
  • loading

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return