基于双向跨模态注意力与自适应图文融合的电力智能巡检缺陷分类方法
CSTR:
作者:
作者单位:

浙江大学电气工程学院,浙江 杭州 310000)

作者简介:

通讯作者:

中图分类号:

基金项目:

国家自然科学基金项目资助(62476242)


An intelligent power inspection defect classification method based on bidirectional cross-modal attention and adaptive image-text fusion
Author:
Affiliation:

College of Electrical Engineering, Zhejiang University, Hangzhou 310000, China

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    针对电力设备智能巡检中异构信息整合困难、语义对齐不精与模态缺失下表现不稳的挑战,提出一种基于双向跨模态注意力机制的自适应图文融合缺陷分类框架。该框架采用视觉Transformer(vision Transformer, ViT)与基于变换器的双向编码器表示(bidirectional encoder representation from transformers, BERT)分别构建高维视觉特征与深层语义表征,并引入基于对比学习的损失,在统一语义空间内实现图文嵌入的显式对齐。设计双向跨模态注意力机制,通过“文本→图像”与“图像→文本”的交互,自适应强化缺陷相关特征、抑制环境干扰。同时构建多分支分类策略,确保模型在单模态缺失等复杂工况下的鲁棒性。实验表明,在电力智能巡检场景12种典型缺陷任务上,该方法准确率达94.5%,F1值达93.9%。与现有主流单模态及多模态融合方法相比,所提方法在特征可分性与分类精度上均有显著提升,为电力巡检多模态处理提供了有效支撑。

    Abstract:

    To address the challenges in intelligent power equipment inspection, including heterogeneous information integration, imprecise semantic alignment, and unstable performance under modality deficiencies, this paper proposes an adaptive image-text fusion defect classification framework based on a bidirectional cross-modal attention mechanism. This framework employs vision Transformer (ViT) and bidirectional encoder representation from transformers (BERT) to construct high-dimension visual features and deep semantic representations, respectively. Contrastive learning losses are then introduced to achieve explicit alignment of image and text embeddings within a unified semantic space. The bidirectional cross-modal attention mechanism is designed to adaptively enhance defect-related features and suppress environmental noise through “text→image” and “image→text” interactions. Concurrently, a multi-branch classification strategy is constructed to ensure model robustness under complex conditions such as single-modality deficiency. Experiments demonstrate that the proposed approach achieves an accuracy of 94.5% and an F1 score of 93.9% across 12 typical defect detection tasks in intelligent power inspection scenarios. Compared to existing mainstream unimodal and multimodal fusion methods, the proposed approach exhibits significant improvements in both feature separability and classification accuracy, providing effective support for multimodal processing in power inspection.

    参考文献
    相似文献
    引证文献
引用本文

周予昊,闫云凤,顾杨昊,等.基于双向跨模态注意力与自适应图文融合的电力智能巡检缺陷分类方法[J].电力系统保护与控制,2026,54(19):177-187.[ZHOU Yuhao, YAN Yunfeng, GU Yanghao, et al. An intelligent power inspection defect classification method based on bidirectional cross-modal attention and adaptive image-text fusion[J]. Power System Protection and Control,2026,V54(19):177-187]

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-01-29
  • 最后修改日期:2026-03-31
  • 录用日期:
  • 在线发布日期: 2026-09-28
  • 出版日期:
文章二维码
关闭
关闭