Abstract:To address the challenges in intelligent power equipment inspection, including heterogeneous information integration, imprecise semantic alignment, and unstable performance under modality deficiencies, this paper proposes an adaptive image-text fusion defect classification framework based on a bidirectional cross-modal attention mechanism. This framework employs vision Transformer (ViT) and bidirectional encoder representation from transformers (BERT) to construct high-dimension visual features and deep semantic representations, respectively. Contrastive learning losses are then introduced to achieve explicit alignment of image and text embeddings within a unified semantic space. The bidirectional cross-modal attention mechanism is designed to adaptively enhance defect-related features and suppress environmental noise through “text→image” and “image→text” interactions. Concurrently, a multi-branch classification strategy is constructed to ensure model robustness under complex conditions such as single-modality deficiency. Experiments demonstrate that the proposed approach achieves an accuracy of 94.5% and an F1 score of 93.9% across 12 typical defect detection tasks in intelligent power inspection scenarios. Compared to existing mainstream unimodal and multimodal fusion methods, the proposed approach exhibits significant improvements in both feature separability and classification accuracy, providing effective support for multimodal processing in power inspection.