[1]阮 楠,张冬旭,王普雄.双重语义增强蒸馏的图书错序云边检测[J].机械与电子,2026,44(06):63-70.
 RUAN Nan,ZHANG Dongxu,WANG Puxiong.Dual Semantic Enhanced Distillation for Cloud-edge Detection of Misplaced Books[J].Machinery & Electronics,2026,44(06):63-70.
点击复制

双重语义增强蒸馏的图书错序云边检测()
分享到:

《机械与电子》[ISSN:1001-2257/CN:52-1052/TH]

卷:
44
期数:
2026年06期
页码:
63-70
栏目:
智能检测
出版日期:
2026-06-27

文章信息/Info

Title:
Dual Semantic Enhanced Distillation for Cloud-edge Detection of Misplaced Books
文章编号:
1001-2257 ( 2026 ) 06-0063-08
作者:
阮 楠 1 张冬旭 2 王普雄 1
1. 西安工程大学图书馆,陕西 西安 710048 ;
2. 西安工程大学计算机科学学院,陕西 西安 710048
Author(s):
RUAN Nan1 ZHANG Dongxu2 WANG Puxiong1
( 1.Library of Xi ’ an Polytechnic University , Xi ’ an 710048 , China ;
2.School of Computer Science , Xi ’ an Polytechnic University , Xi ’ an 710048 , China )
关键词:
图书错序检测目标检测知识蒸馏DINOv3 云边协同视觉语言模型
Keywords:
book disorder detection object detection heterogeneous knowledge distillation DINOv3 cloud – edge collaboration vision-language model
分类号:
TP391.41 ;G250.7
文献标志码:
A
摘要:
针对高校图书馆智能化转型中在架图书错序管理面临的图书密集排列、纹理高度相似及光照环境复杂等多重挑战,提出一种基于异构骨干网络与双重语义增强蒸馏的错序检测方法。首先,构建以 ResNet101 为教师网络、 MobileNetV2 为学生网络的异构蒸馏架构,通过多尺度特征空间的严格投影对齐,实现深层语义表征向轻量化模型的高效迁移。其次,引入书籍特征增强模块,借助三级级联机制将 DINOv3 的全局语义先验无缝注入骨干网络,突破传统检测器在细粒度特征提取方面的局限。同时,设计基于概率特征分布的颜色 形状知识蒸馏策略,并结合云边协同的视觉语言模型推理框架,实现从像素级定位到语义级逻辑判别的全链路优化。实验结果表明,该方法在复杂场景下多次实验达到的平均检测精度( mAP@0.5 : 0.95 )达 82.1% ,边缘端推理速度达 124.3 帧/ s ,显著优于现有实时目标检测模型,表现出良好的性能与实用价值。
Abstract:
To address the multiple challenges in on shelf book misplacement management during the intelligent transformation of university libraries , such as denselyarranged books , highly similar textures , and complex lighting conditions , a novel misplacement detection method based on heterogeneous backbone networks and dual semantic enhanced distillation is proposed.First , a heterogeneous distillation architecture is constructedwith ResNet101 as the teacher network and MobileNetV2 as the student network , achieving efficient transfer of deep semantic representations to a lightweight inference model through strict projection alignment in multi-scale feature spaces.Subsequently , a Book Feature Enhancement Module ( BFEM ) is introduced to seamlessly inject the global semantic priors of DINOv3 into the backbone network through a three level cascading mechanism , overcoming the limitations of traditional detectors in fine grained feature extraction.Additionally , a color shape knowledge distillation strategy based on probabilistic feature distribution is designed and combined with a cloud edge collaborative Vision-Language Model( VLM ) inference framework , enablingfull chain optimization from pixel level localization to semantic level logical discrimination.Experimental results demonstrate that the proposed method achieves an average detection accuracy ( mAP@0.5 : 0.95 ) of 82.1% through multiple experiments in complex scenarios , with an edge-side inference speed of 124.3 FPS , significantly outperforming existing real-time object detectors and exhibiting superior performance and practical value.

参考文献/References:

[ 1 ] 王红芳,刘泽远,李英健,等 . 基于改进的 YOLOv3-Tiny 深度网络在架图书错序检测方法[ J ] . 现代电子技术,2022 , 45 ( 22 ): 164-170.

[ 2 ] 王红芳,武薛,宣静雯 . 基于图书索书号识别的在架图书错序检测方法[ J ] . 新世纪图书馆, 2023 ( 1 ): 31-36.
[ 3 ] 王红芳,武薛,宣静雯,等 . 图书馆在架图书乱架错序检测系统设计与实现[ J ] . 微型电脑应用, 2023 , 39 ( 8 ): 26-28.
[ 4 ] Zhao Yi ’ an , Lv Wenyu , Xu Shangliang , et al.DETRS beat YOLOs on real-time object detection [ C ] ∥2024 IEEE / CVF Conference on Computer Vision and Pattern Recognition ( CVPR ) .New York : IEEE , 2024 :16965-16974.
[ 5 ] Peng Yansong , Li Hebei , Wu Peixi , et al.D-FINE : redefine regression task in DETRs as fine-grained distribution refinement [ C ] ∥ International Conference on Learning Representations 2025 ( ICLR 2025 ), 2025.
[ 6 ] Huang Sihua , Lu Zhichao , Cun Xiaodong , et al.DEIM : DETR with improved matching for fast convergence [ C ] ∥2025 IEEE / CVF Conference on Computer Vision and Pattern Recognition ( CVPR ) .New York : IEEE , 2025 : 15162-15171.
[ 7 ] Zhou Yang , Gao Xu , Chen Zichong , et al.Attention distillation : a unified approach to visual characteristics transfer [ C ] ∥2025 Computer Vision and Pattern Recognition Conference ( CVPR ) .New York : IEEE , 2025 : 18270-18280.
[ 8 ] Shu Changyong , Liu Yifan , Gao Jianfei , et al.Channel wise knowledge distillation for dense prediction [ C ] ∥ 2021 IEEE / CVF International Conference on Computer Vision ( ICCV ) .New York : IEEE , 2021 : 5311-5320.
[ 9 ] Yang Zhendong , Li Zhe , Shao Mingqi , et al.Masked generative distillation [ C ] ∥European Conference on Computer Vision ( ECCV 2022 ) .Cham : Springer Nature Switzerland , 2022 : 53-69.
[ 10 ] Li Wentao , Zhao Danpei , Yuan Bo , et al.PETDet : Proposal enhancement for two-stage fine-grained object detection [ J ] .IEEE Transactions on Geoscience and Remote Sensing , 2023 , 62 : 1-14.
[ 11 ] Xie Xingxing , Cheng Gong , Li Wenbo , et al.Learning discriminative representation for fine-grained object detection in remote sensing images [ J ] .IEEE Transactions on Circuits and Systems for Video Technology , 2025 , 35 : 8197-8208.
[ 12 ] Siméoni O , Vo H V , Seitzer M , et al.DINOv3 [ PP / OL ] .V1.arXiv ( 2025-08-13 )[ 2026-01-03 ] .https : ∥doi.org / 10.48550 / arXiv.2508.1010.
[ 13 ] Bai Shuai , Cai Yuxuan , Chen Ruizhe , et al.Qwen3 VL technical report [ PP / OL ] .V2.arXiv( 2025-11-27 )[ 2026-01-03 ] .https : ∥doi.org / 10.48550 / arXiv.2511.21631.
[ 14 ] 张凯兵,周茜茜,吴昊,等 . 融合全局感知与坐标注意力的道路缺陷检测[ J ] . 西安工程大学学报, 2025 , 39( 4 ): 97-106.
[ 15 ] Varghese R , Sambath M.YOLOv8 : a novel object detection algorithm with enhanced performance and robustness [ C ] ∥2024 International Conference on Advances in Data Engineering and Intelligent Computing Systems ( ADICS ) .New York : IEEE , 2024 : 1-6.
[ 16 ] Wang Ao , Chen Hui , Liu Lihao , et al.YOLOv10 : real time end-to-end object detection [ C ] ∥Advances in Neural Information Processing Systems 37 ( NeurIPS 2024 ), 2024 : 107984-108011.
[ 17 ] Rahima K , Hussain M.YOLOv11 : an overview of the key architectural enhancements [ PP / OL ] .V1.arXiv ( 2024-10-23 )[ 2026-01-03 ] .https : ∥doi.org / 10.48550 / arXiv.2410.17725.
[ 18 ] Tian Yunjie , Ye Qixiang , Doermann D.YOLOv12 : attention-centric real-time object detectors [ PP / OL ] . V1.arXiv ( 2025 02 18 )[ 2026 01 03 ] .https : ∥ doi.org / 10.48550 / arXiv.2502.12524.
[ 19 ] He Kaiming , Zhang Xiangyu , Ren Shaoqing , et al.Deep residual learning for image recognition [ C ] ∥2016 IEEE Conference on Computer Vision and Pattern Recognition ( CVPR ) .New York : IEEE , 2016 : 770-778.
[ 20 ] Liu Zhuang , Mao Hanzi , Wu Chaoyuan , et al.A convnet for the 2020s [ C ] ∥2022 IEEE Conference on Computer Vision and Pattern Recognition ( CVPR ) . New York : IEEE , 2022 : 11976-11986.

相似文献/References:

[1]卞越洋,高晓科,张伟军.机器人带电作业中的视觉定位与优化策略[J].机械与电子,2021,(05):57.
 BIAN Yueyang,GAO Xiaoke,ZHANG Weijun.Vision Positioning and Optimization Strategyin Robot Live Work[J].Machinery & Electronics,2021,(06):57.
[2]刘瑞昊,于振中,孙 强.改进多尺度特征融合的工业现场目标检测算法[J].机械与电子,2022,(11):40.
 LIU Ruihao,YU Zhenzhong,SUN Qiang.Improved Multi-scale Feature Fusion for Industrial Field Object Detection Algorithm[J].Machinery & Electronics,2022,(06):40.
[3]范 涛,王明泉,张俊生,等.基于轻量化 YOLOv4 的轮毂内部缺陷检测算法[J].机械与电子,2023,41(02):3.
 FAN Tao,WANG Mingquan,ZHANG Junsheng,et al.Internal Defect Detection Algorithm of Wheel Hub Based on Lightweight YOLOv4[J].Machinery & Electronics,2023,41(06):3.
[4]张晋钊,刘新妹,等.改进 YOLOv5s 的 PCB 元器件检测技术[J].机械与电子,2024,42(10):35.
 ZHANG Jinzhao,LIU Xinmei,et al.Improved YOLOv5s Component Detection Technology for PCB[J].Machinery & Electronics,2024,42(06):35.
[5]黄心玥,王明泉,耿宇杰,等.基于 YOLOv8 的铝合金铸造轮毂缺陷检测算法研究与分析[J].机械与电子,2025,(02):26.
 HUANG Xinyue,WANG Mingquan,GENG Yujie,et al.Research and Analysis of Defect Detection Algorithm for Aluminum Alloy Casting Wheels Based on YOLOv8[J].Machinery & Electronics,2025,(06):26.
[6]郑卓纹,吴攀超,王婷婷,等.基于轻量化目标检测算法的指针仪表读数识别[J].机械与电子,2025,(05):10.
 ZHENG Zhuowen,WU Panchao,WANG Tingting,et al.Pointer Gauge Reading Recognition Based on Lightweight Target Detection Algorithm[J].Machinery & Electronics,2025,(06):10.
[7]郭 庆,陈 川,季强东,等.基于改进 YOLOv11 的复杂环境下火焰目标检测[J].机械与电子,2025,(11):33.
 GUO Qing,CHEN Chuan,JI Qiangdong,et al.Flame Target Detection in Complex Environments Based on Improved YOLOv11[J].Machinery & Electronics,2025,(06):33.
[8]兰海麟,吉琳娜,杨风暴,等.基于生成对抗网络的红外与可见光视频语义特征驱动融合方法[J].机械与电子,2026,44(01):35.
 LAN Hailin,JI Linna,YANG Fengbao,et al.Semantic Feature Driven Fusion Method for Infrared and Visible Videos Based on Generative Adversarial Networks[J].Machinery & Electronics,2026,44(06):35.
[9]张思思,滑文强. 基于时频双域协同与语义增强的复杂水域漂浮物检测方法[J].机械与电子,2026,44(03):47.
 ZHANG Sisi,HUA Wenqiang. A Dual-domain Synergistic and Semantically Enhanced Method for Floating Object Detection in Complex Water Environments[J].Machinery & Electronics,2026,44(06):47.

备注/Memo

备注/Memo:
收稿日期: 2026-03-23
基金项目:西安市 2024 年度社会科学规划基金项目( 24YZ17 );国家图书馆革命文献与民国时期文献保护计划基金项目( 20210068 )
作者简介:阮 楠 ( 1987- ),女,陕西韩城人,硕士,馆员,研究方向为图书智能管理及智慧图书馆建设;张冬旭 ( 2001- ),男,河南南阳人,硕士研究生,研究方向为电子信息;王普雄 ( 1984- ),男,陕西西安人,硕士,馆员,研究方向为智慧图书馆建设。
更新日期/Last Update: 2026-08-26