<?xml version="1.1" encoding="utf-8"?>
<article xsi:noNamespaceSchemaLocation="http://jats.nlm.nih.gov/publishing/1.1/xsd/JATS-journalpublishing1-mathml3.xsd" dtd-version="1.1" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"><front><journal-meta><journal-id journal-id-type="publisher-id">JERA</journal-id><journal-title-group><journal-title>Journal of Electronic Research and Application</journal-title></journal-title-group><issn>2208-3502</issn><eissn>2208-3510</eissn><publisher><publisher-name>Bio-Byword Scientific Publishing Pty. Ltd.</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.26689/jera.v9i6.13152</article-id><article-categories><subj-group subj-group-type="heading"><subject>Article</subject></subj-group></article-categories><title>Research on Optimization Algorithm for Video Recognition Accuracy in Smart Construction Sites</title><url>https://artdesignp.com/journal/JERA/9/6/10.26689/jera.v9i6.13152</url><author>GaoXiang,NieLei,WangYinan,ZhangChi</author><pub-date pub-type="publication-year"><year>2025</year></pub-date><volume>9</volume><issue>6</issue><history><date date-type="pub"><published-time>2025-12-16</published-time></date></history><abstract>Aiming at the problems faced by construction site video management in the recognition of cigarette butts, reflective vests, and other objects, such as small target confusion, high-brightness false alarms, occlusion missed detections, and poor adaptability to complex environments, this study proposes a recognition accuracy optimization algorithm based on multimodal fusion. The research constructs a dataset containing three modalities of data: visible light, infrared, and millimeter-wave. The Dust-GAN algorithm is adopted to realize dust removal and enhancement of dusty images, and the SAA module is introduced into YOLOv8-s to improve the small target recall rate. Meanwhile, three-modal feature fusion is achieved, and channel pruning and quantization-aware training are used to realize algorithm lightweighting. The algorithm was deployed and operated on-site for 3 months, effectively reducing the construction site safety accident rate by 65%, which provides a solution for safety management and control in smart construction sites under complex environments.</abstract><keywords/></article-meta></front><body/><back><ref-list><ref id="B1" content-type="article"><label>1</label><element-citation publication-type="journal"><p>Jiang X, Wang B, Xia Y, et al., 2022, Smoking Behavior Detection Based on Human Key Points and YOLOv4. Journal of Shaanxi Normal University (Natural Science Edition), 50(3): 96–103.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B2" content-type="article"><label>2</label><element-citation publication-type="journal"><p>Wang D, Bai C, Wu K, 2021, Review of Video Object Detection Based on Deep Learning. Journal of Computer Science and Exploration, 2021: 1–15.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B3" content-type="article"><label>3</label><element-citation publication-type="journal"><p>Varghese R, Sambath M, 2024, YOLOv8: A Novel Object Detection Algorithm with Enhanced Performance and Robustness. International Conference on Advances in Data Engineering and Intelligent Computing Systems (ADICS), 2024.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B4" content-type="article"><label>4</label><element-citation publication-type="journal"><p>Chen S, Ma H, Wang T, et al., 2022, Video Sentiment Analysis Technology Based on Multimodal Fusion. Journal of Chengdu University of Information Technology, 2022(6): 656–661.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B5" content-type="article"><label>5</label><element-citation publication-type="journal"><p>Guo N, Jiang L, 2021, Processing of Multimodal Video Captions Based on Hard Attention Mechanism. Application Research of Computers, 38(3): 956–960.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B6" content-type="article"><label>6</label><element-citation publication-type="journal"><p>Pan W, Wei C, Qian C, et al., 2024, Improved YOLOv8s Model for Small Object Detection from UAV Perspective. Computer Engineering and Applications, 60(9): 142–150.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B7" content-type="article"><label>7</label><element-citation publication-type="journal"><p>Wang Y, Li M, Sun H, 2024, External Knowledge-Based VQA Integrating Cross-Modal Transformer. Science Technology and Engineering, 24(20): 8577–8586.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B8" content-type="article"><label>8</label><element-citation publication-type="journal"><p>He Y, Zhang X, Sun J, 2017, Channel Pruning for Accelerating Very Deep Neural Networks. Proceedings of the IEEE International Conference on Computer Vision, 2017: 1389–1397.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B9" content-type="article"><label>9</label><element-citation publication-type="journal"><p>Li H, Wang L, Zhang J, 2023, Research on Multimodal Data Acquisition and Synchronization System for Smart Construction Sites. Automation &amp; Instrumentation, 2023(8): 145–149.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B10" content-type="article"><label>10</label><element-citation publication-type="journal"><p>Zhao Y, Wang T, Li T, 2021, Image Rendering and Data Augmentation Technology for Reflective Vests in Complex Lighting Environments. Journal of Graphics, 42(5): 825–832.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B11" content-type="article"><label>11</label><element-citation publication-type="journal"><p>Chen J, 2019, Design and Research of Dust Concentration Detection Based on Image Method, thesis, China Jiliang University.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B12" content-type="article"><label>12</label><element-citation publication-type="journal"><p>Zhang L, Tian Y, 2024, Multi-Scale Lightweight Vehicle Object Detection Algorithm Based on Improved YOLOv8. Computer Engineering and Applications, 60(3): 129–137.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B13" content-type="article"><label>13</label><element-citation publication-type="journal"><p>Yue M, Shu K, Zhang C, et al., 2024, Research on Infrared Small Target Detection Algorithm Based on Improved YOLOv8. Infrared Technology, 2024(11): 1286–1292.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B14" content-type="article"><label>14</label><element-citation publication-type="journal"><p>Ju R, Chien C, Chiang J, 2024, YOLOv8-ResCBAM: YOLOv8 Based on an Effective Attention Module for Pediatric Wrist Fracture Detection. arXiv. https://doi.org/10.48550/arXiv.2409.18826</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B15" content-type="article"><label>15</label><element-citation publication-type="journal"><p>Wang M, Yao G, Yang Y, et al., 2023, Deep Learning-Based Object Detection for Visible Dust and Prevention Measures on Construction Sites. Developments in the Built Environment, 2023: 16.</p><pub-id pub-id-type="doi"/></element-citation></ref></ref-list></back></article>
