<?xml version="1.1" encoding="utf-8"?>
<article xsi:noNamespaceSchemaLocation="http://jats.nlm.nih.gov/publishing/1.1/xsd/JATS-journalpublishing1-mathml3.xsd" dtd-version="1.1" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"><front><journal-meta><journal-id journal-id-type="publisher-id">JERA</journal-id><journal-title-group><journal-title>Journal of Electronic Research and Application</journal-title></journal-title-group><issn>2208-3502</issn><eissn>2208-3510</eissn><publisher><publisher-name>Bio-Byword Scientific Publishing Pty. Ltd.</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.26689/jera.v9i1.9457</article-id><article-categories><subj-group subj-group-type="heading"><subject>Article</subject></subj-group></article-categories><title>A Multi-Scale Attention-Based Pedestrian Detection Method for Roadways Using the YOLOv5 Framework</title><url>https://artdesignp.com/journal/JERA/9/1/10.26689/jera.v9i1.9457</url><author>WangRuihan,LiuBoling,LiaoTingyu</author><pub-date pub-type="publication-year"><year>2025</year></pub-date><volume>9</volume><issue>1</issue><history><date date-type="pub"><published-time>2025-02-13</published-time></date></history><abstract>Due to multi-scale variations and occlusion problems, accurate traffic road pedestrian detection faces great challenges. This paper proposes an improved pedestrian detection method called Multi Scales Attention-YOLOv5x (MSA-YOLOv5x) based on the YOLOv5x framework. Firstly, by replacing the first convolutional operation of the backbone network with the Focus module, this method expands the number of image input channels to enhance feature expressiveness. Secondly, we construct C3_CBAM module instead of the original C3 module for better feature fusion. In this way, the learning process could achieve more multi-scale features and occluded pedestrian target features through channel attention and spatial attention. Additionally, a new feature pyramid detection layer and a new detection channel are embedded in the feature fusion part for enhancing multi-scale pedestrian detection accuracy. Compared with the baseline methods, experimental results on a public dataset demonstrate that the proposed method achieves optimal detection accuracy for traffic road pedestrian detection.</abstract><keywords/></article-meta></front><body/><back><ref-list><ref id="B1" content-type="article"><label>1</label><element-citation publication-type="journal"><p>Dalal N, Triggs B, 2005, Histograms of Oriented Gradients for Human Detection. 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), Vol. 1. IEEE, 2005: 886–893.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B2" content-type="article"><label>2</label><element-citation publication-type="journal"><p>Ahonen T, Hadid A, Pietik¨ainen M, 2004, Face Recognition with Local Binary Patterns. Computer Vision-ECCV 2004: 8th European Conference on Computer Vision, Proceedings, Part I 8, Springer, 2004: 469–481.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B3" content-type="article"><label>3</label><element-citation publication-type="journal"><p>Lowe DG, 2004, Distinctive Image Features from Scale-Invariant Keypoints. International Journal of Computer Vision, 60: 91–110.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B4" content-type="article"><label>4</label><element-citation publication-type="journal"><p>Girshick R, Donahue J, Darrell T, et al., 2014, Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014: 580–587.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B5" content-type="article"><label>5</label><element-citation publication-type="journal"><p>Girshick R, 2015, Fast R-CNN. Proceedings of the IEEE International Conference on Computer Vision, 2015: 1440–1448.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B6" content-type="article"><label>6</label><element-citation publication-type="journal"><p>Ren S, He K, Girshick R, et al., 2015, Faster R-CNN: Towards Realtime Object Detection with Region Proposal Networks. Advances in Neural Information Processing Systems, 28: 2015.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B7" content-type="article"><label>7</label><element-citation publication-type="journal"><p>He K, Gkioxari G, Doll´ar P, et al., 2017, Mask R-CNN. Proceedings of the IEEE International Conference on Computer Vision, 2017: 2961–2969.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B8" content-type="article"><label>8</label><element-citation publication-type="journal"><p>Redmon J, Divvala S, Girshick R, et al., 2016, You Only Look Once: Unified, Real-Time Object Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016: 779–788.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B9" content-type="article"><label>9</label><element-citation publication-type="journal"><p>Liu W, Anguelov D, Erhan D, et al., 2016, SSD: Single Shot Multibox Detector. Computer Vision–ECCV 2016: 14th European Conference Proceedings, Part I 14. Springer, 2016: 21–37.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B10" content-type="article"><label>10</label><element-citation publication-type="journal"><p>Tian Y, Luo P, Wang X, et al., 2015, Deep Learning Strong Parts for Pedestrian Detection. Proceedings of the IEEE International Conference on Computer Vision, 2015: 1904–1912.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B11" content-type="article"><label>11</label><element-citation publication-type="journal"><p>Li Q, Su Y, Gao Y, et al., 2022, Oaf-Net: An Occlusion-Aware Anchor-Free Network for Pedestrian Detection in a Crowd. IEEE Transactions on Intelligent Transportation Systems, 23(11): 21291–21300.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B12" content-type="article"><label>12</label><element-citation publication-type="journal"><p>Fei C, Liu B, Chen Z, et al., 2019, Learning Pixel-Level and Instance-Level Context-Aware Features for Pedestrian Detection in Crowds. IEEE Access, 7: 94944–94953.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B13" content-type="article"><label>13</label><element-citation publication-type="journal"><p>Xie J, Pang Y, Khan MH, et al., 2020, Mask-Guided Attention Network and Occlusion-Sensitive Hard Example Mining for Occluded Pedestrian Detection. IEEE Transactions on Image Processing, 30: 3872–3884.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B14" content-type="article"><label>14</label><element-citation publication-type="journal"><p>Xie H, Chen Y, Shin H, 2019, Context-Aware Pedestrian Detection Especially for Small-Sized Instances with Deconvolution Integrated Faster RCNN (dif R-CNN). Applied Intelligence, 49: 1200–1211.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B15" content-type="article"><label>15</label><element-citation publication-type="journal"><p>Lin C, Lu J, Wang G, et al., 2018, Graininess-Aware Deep Feature Learning for Pedestrian Detection. Proceedings of the European conference on Computer Vision (ECCV), 2018: 732–747.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B16" content-type="article"><label>16</label><element-citation publication-type="journal"><p>Yan C, Zhang H, Li X, et al., 2022, R-SSD: Refined Single Shot Multibox Detector for Pedestrian Detection. Applied Intelligence, 52(9): 10430–10447.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B17" content-type="article"><label>17</label><element-citation publication-type="journal"><p>Wang CY, Liao HYM, Wu YH, et al., 2020, CSPNet: A New Backbone that can Enhance Learning Capability of CNN. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020: 390–391.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B18" content-type="article"><label>18</label><element-citation publication-type="journal"><p>Felzenszwalb PF, Girshick RB, McAllester D, et al., 2009, Object Detection with Discriminatively Trained Part-Based Models. IEEE Transactions on Pattern Analysis and Machine Intelligence, 32(9): 1627–1645.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B19" content-type="article"><label>19</label><element-citation publication-type="journal"><p>Woo S, Park J, Lee JY, et al., 2018, CBAM: Convolutional Block Attention Module. Proceedings of the European Conference on Computer Vision (ECCV), 2018: 3–19.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B20" content-type="article"><label>20</label><element-citation publication-type="journal"><p>Redmon J, Farhadi A, 2018, Yolov3: An Incremental Improvement. arXiv Preprint, arXiv:1804.02767.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B21" content-type="article"><label>21</label><element-citation publication-type="journal"><p>He K, Zhang X, Ren S, et al., 2016, Deep Residual Learning for Image Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016: 770–778.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B22" content-type="article"><label>22</label><element-citation publication-type="journal"><p>He K, Zhang X, Ren S, et al., 2015, Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 37(9): 1904–1916.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B23" content-type="article"><label>23</label><element-citation publication-type="journal"><p>Zhang S, Benenson R, Schiele B, 2017, Citypersons: A Diverse Dataset for Pedestrian Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017: 3213–3221.</p><pub-id pub-id-type="doi"/></element-citation></ref></ref-list></back></article>
