<?xml version="1.1" encoding="utf-8"?>
<article xsi:noNamespaceSchemaLocation="http://jats.nlm.nih.gov/publishing/1.1/xsd/JATS-journalpublishing1-mathml3.xsd" dtd-version="1.1" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"><front><journal-meta><journal-id journal-id-type="publisher-id">JERA</journal-id><journal-title-group><journal-title>Journal of Electronic Research and Application</journal-title></journal-title-group><issn>2208-3502</issn><eissn>2208-3510</eissn><publisher><publisher-name>Bio-Byword Scientific Publishing Pty. Ltd.</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.26689/jera.v10i5.15270</article-id><article-categories><subj-group subj-group-type="heading"><subject>Article</subject></subj-group></article-categories><title>PSR-DETR: An Improved RT-DETR for Rail Surface Defect Detection</title><url>https://artdesignp.com/journal/JERA/10/5/10.26689/jera.v10i5.15270</url><author>LiGuangzhuo,JiaShijie</author><pub-date pub-type="publication-year"><year>2026</year></pub-date><volume>10</volume><issue>5</issue><history><date date-type="pub"><published-time>2026-06-29</published-time></date></history><abstract>This paper addresses low detection accuracy in rail surface defect detection. The problem comes from many defect types, large scale changes, and small dense targets. Hence, an improved model based on RT-DETR is proposed namely PSR-DETR. The PR_BasicBlock module first simplifies the model structure. It reduces parameters and computation cost. Meanwhile, it maintains satisfactory detection performance. Consequently, the network becomes more lightweight. After that, the RetC3 module adds a new attention mechanism. It enhances feature integration. It also strengthens the model’s capability to represent and distinguish targets of different scales. Finally, the SSFF module adds extra feature fusion paths. It helps the model emphasize critical regions. As a result, the detection performance is further improved. Experimental results show clear improvements, where the model does not greatly increase parameters or computation. The mAP@0.5 achieves 68.0%. The mAP@0.5:0.95 attains 44.7%, which are improvements of 6.3% and 2.7% over the original model. These findings show that the proposed method is effective and practical for enhancing detection performance.</abstract><keywords/></article-meta></front><body/><back><ref-list><ref id="B1" content-type="article"><label>1</label><element-citation publication-type="journal"><p>Yang F, Tu W, Wei Z, et al., 2023, Review on the Development Of Railway Civil, Electrical and Power Inspection Equipment. Journal of Traffic and Transportation Engineering, 23(1): 47–69.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B2" content-type="article"><label>2</label><element-citation publication-type="journal"><p>Gong W, Akbar M, Jawad G, et al., 2022, Nondestructive Testing Technologies for Rail Inspection: A Review. Coatings, 12(11): 1790.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B3" content-type="article"><label>3</label><element-citation publication-type="journal"><p>He Z, Wang Y, Liu J, et al., 2016, High-Speed Rail Surface Defect Image Segmentation Based On Background Subtraction. Chinese Journal of Scientific Instrument, 37(3): 640–649 .</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B4" content-type="article"><label>4</label><element-citation publication-type="journal"><p>Cao Y, Duan Y, Wu D, 2020, 2D-Otsu Rail Defect Image Segmentation based on WFSOA. Computer Science, 47(5): 154–160 .</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B5" content-type="article"><label>5</label><element-citation publication-type="journal"><p>Taştimur C, Karaköse M, Akın E, et al., 2016, Rail Defect Detection with Real-Time Image Processing Technique, 2016 IEEE 14th International Conference on Industrial Informatics (INDIN), 411–415.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B6" content-type="article"><label>6</label><element-citation publication-type="journal"><p>Gan J, Wang J, Yu H, et al., 2018, Online Rail Surface Inspection Utilizing Spatial Consistency and Continuity. IEEE Transactions on Systems, Man, and Cybernetics: Systems, 50(7): 2741–2751.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B7" content-type="article"><label>7</label><element-citation publication-type="journal"><p>Zhou M, Tang Q, Shi T, et al., 2023, Rail Surface Crack Detection Algorithm based on Improved Yolov5s. Liquid Crystals &amp; Displays, 38(5): 666–679.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B8" content-type="article"><label>8</label><element-citation publication-type="journal"><p>Yuan H, Chen H, Liu S, et al., 2019, A Deep Convolutional Neural Network For Detection Of Rail Surface Defect, 2019 IEEE Vehicle Power and Propulsion Conference (VPPC), 1–4.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B9" content-type="article"><label>9</label><element-citation publication-type="journal"><p>Feng J, Yuan H, Hu Y, et al., 2020, Research on Deep Learning Method For Rail Surface Defect Detection. IET Electrical Systems in Transportation, 10(4): 436–442.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B10" content-type="article"><label>10</label><element-citation publication-type="journal"><p>Zhang C, Xu D, Zhang L, et al., 2023, Rail Surface Defect Detection based on Image Enhancement and Improved YOLOX. Electronics, 12(12): 2672.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B11" content-type="article"><label>11</label><element-citation publication-type="journal"><p>Wu Y, Cui C, He Y, 2024, Rail Surface Defect Detection Method based on Semantic Augmentation and Yolov8. Journal of Railway Science and Engineering, 21(1): 1–12.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B12" content-type="article"><label>12</label><element-citation publication-type="journal"><p>Girshick R, Donahue J, Darrell T, et al., 2014, Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 580–587.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B13" content-type="article"><label>13</label><element-citation publication-type="journal"><p>Girshick R, 2015, Fast R-CNN, Proceedings of the IEEE International Conference on Computer Vision, 1440–1448.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B14" content-type="article"><label>14</label><element-citation publication-type="journal"><p>He K, Gkioxari G, Dollár P, et al., 2017, Mask R-CNN, Proceedings of the IEEE International Conference on Computer Vision, 2961–2969.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B15" content-type="article"><label>15</label><element-citation publication-type="journal"><p>Xu Y, Yu G, Wang Y, et al., 2017, Car Detection from Low-Altitude UAV Imagery with the Faster R-CNN. Journal of Advanced Transportation, 2017: 2823617.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B16" content-type="article"><label>16</label><element-citation publication-type="journal"><p>Avola D, Cinque L, Diko A, et al., 2021, MS-Faster R-CNN: Multi-Stream Backbone for Improved Faster R-CNN Object Detection and Aerial Tracking from UAV Images. Remote Sensing, 2021(13): 1670.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B17" content-type="article"><label>17</label><element-citation publication-type="journal"><p>Redmon J, Farhadi A, 2018, YOLOv3: An Incremental Improvement. arXiv, arXiv:1804.02767.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B18" content-type="article"><label>18</label><element-citation publication-type="journal"><p>Bochkovskiy A, Wang C, Liao H, 2020, Yolov4: Optimal Speed and Accuracy of Object Detection, arXiv, arXiv:2004.10934.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B19" content-type="article"><label>19</label><element-citation publication-type="journal"><p>Jocher G, Chaurasia A, Stoken A, et al., 2022, Ultralytics YOLOv5, Zenodo, viewed August 4, 2024, https://github.com/ultralytics/yolov5</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B20" content-type="article"><label>20</label><element-citation publication-type="journal"><p>Wang C, Bochkovskiy A, Liao H, 2023, Yolov7: Trainable Bag-of-Freebies Sets New State-of-The-Art for Real-Time Object Detectors, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 7464–7475.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B21" content-type="article"><label>21</label><element-citation publication-type="journal"><p>Jocher G, Chaurasia A, Qiu J, 2023, Ultralytics YOLOv8, viewed July 18, 2025, https://github.com/ultralytics/ultralytics</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B22" content-type="article"><label>22</label><element-citation publication-type="journal"><p>Carion N, Massa F, Synnaeve G, et al., 2020, End-to-End Object Detection with Transformers, Proceedings of the European Conference on Computer Vision, 213–229.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B23" content-type="article"><label>23</label><element-citation publication-type="journal"><p>Zhu X, Su W, Lu L, et al., 2020, Deformable DETR: Deformable Transformers for End-to-End Object Detection, arXiv, arXiv:2010.04159.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B24" content-type="article"><label>24</label><element-citation publication-type="journal"><p>Zhang H, Li F, Liu S, et al., 2022, DINO: DETR with Improved Denoising Anchor Boxes for End-to-End Object Detection, arXiv, arXiv:2203.03605.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B25" content-type="article"><label>25</label><element-citation publication-type="journal"><p>Zhao Y, Lv W, Xu S, et al., 2023, Detrs Beat Yolos on Real-Time Object Detection, arXiv, arXiv:2304.08069.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B26" content-type="article"><label>26</label><element-citation publication-type="journal"><p>Vaswani A, Shazeer N, Parmar N, et al., 2017, Attention is All you Need, Advances in Neural Information Processing Systems.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B27" content-type="article"><label>27</label><element-citation publication-type="journal"><p>Liu Z, Lin Y, Cao Y, et al., 2021, Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows, Proceedings of the IEEE/CVF International Conference on Computer Vision, 10012–10022.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B28" content-type="article"><label>28</label><element-citation publication-type="journal"><p>Chen J, Kao S, He H, et al., 2023, Run, Don’t Walk: Chasing Higher FLOPS for Faster Neural Networks, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 12021–12031.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B29" content-type="article"><label>29</label><element-citation publication-type="journal"><p>Ding X, Zhang X, Ma N, et al., 2021, Repvgg: Making VGG-Style Convnets Great Again, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 13733–13742.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B30" content-type="article"><label>30</label><element-citation publication-type="journal"><p>Kang M, Ting M, Ting F, et al., 2024, ASF-YOLO: A Novel YOLO Model with Attentional Scale Sequence Fusion for Cell Instance Segmentation. Image and Vision Computing, 2024(147): 105057.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B31" content-type="article"><label>31</label><element-citation publication-type="journal"><p>Fan Q, Huang H, Chen M, et al., 2024, RMT: Retentive Networks Meet Vision Transformers, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 5641–5651.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B32" content-type="article"><label>32</label><element-citation publication-type="journal"><p>Zhu L, Wang X, Ke Z, et al., 2023, BiFormer: Vision Transformer with Bi-Level Routing Attention, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10323–10333.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B33" content-type="article"><label>33</label><element-citation publication-type="journal"><p>Yang X, Li Y, Li Y, et al., 2025, Lightweight Rail Surface Defect Detection Algorithm based on an Improved Yolov8. Measurement, 2025(242): 115922.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B34" content-type="article"><label>34</label><element-citation publication-type="journal"><p>Baidu PaddlePaddle Team, 2022, PP-YOLOE Object Detection Model, viewed June 8, 2025, https://github.com/PaddlePaddle/PaddleDetection</p><pub-id pub-id-type="doi"/></element-citation></ref></ref-list></back></article>
