<?xml version="1.1" encoding="utf-8"?>
<article xsi:noNamespaceSchemaLocation="http://jats.nlm.nih.gov/publishing/1.1/xsd/JATS-journalpublishing1-mathml3.xsd" dtd-version="1.1" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"><front><journal-meta><journal-id journal-id-type="publisher-id">JERA</journal-id><journal-title-group><journal-title>Journal of Electronic Research and Application</journal-title></journal-title-group><issn>2208-3502</issn><eissn>2208-3510</eissn><publisher><publisher-name>Bio-Byword Scientific Publishing Pty. Ltd.</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.26689/jera.v9i3.10597</article-id><article-categories><subj-group subj-group-type="heading"><subject>Article</subject></subj-group></article-categories><title>The Latest Research Progress of Attention Mechanism in Deep Learning</title><url>https://artdesignp.com/journal/JERA/9/3/10.26689/jera.v9i3.10597</url><author>JiangXu,BaiXiaoling,YinLifeng</author><pub-date pub-type="publication-year"><year>2025</year></pub-date><volume>9</volume><issue>3</issue><history><date date-type="pub"><published-time>2025-05-29</published-time></date></history><abstract>With the development of artificial intelligence and deep learning, the attention mechanism has become a key technology for enhancing the performance of complex tasks. This paper reviews the evolution of attention mechanisms, including soft attention, hard attention, and recent innovations such as multi-head latent attention and cross-attention. It focuses on the latest research outcomes, such as lightning attention, the PADRe polynomial attention replacement algorithm, the context anchor attention module, and improvements in attention mechanisms for large models. These advancements improve the efficiency and accuracy of models, expanding the application potential of attention mechanisms in fields such as computer vision, natural language processing, and remote sensing object detection, aiming to provide readers with a comprehensive understanding and stimulate innovative thinking.</abstract><keywords/></article-meta></front><body/><back><ref-list><ref id="B1" content-type="article"><label>1</label><element-citation publication-type="journal"><p>Vaswani A, Shazeer N, Parmar N, et al., 2017, Attention is All You Need. Advances in Neural Information Processing Systems, 30: 5998–6008.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B2" content-type="article"><label>2</label><element-citation publication-type="journal"><p>Hu J, Shen L, Sun G, 2018, Squeeze and Excitation Networks, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 7132–7141.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B3" content-type="article"><label>3</label><element-citation publication-type="journal"><p>Devlin J, Chang MW, Lee K, et al., 2019, Bert: Pre-training of Deep Bidirectional Transformers for Language Understanding, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (long and short papers), Minneapolis, Minnesota, 4171–4186.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B4" content-type="article"><label>4</label><element-citation publication-type="journal"><p>Wang X, Girshick R, Gupta A, et al., 2018, Non-local Neural Networks, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 7794–7803.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B5" content-type="article"><label>5</label><element-citation publication-type="journal"><p>Anderson P, He X, Buehler C, et al., 2018, Bottom-up and top-down Attention for Image Captioning and Visual Question Answering, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6077–6086.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B6" content-type="article"><label>6</label><element-citation publication-type="journal"><p>Bahdanau D, Cho K, Bengio Y, 2014, Neural Machine Translation by Jointly Learning to Align and Translate. https://doi.org/10.48550/arXiv.1409.0473</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B7" content-type="article"><label>7</label><element-citation publication-type="journal"><p>Luong MT, Pham H, Manning CD, 2015, Effective Approaches to Attention-based Neural Machine Translation. https://doi.org/10.48550/arXiv.1508.04025</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B8" content-type="article"><label>8</label><element-citation publication-type="journal"><p>Sukhbaatar S, Weston J, Fergus R, 2015, End-to-end Memory Networks. Advances in Neural Information Processing Systems, 28.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B9" content-type="article"><label>9</label><element-citation publication-type="journal"><p>Yao L, Torabi A, Cho K, et al., 2015, Describing Videos by Exploiting Temporal Structure, Proceedings of the IEEE International Conference on Computer Vision, 4507–4515.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B10" content-type="article"><label>10</label><element-citation publication-type="journal"><p>Martins A, Astudillo R, 2016, From Softmax to Sparsemax: A Sparse Model of Attention and Multi-label Classification, International Conference on Machine Learning. PMLR, 1614–1623.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B11" content-type="article"><label>11</label><element-citation publication-type="journal"><p>Yang Z, Yang D, Dyer C, et al., 2016, Hierarchical Attention Networks for Document Classification. Association for Computational Linguistics, 2016: 1480–1489.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B12" content-type="article"><label>12</label><element-citation publication-type="journal"><p>Lu J, Xiong C, Parikh D, et al., 2017, Knowing When to Look: Adaptive Attention via a Visual Sentinel for Image Captioning, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 375–383.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B13" content-type="article"><label>13</label><element-citation publication-type="journal"><p>Gheini M, Ren X, May J, 2021, Cross-attention is All You Need: Adapting Pretrained Transformers for Machine Translation. https://arxiv.org/abs/2104.08771</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B14" content-type="article"><label>14</label><element-citation publication-type="journal"><p>Qin Z, Sun W, Li D, et al., 2024, Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models. https://arxiv.org/abs/2401.04658</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B15" content-type="article"><label>15</label><element-citation publication-type="journal"><p>Liu A, Feng B, Wang B, et al., 2024, Deepseek-v2: A Strong, Economical, and Efficient Mixture-of-experts Language Model. https://arxiv.org/abs/2405.04434</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B16" content-type="article"><label>16</label><element-citation publication-type="journal"><p>Yuan J, Gao H, Dai D, et al., 2025, Native Sparse Attention: Hardware-aligned and Natively Trainable Sparse Attention. https://arxiv.org/abs/2502.11089</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B17" content-type="article"><label>17</label><element-citation publication-type="journal"><p>Cai X, Lai Q, Wang Y, et al., 2024, Poly Kernel Inception Network for Remote Sensing Detection, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 27706–27716.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B18" content-type="article"><label>18</label><element-citation publication-type="journal"><p>Letourneau PD, Singh MK, Cheng HP, et al., 2024, Padre: A Unifying Polynomial Attention Drop-in Replacement for Efficient Vision Transformer. https://arxiv.org/abs/2407.11306</p><pub-id pub-id-type="doi"/></element-citation></ref></ref-list></back></article>
