<?xml version="1.1" encoding="utf-8"?>
<article xsi:noNamespaceSchemaLocation="http://jats.nlm.nih.gov/publishing/1.1/xsd/JATS-journalpublishing1-mathml3.xsd" dtd-version="1.1" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"><front><journal-meta><journal-id journal-id-type="publisher-id">JWA</journal-id><journal-title-group><journal-title>Journal of World Architecture</journal-title></journal-title-group><issn>2208-3480</issn><eissn>2208-3499</eissn><publisher><publisher-name>Bio-Byword Scientific Publishing Pty. Ltd.</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.26689/jwa.v9i6.13409</article-id><article-categories><subj-group subj-group-type="heading"><subject>Article</subject></subj-group></article-categories><title>Optimal Control Strategy for Unit Operation Based on Reinforcement Learning</title><url>https://artdesignp.com/journal/JWA/9/6/10.26689/jwa.v9i6.13409</url><author>LuoGuangming,ShiTaotao,ZhangChao,JiangJunying,LiHui</author><pub-date pub-type="publication-year"><year>2025</year></pub-date><volume>9</volume><issue>6</issue><history><date date-type="pub"><published-time>2025-12-31</published-time></date></history><abstract>Power system operation optimization faces dual challenges from energy structure transformation and extreme environmental conditions. Traditional unit control methods demonstrate limitations in addressing renewable energy volatility, load demand uncertainty, and sudden system disturbances. Deep reinforcement learning, through constructing a state-action-reward decision framework, effectively handles the time-varying, nonlinear, and uncertain characteristics of complex systems, providing new technical pathways for unit operation optimization. Studies show that applications of voltage regulation frameworks based on gated Markov decision processes and reinforcement learning in optimizing high-pressure feedwater heater operations, along with the integration of Hooke-Jeeves algorithm and deep deterministic strategy gradient methods in air handling unit control, all validate deep reinforcement learning’s unique advantages in solving multi-objective optimization problems for power generation units.</abstract><keywords/></article-meta></front><body/><back><ref-list><ref id="B1" content-type="article"><label>1</label><element-citation publication-type="journal"><p>Zou Y, Ji Y, Li W, et al., 2024, Research on Deep Learning-Based Optimization Modeling Method for Combined Cycle Units. Power Equipment Management, 2024(18): 280–282.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B2" content-type="article"><label>2</label><element-citation publication-type="journal"><p>Zhou N, Liang X, Yu X, et al., 2023, Research on Optimal Operation of Integrated Energy Systems Using DRL. Power Big Data, 26(6): 49–57.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B3" content-type="article"><label>3</label><element-citation publication-type="journal"><p>Li J, Yao Y, Liu B, 2022, Optimal Operation of Integrated Energy Systems Based on Comprehensive Evaluation Indicators. Journal of Guangxi University (Natural Science Edition), 47(6): 1518–1531.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B4" content-type="article"><label>4</label><element-citation publication-type="journal"><p>Liu G, Jin Y, Cao X, et al., 2022, Thermal Power Load Optimization Allocation for Gas Turbine Units Based on Deep Learning and Chaos Optimization. Journal of Thermal Power Generation, 51(2): 178–182.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B5" content-type="article"><label>5</label><element-citation publication-type="journal"><p>Nie C, An L, Xu G, et al., 2021, Real-Time Optimization Strategy for Air-Cooled Island Operation in Coal-Fired Power Stations based on Big Data. Journal of Power Engineering, 41(9): 713–720.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B6" content-type="article"><label>6</label><element-citation publication-type="journal"><p>Lü J, 2024, Machine Learning-Based Optimization of Energy Efficiency Parameters for Coal-Fired Power Units, thesis, Northeast Electric Power University.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B7" content-type="article"><label>7</label><element-citation publication-type="journal"><p>Pan L, 2017, Research on Fault Diagnosis of Key Components in Wind Turbine Drive Systems Using Deep Learning Networks, thesis, Shanghai Dianji University.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B8" content-type="article"><label>8</label><element-citation publication-type="journal"><p>Zhang L, Wu H, Li Z, et al., 2024, Design of an AI-based Maintenance System for Thermal Power Plant Units. Mold Manufacturing, 24(11): 207–209.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B9" content-type="article"><label>9</label><element-citation publication-type="journal"><p>Zhang Y, Wang L, Liu Y, et al., 2024, A Multi-Turbine Operation Monitoring Method Based on Balanced Distribution Adaptive Transfer Learning. Renewable Energy, 42(8): 1068–1073.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B10" content-type="article"><label>10</label><element-citation publication-type="journal"><p>Tang H, Yan Z, Fang D, et al., n.d., Deep Transfer Reinforcement Learning-Based Optimization Method for Flexible Resource Grid Dispatching. Control Engineering, 1–13.</p><pub-id pub-id-type="doi"/></element-citation></ref></ref-list></back></article>
