<?xml version="1.1" encoding="utf-8"?>
<article xsi:noNamespaceSchemaLocation="http://jats.nlm.nih.gov/publishing/1.1/xsd/JATS-journalpublishing1-mathml3.xsd" dtd-version="1.1" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"><front><journal-meta><journal-id journal-id-type="publisher-id">EIR</journal-id><journal-title-group><journal-title>Educational Innovation Research</journal-title></journal-title-group><issn>3029-1844</issn><eissn>3029-1852</eissn><publisher><publisher-name>Bio-Byword Scientific Publishing Pty. Ltd.</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.18063/EIR.v4i4.1995</article-id><article-categories><subj-group subj-group-type="heading"><subject>Article</subject></subj-group></article-categories><title>Research on an Automated Intraday Liquidity Scheduling Strategy for Finance Companies Based on Deep Reinforcement Learning</title><url>https://artdesignp.com/journal/EIR/4/4/10.18063/EIR.v4i4.1995</url><author>GeBin</author><pub-date pub-type="publication-year"><year>2026</year></pub-date><volume>4</volume><issue>4</issue><history><date date-type="pub"><published-time>2026-04-26</published-time></date></history><abstract>This study rigorously formulates the complex fund-scheduling problem as a Markov decision process (MDP). It constructs a state space that integrates real-time and forecast information, an atomic action space that conforms to business logic, and a reward function that balances long-term returns against immediate risk. To address the curse of dimensionality and the credit-assignment problem in coordinated scheduling among multiple fund units, a multi-agent deep deterministic policy gradient (MADDPG) algorithm is adopted. Under a centralized-training and decentralized-execution framework, the algorithm reconciles global optimization with decentralized decision-making. In addition, a difference-reward mechanism and Kalman filtering are used to accurately measure each agent&amp;rsquo;s individual contribution and reduce the impact of environmental noise on reward signals. The results show that, compared with a static rule engine and a conventional linear programming method, the proposed deep reinforcement learning strategy reduces average daily funding costs by 50.4%, lowers the payment failure rate to 0.002%, and maintains a high liquidity buffer adequacy ratio. The strategy also demonstrates clear advantages in decision timeliness, collaborative handling of complex instructions, and self-adaptation potential, thereby providing an innovative pathway for finance-company fund scheduling to progress from intelligentization to automation.</abstract><keywords>fund scheduling, Markov decision process, reward function, Kalman filtering, finance company</keywords></article-meta></front><body/><back><ref-list><ref id="B1" content-type="article"><label>1</label><element-citation publication-type="journal"><p>[1] Wang YZ, Song X B, Sun Y, et al., 2025, A Survey of Capital Efficiency and Financial Risk among Chinese Listed Companies: 2024. Accounting Research, 2025(12): 180-188.
[2] Jiang AL, 2025, Practice of Building a World-Class Financial Management System in Company A, a Financial Leasing Company. Finance &amp;amp; Accounting, 2025(15): 27-29.
[3] Zhao M, Xie L, Lin WJ, et al., 2024, A Deep Reinforcement Learning Portfolio Model Based on a Dynamic Selection Predictor. Computer Science, 51(4): 344-352.
[4] Long J, Xie L, Xu HJ, 2024, An Ensemble Deep Reinforcement Learning Portfolio Model. Journal of Computer Applications, 44(1): 300-310.
[5] Bu Z, Zhang SF, Li XY, et al., 2023, Adaptive Stock Index Forecasting Based on Deep Reinforcement Learning. Journal of Management Sciences in China, 26(4): 148-174.</p><pub-id pub-id-type="doi"/></element-citation></ref></ref-list></back></article>
