<?xml version="1.1" encoding="utf-8"?>
<article xsi:noNamespaceSchemaLocation="http://jats.nlm.nih.gov/publishing/1.1/xsd/JATS-journalpublishing1-mathml3.xsd" dtd-version="1.1" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"><front><journal-meta><journal-id journal-id-type="publisher-id">JERA</journal-id><journal-title-group><journal-title>Journal of Electronic Research and Application</journal-title></journal-title-group><issn>2208-3502</issn><eissn>2208-3510</eissn><publisher><publisher-name>Bio-Byword Scientific Publishing Pty. Ltd.</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.26689/jera.v9i4.11459</article-id><article-categories><subj-group subj-group-type="heading"><subject>Article</subject></subj-group></article-categories><title>Real-Time Sound Source Localization Method Based on Selective SRP-PHAT and Vision Fusion</title><url>https://artdesignp.com/journal/JERA/9/4/10.26689/jera.v9i4.11459</url><author>HuangJinde</author><pub-date pub-type="publication-year"><year>2025</year></pub-date><volume>9</volume><issue>4</issue><history><date date-type="pub"><published-time>2025-08-07</published-time></date></history><abstract>Aiming at the problem that the traditional SRP-PHAT sound source localization method performs intensive search in a 360-degree space, resulting in high computational complexity and difficulty in meeting real-time requirements, an innovative high-precision sound source localization method is proposed. This method combines the selective SRP-PHAT algorithm with real-time visual analysis. Its core innovations include using face detection to dynamically determine the scanning angle range to achieve visually guided selective scanning, distinguishing face sound sources from background noise through a sound source classification mechanism, and implementing intelligent background orientation selection to ensure comprehensive monitoring of environmental noise. Experimental results show that the method achieves a positioning accuracy of ±5 degrees and a processing speed of more than 10FPS in complex real environments, and its performance is significantly better than the traditional full-angle scanning method.</abstract><keywords/></article-meta></front><body/><back><ref-list><ref id="B1" content-type="article"><label>1</label><element-citation publication-type="journal"><p>Schmidt RO, 1986, Multiple Emitter Location and Signal Parameter Estimation. IEEE Transactions on Antennas and Propagation, 34(3).</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B2" content-type="article"><label>2</label><element-citation publication-type="journal"><p>Omologo M, Svaizer P, 1997, Use of the Crosspower-Spectrum Phase in Acoustic Event Location. IEEE Trans Speech Audio Process, 5(3): 288–292.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B3" content-type="article"><label>3</label><element-citation publication-type="journal"><p>Brumann K, Doclo S, 2024, Steered Response Power-Based Direction-of-Arrival Estimation Exploiting an Auxiliary Microphone. European Signal Processing Conference, 917–921.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B4" content-type="article"><label>4</label><element-citation publication-type="journal"><p>Li C, Hendriks RC, 2023, Alternating Least-Squares-Based Microphone Array Parameter Estimation for a Single-Source Reverberant and Noisy Acoustic Scenario, in IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31, 3922–3934.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B5" content-type="article"><label>5</label><element-citation publication-type="journal"><p>Diaz-Guerra D, Miguel A, JR Beltran JR, 2021, Robust Sound Source Tracking Using SRP-PHAT and 3D Convolutional Neural Networks, in IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29, 300–311.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B6" content-type="article"><label>6</label><element-citation publication-type="journal"><p>Alghareb FS, Hasan BT, 2025, Multitask Learning-Based Pipeline-Parallel Computation Offloading Architecture for Deep Face Analysis. Computers, 14: 29.</p><pub-id pub-id-type="doi"/></element-citation></ref></ref-list></back></article>
