Zum Hauptinhalt springen

Showing 1–26 of 26 results for author: Naylor, P A

.
  1. arXiv:2407.06342  [pdf, other

    eess.AS

    XANE Background Acoustic Embeddings: Ablation and Clustering Analysis

    Authors: Dushyant Sharma, James Fosburgh, Sri Harsha Dumpala, Chandramouli Shama Sastri, Stanislav Yu. Kruchinin, Patrick A. Naylor

    Abstract: We explore the recently proposed explainable acoustic neural embedding~(XANE) system that models the background acoustics of a speech signal in a non-intrusive manner. The XANE embeddings are used to estimate specific parameters related to the background acoustic properties of the signal which allows the embeddings to be explainable in terms of those parameters. We perform ablation studies on the… ▽ More

    Submitted 8 July, 2024; originally announced July 2024.

    Comments: arXiv admin note: substantial text overlap with arXiv:2406.05199

  2. arXiv:2406.05199  [pdf, other

    eess.AS cs.SD

    XANE: eXplainable Acoustic Neural Embeddings

    Authors: Sri Harsha Dumpala, Dushyant Sharma, Chandramouli Shama Sastri, Stanislav Kruchinin, James Fosburgh, Patrick A. Naylor

    Abstract: We present a novel method for extracting neural embeddings that model the background acoustics of a speech signal. The extracted embeddings are used to estimate specific parameters related to the background acoustic properties of the signal in a non-intrusive manner, which allows the embeddings to be explainable in terms of those parameters. We illustrate the value of these embeddings by performin… ▽ More

    Submitted 7 June, 2024; originally announced June 2024.

  3. arXiv:2405.02991  [pdf, other

    cs.SD eess.AS

    Steered Response Power for Sound Source Localization: A Tutorial Review

    Authors: Eric Grinstein, Elisa Tengan, Bilgesu Çakmak, Thomas Dietzen, Leonardo Nunes, Toon van Waterschoot, Mike Brookes, Patrick A. Naylor

    Abstract: In the last three decades, the Steered Response Power (SRP) method has been widely used for the task of Sound Source Localization (SSL), due to its satisfactory localization performance on moderately reverberant and noisy scenarios. Many works have analyzed and extended the original SRP method to reduce its computational cost, to allow it to locate multiple sources, or to improve its performance i… ▽ More

    Submitted 9 May, 2024; v1 submitted 5 May, 2024; originally announced May 2024.

  4. arXiv:2403.09455  [pdf, other

    cs.SD eess.AS

    The Neural-SRP method for positional sound source localization

    Authors: Eric Grinstein, Toon van Waterschoot, Mike Brookes, Patrick A. Naylor

    Abstract: Steered Response Power (SRP) is a widely used method for the task of sound source localization using microphone arrays, showing satisfactory localization performance on many practical scenarios. However, its performance is diminished under highly reverberant environments. Although Deep Neural Networks (DNNs) have been previously proposed to overcome this limitation, most are trained for a specific… ▽ More

    Submitted 14 March, 2024; originally announced March 2024.

    Comments: Presented at Asilomar Conference on Signals, Systems, and Computers

  5. arXiv:2403.05393  [pdf, other

    eess.AS

    Binaural Speech Enhancement Using Deep Complex Convolutional Transformer Networks

    Authors: Vikas Tokala, Eric Grinstein, Mike Brookes, Simon Doclo, Jesper Jensen, Patrick A. Naylor

    Abstract: Studies have shown that in noisy acoustic environments, providing binaural signals to the user of an assistive listening device may improve speech intelligibility and spatial awareness. This paper presents a binaural speech enhancement method using a complex convolutional neural network with an encoder-decoder architecture and a complex multi-head attention transformer. The model is trained to est… ▽ More

    Submitted 8 March, 2024; originally announced March 2024.

    Comments: Accepted to ICASSP 2024

  6. arXiv:2312.16763  [pdf, other

    eess.AS cs.SD

    Uncertainty Quantification in Machine Learning for Joint Speaker Diarization and Identification

    Authors: Simon W. McKnight, Aidan O. T. Hogg, Vincent W. Neo, Patrick A. Naylor

    Abstract: This paper studies modulation spectrum features ($Φ$) and mel-frequency cepstral coefficients ($Ψ$) in joint speaker diarization and identification (JSID). JSID is important as speaker diarization on its own to distinguish speakers is insufficient for many applications, it is often necessary to identify speakers as well. Machine learning models are set up using convolutional neural networks (CNNs)… ▽ More

    Submitted 30 December, 2023; v1 submitted 27 December, 2023; originally announced December 2023.

    Comments: 12 pages, 7 figures

  7. arXiv:2311.18689  [pdf, other

    eess.AS cs.SD eess.SP

    Subspace Hybrid MVDR Beamforming for Augmented Hearing

    Authors: Sina Hafezi, Alastair H. Moore, Pierre H. Guiraud, Patrick A. Naylor, Jacob Donley, Vladimir Tourbabin, Thomas Lunner

    Abstract: Signal-dependent beamformers are advantageous over signal-independent beamformers when the acoustic scenario - be it real-world or simulated - is straightforward in terms of the number of sound sources, the ambient sound field and their dynamics. However, in the context of augmented reality audio using head-worn microphone arrays, the acoustic scenarios encountered are often far from straightforwa… ▽ More

    Submitted 30 November, 2023; originally announced November 2023.

    Comments: 14 pages, 10 figures, submitted for IEEE/ACM Transactions on Audio, Speech, and Language Processing on 23-Nov-2023

  8. arXiv:2308.04169  [pdf, other

    cs.SD cs.LG eess.AS

    Dual input neural networks for positional sound source localization

    Authors: Eric Grinstein, Vincent W. Neo, Patrick A. Naylor

    Abstract: In many signal processing applications, metadata may be advantageously used in conjunction with a high dimensional signal to produce a desired output. In the case of classical Sound Source Localization (SSL) algorithms, information from a high dimensional, multichannel audio signals received by many distributed microphones is combined with information describing acoustic properties of the scene, s… ▽ More

    Submitted 8 August, 2023; originally announced August 2023.

  9. arXiv:2306.16081  [pdf, other

    cs.SD eess.AS

    Graph neural networks for sound source localization on distributed microphone networks

    Authors: Eric Grinstein, Mike Brookes, Patrick A. Naylor

    Abstract: Distributed Microphone Arrays (DMAs) present many challenges with respect to centralized microphone arrays. An important requirement of applications on these arrays is handling a variable number of input channels. We consider the use of Graph Neural Networks (GNNs) as a solution to this challenge. We present a localization method using the Relation Network GNN, which we show shares many similariti… ▽ More

    Submitted 28 June, 2023; originally announced June 2023.

    Comments: Presented as a poster at ICASSP 2023

  10. arXiv:2306.16071  [pdf, other

    eess.AS cs.CL cs.SD

    Long-term Conversation Analysis: Exploring Utility and Privacy

    Authors: Francesco Nespoli, Jule Pohlhausen, Patrick A. Naylor, Joerg Bitzer

    Abstract: The analysis of conversations recorded in everyday life requires privacy protection. In this contribution, we explore a privacy-preserving feature extraction method based on input feature dimension reduction, spectral smoothing and the low-cost speaker anonymization technique based on McAdams coefficient. We assess the utility of the feature extraction methods with a voice activity detection and a… ▽ More

    Submitted 28 June, 2023; originally announced June 2023.

    Comments: Submitted to ITG Conference on Speech Communication, 2023

  11. arXiv:2306.16069  [pdf, other

    eess.AS cs.SD eess.SP

    Two-Stage Voice Anonymization for Enhanced Privacy

    Authors: Francesco Nespoli, Daniel Barreda, Joerg Bitzer, Patrick A. Naylor

    Abstract: In recent years, the need for privacy preservation when manipulating or storing personal data, including speech , has become a major issue. In this paper, we present a system addressing the speaker-level anonymization problem. We propose and evaluate a two-stage anonymization pipeline exploiting a state-of-the-art anonymization model described in the Voice Privacy Challenge 2022 in combination wit… ▽ More

    Submitted 28 June, 2023; originally announced June 2023.

    Comments: submitted to INTERSPEECH

  12. arXiv:2303.08967  [pdf, other

    eess.AS eess.SP

    Subspace Hybrid Beamforming for Head-worn Microphone Arrays

    Authors: Sina Hafezi, Alastair H. Moore, Pierre Guiraud, Patrick A. Naylor, Jacob Donley, Vladimir Tourbabin, Thomas Lunner

    Abstract: A two-stage multi-channel speech enhancement method is proposed which consists of a novel adaptive beamformer, Hybrid Minimum Variance Distortionless Response (MVDR), Isotropic-MVDR (Iso), and a novel multi-channel spectral Principal Components Analysis (PCA) denoising. In the first stage, the Hybrid-MVDR performs multiple MVDRs using a dictionary of pre-defined noise field models and picks the mi… ▽ More

    Submitted 15 March, 2023; originally announced March 2023.

    Comments: 5 pages, 4 figures, accepted for ICASSP 2023

  13. arXiv:2212.01306  [pdf, other

    eess.AS cs.SD

    Relative Acoustic Features for Distance Estimation in Smart-Homes

    Authors: Francesco Nespoli, Daniel Barreda, Patrick A. Naylor

    Abstract: Any audio recording encapsulates the unique fingerprint of the associated acoustic environment, namely the background noise and reverberation. Considering the scenario of a room equipped with a fixed smart speaker device with one or more microphones and a wearable smart device (watch, glasses or smartphone), we employed the improved proportionate normalized least mean square adaptive filter to est… ▽ More

    Submitted 2 December, 2022; originally announced December 2022.

    Journal ref: Interspeech 2022

  14. arXiv:2209.15472  [pdf, other

    eess.AS eess.SP

    Binaural Speech Enhancement Using STOI-Optimal Masks

    Authors: Vikas Tokala, Mike Brookes, Patrick A. Naylor

    Abstract: STOI-optimal masking has been previously proposed and developed for single-channel speech enhancement. In this paper, we consider the extension to the task of binaural speech enhancement in which spatial information is known to be important to speech understanding and therefore should be preserved by the enhancement processing. Masks are estimated for each of the binaural channels individually and… ▽ More

    Submitted 30 September, 2022; originally announced September 2022.

    Comments: Accepted at IWAENC 2022

  15. arXiv:2203.13919  [pdf

    eess.AS cs.AI

    Spatial Processing Front-End For Distant ASR Exploiting Self-Attention Channel Combinator

    Authors: Dushyant Sharma, Rong Gong, James Fosburgh, Stanislav Yu. Kruchinin, Patrick A. Naylor, Ljubomir Milanovic

    Abstract: We present a novel multi-channel front-end based on channel shortening with theWeighted Prediction Error (WPE) method followed by a fixed MVDR beamformer used in combination with a recently proposed self-attention-based channel combination (SACC) scheme, for tackling the distant ASR problem. We show that the proposed system used as part of a ContextNet based end-to-end (E2E) ASR system outperforms… ▽ More

    Submitted 25 March, 2022; originally announced March 2022.

    Comments: to be presented at ICASSP 2022

  16. arXiv:2004.12745  [pdf, other

    eess.AS cs.SD

    Time-Frequency Analysis and Parameterisation of Knee Sounds for Non-invasive Detection of Osteoarthritis

    Authors: Costas Yiallourides, Patrick A. Naylor

    Abstract: Objective: In this work the potential of non-invasive detection of knee osteoarthritis is investigated using the sounds generated by the knee joint during walking. Methods: The information contained in the time-frequency domain of these signals and its compressed representations is exploited and their discriminant properties are studied. Their efficacy for the task of normal vs abnormal signal cla… ▽ More

    Submitted 27 April, 2020; originally announced April 2020.

    Comments: Submitted to IEEE Transactions on Biomedical Engineering

  17. arXiv:1901.05852  [pdf, other

    eess.AS cs.SD

    Detecting Sound-Absorbing Materials in a Room from a Single Impulse Response using a CRNN

    Authors: Constantinos Papayiannis, Christine Evers, Patrick A. Naylor

    Abstract: The materials of surfaces in a room play an important room in shaping the auditory experience within them. Different materials absorb energy at different levels. The level of absorption also varies across frequencies. This paper investigates how cues from a measured impulse response in the room can be exploited by machines to detect the materials present. With this motivation, this paper proposes… ▽ More

    Submitted 27 October, 2019; v1 submitted 17 January, 2019; originally announced January 2019.

    Comments: Submitted for review for IEEE ICASSP 2020

  18. arXiv:1901.03257  [pdf, other

    eess.AS cs.SD

    Data Augmentation of Room Classifiers using Generative Adversarial Networks

    Authors: Constantinos Papayiannis, Christine Evers, Patrick A. Naylor

    Abstract: The classification of acoustic environments allows for machines to better understand the auditory world around them. The use of deep learning in order to teach machines to discriminate between different rooms is a new area of research. Similarly to other learning tasks, this task suffers from the high-dimensionality and the limited availability of training data. Data augmentation methods have prov… ▽ More

    Submitted 4 December, 2020; v1 submitted 10 January, 2019; originally announced January 2019.

    Comments: Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing

  19. arXiv:1812.09324  [pdf, other

    eess.AS cs.SD

    End-to-End Classification of Reverberant Rooms using DNNs

    Authors: Constantinos Papayiannis, Christine Evers, Patrick A. Naylor

    Abstract: Reverberation is present in our workplaces, our homes, concert halls and theatres. This paper investigates how deep learning can use the effect of reverberation on speech to classify a recording in terms of the room in which it was recorded. Existing approaches in the literature rely on domain expertise to manually select acoustic parameters as inputs to classifiers. Estimation of these parameters… ▽ More

    Submitted 1 November, 2020; v1 submitted 21 December, 2018; originally announced December 2018.

    Comments: Accepted for publication in IEEE/ACM Transactions on Audio, Speech, and Language Processing

  20. arXiv:1811.08482   

    eess.AS cs.SD

    Proceedings of the LOCATA Challenge Workshop -- a satellite event of IWAENC 2018

    Authors: Heinrich W. Loellmann, Christine Evers, Alexander Schmidt, Hendrik Barfuss, Patrick A. Naylor, Walter Kellermann

    Abstract: Algorithms for acoustic source localization and tracking provide estimates of the positional information about active sound sources in acoustic environments and are essential for a wide range of applications such as personal assistants, smart homes, tele-conferencing systems, hearing aids, or autonomous systems. The aim of the IEEE-AASP Challenge on sound source localization and tracking (LOCATA)… ▽ More

    Submitted 20 August, 2019; v1 submitted 20 November, 2018; originally announced November 2018.

    Comments: Workshop Proceedings

  21. arXiv:1606.03365  [pdf, other

    cs.SD

    Acoustic Characterization of Environments (ACE) Challenge Results Technical Report

    Authors: James Eaton, Nikolay D. Gaubitch, Alastair H. Moore, Patrick A. Naylor

    Abstract: This document provides the results of the tests of acoustic parameter estimation algorithms on the Acoustic Characterization of Environments (ACE) Challenge Evaluation dataset which were subsequently submitted and written up into papers for the Proceedings of the ACE Challenge. This document is supporting material for a forthcoming journal paper on the ACE Challenge which will provide further anal… ▽ More

    Submitted 27 June, 2017; v1 submitted 17 December, 2015; originally announced June 2016.

    Comments: Supporting material for Proceedings of the ACE Challenge Workshop - a satellite event of IEEE-WASPAA 2015 (arXiv:1510.00383)

  22. arXiv:1510.07546  [pdf, other

    cs.SD

    Direct-to-Reverberant Ratio Estimation on the ACE Corpus Using a Two-channel Beamformer

    Authors: James Eaton, Patrick A. Naylor

    Abstract: Direct-to-Reverberant Ratio (DRR) is an important measure for characterizing the properties of a room. The recently proposed DRR Estimation using a Null-Steered Beamformer (DENBE) algorithm was originally tested on simulated data where noise was artificially added to the speech after convolution with impulse responses simulated using the image-source method. This paper evaluates the performance of… ▽ More

    Submitted 26 October, 2015; originally announced October 2015.

    Comments: In Proceedings of the ACE Challenge Workshop - a satellite event of IEEE-WASPAA 2015 (arXiv:1510.00383). arXiv admin note: text overlap with arXiv:1510.01193

    Report number: ACEChallenge/2015/03

  23. arXiv:1510.04616  [pdf, ps, other

    cs.SD

    Evaluating the Non-Intrusive Room Acoustics Algorithm with the ACE Challenge

    Authors: Pablo Peso Parada, Dushyant Sharma, Toon van Waterschoot, Patrick A. Naylor

    Abstract: We present a single channel data driven method for non-intrusive estimation of full-band reverberation time and full-band direct-to-reverberant ratio. The method extracts a number of features from reverberant speech and builds a model using a recurrent neural network to estimate the reverberant acoustic parameters. We explore three configurations by including different data and also by combining t… ▽ More

    Submitted 15 October, 2015; originally announced October 2015.

    Comments: In Proceedings of the ACE Challenge Workshop - a satellite event of IEEE-WASPAA 2015 (arXiv:1510.00383)

    Report number: ACEChallenge/2015/06

  24. arXiv:1510.01193  [pdf, other

    cs.SD

    Reverberation time estimation on the ACE corpus using the SDD method

    Authors: James Eaton, Patrick A. Naylor

    Abstract: Reverberation Time (T60) is an important measure for characterizing the properties of a room. The author's T60 estimation algorithm was previously tested on simulated data where the noise is artificially added to the speech after convolution with a impulse responses simulated using the image method. We test the algorithm on speech convolved with real recorded impulse responses and noise from the s… ▽ More

    Submitted 5 October, 2015; originally announced October 2015.

    Comments: In Proceedings of the ACE Challenge Workshop - a satellite event of IEEE-WASPAA 2015 (arXiv:1510.00383)

    Report number: ACEChallenge/2015/02

  25. arXiv:1510.00383  other

    cs.SD

    Proceedings of the ACE Challenge Workshop - a satellite event of IEEE-WASPAA (2015)

    Authors: James Eaton, Nikolay D. Gaubitch, Alastair H. Moore, Patrick A. Naylor

    Abstract: Several established parameters and metrics have been used to characterize the acoustics of a room. The most important are the Direct-To-Reverberant Ratio (DRR), the Reverberation Time (T60) and the reflection coefficient. The acoustic characteristics of a room based on such parameters can be used to predict the quality and intelligibility of speech signals in that room. Recently, several important… ▽ More

    Submitted 1 October, 2015; originally announced October 2015.

    Comments: New Paltz, New York, USA

  26. Source Coding in Networks with Covariance Distortion Constraints

    Authors: Adel Zahedi, Jan Østergaard, Søren Holdt Jensen, Patrick A. Naylor, Søren Bech

    Abstract: We consider a source coding problem with a network scenario in mind, and formulate it as a remote vector Gaussian Wyner-Ziv problem under covariance matrix distortions. We define a notion of minimum for two positive-definite matrices based on which we derive an explicit formula for the rate-distortion function (RDF). We then study the special cases and applications of this result. We show that two… ▽ More

    Submitted 27 September, 2016; v1 submitted 5 April, 2015; originally announced April 2015.