Abstract
The separation of unknown sources from an acoustic scene remains one of the main challenges in the analysis of audio signals. In this paper we address the problem of source separation in multitalker scenarios, attempting a speaker-independent segmentation of the scene, based on the analysis of relative changes in the grouping features, followed by an ideal binary mask (IBM) estimation. In the presented algorithm, periodicity, power, and spatial features are combined in a general framework to estimate the on- and offsets of acoustic sources. A deep neural network (DNN) is then applied to convert relative feature changes into spectro-temporal contrast maps which are used to extract auditory glimpses, i.e. spectro-temporal segments which are dominated by the same source. By exploiting relative feature changes instead of absolute source properties, the presented glimpse formation is well applicable to unseen acoustic conditions or speakers. Using retrospective labeling, the formed glimpses are then assigned to fixed azimuth locations, providing the final IBM estimate. A systematic evaluation with up to five competing speakers and interfering diffuse noise shows that the segmentation procedure generalizes well to unseen speakers as well as to untrained noise conditions and source numbers. A comparison between the final IBM estimates and the output of a baseline system, which performs the source segregation based on spatial cues only, shows that the contrast based segmentation utilizing multiple features and spectro-temporal context can improve the quality of the estimated IBMs, in particular in conditions with multiple interfering talkers.
| Original language | English |
|---|---|
| Journal | IEEE/ACM Transactions on Audio Speech and Language Processing |
| Volume | 30 |
| Pages (from-to) | 1249-1262 |
| ISSN | 2329-9290 |
| DOIs | |
| Publication status | Published - 2022 |
Fingerprint
Dive into the research topics of 'Segmentation of Multitalker Mixtures Based on Local Feature Contrasts and Auditory Glimpses'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver