Speech Decomposition Based on a Hybrid Speech Model and Optimal Segmentation

Alfredo Esquivel Jaramillo; Jesper Kjær Nielsen; Mads Græsbøll Christensen

doi:10.21437/Interspeech.2021-47

Speech Decomposition Based on a Hybrid Speech Model and Optimal Segmentation

Alfredo Esquivel Jaramillo, Jesper Kjær Nielsen, Mads Græsbøll Christensen

Publikation: Bidrag til bog/antologi/rapport/konference proceeding › Konferenceartikel i proceeding › Forskning › peer review

1 Citationer (Scopus)

Abstract

In a hybrid speech model, both voiced and unvoiced components can coexist in a segment. Often, the voiced speech is regarded as the deterministic component, and the unvoiced speech and additive noise are the stochastic components. Typically, the speech signal is considered stationary within fixed segments of 20-40 ms, but the degree of stationarity varies over time. For decomposing noisy speech into its voiced and unvoiced components, a fixed segmentation may be too crude, and we here propose to adapt the segment length according to the signal local characteristics. The segmentation relies on parameter estimates of a hybrid speech model and the maximum a posteriori (MAP) and log-likelihood criteria as rules for model selection among the possible segment lengths, for voiced and unvoiced speech, respectively. Given the optimal segmentation markers and the estimated statistics, both components are estimated using linear filtering. A codebook-based approach differentiates between unvoiced speech and noise. A better extraction of the components is possible by taking into account the adaptive segmentation, compared to a fixed one. Also, a lower distortion for voiced speech and higher segSNR for both components is possible, as compared to other decomposition methods.

Originalsprog	Engelsk
Titel	Proc. Interspeech 2021
Publikationsdato	2021
Sider	1164-1168
DOI	https://doi.org/10.21437/Interspeech.2021-47
Status	Udgivet - 2021
Begivenhed	22nd Annual Conference of the International Speech Communication Association, INTERSPEECH 2021 - Brno, Tjekkiet Varighed: 30 aug. 2021 → 3 sep. 2021

Konference

Konference	22nd Annual Conference of the International Speech Communication Association, INTERSPEECH 2021
Land/Område	Tjekkiet
By	Brno
Periode	30/08/2021 → 03/09/2021

Adgang til dokumentet

10.21437/Interspeech.2021-47Licens: Ikke-specificeret

https://arxiv.org/pdf/2105.01302.pdf
https://www.isca-speech.org/archive/pdfs/interspeech_2021/jaramillo21_interspeech.pdfLicens: Ikke-specificeret

AUB Link

Søg efter materialet i Aalborg Universitetsbiblioteks søgemaskine

Andre filer og links

Slides_presentation

Citationsformater

@inproceedings{5d0f72b9271d41b0ac7f905fe2d06b1e,

title = "Speech Decomposition Based on a Hybrid Speech Model and Optimal Segmentation",

abstract = "In a hybrid speech model, both voiced and unvoiced components can coexist in a segment. Often, the voiced speech is regarded as the deterministic component, and the unvoiced speech and additive noise are the stochastic components. Typically, the speech signal is considered stationary within fixed segments of 20-40 ms, but the degree of stationarity varies over time. For decomposing noisy speech into its voiced and unvoiced components, a fixed segmentation may be too crude, and we here propose to adapt the segment length according to the signal local characteristics. The segmentation relies on parameter estimates of a hybrid speech model and the maximum a posteriori (MAP) and log-likelihood criteria as rules for model selection among the possible segment lengths, for voiced and unvoiced speech, respectively. Given the optimal segmentation markers and the estimated statistics, both components are estimated using linear filtering. A codebook-based approach differentiates between unvoiced speech and noise. A better extraction of the components is possible by taking into account the adaptive segmentation, compared to a fixed one. Also, a lower distortion for voiced speech and higher segSNR for both components is possible, as compared to other decomposition methods.",

keywords = "speech decomposition, hybrid speech model, optimal segmentation, autoregressive",

author = "{Esquivel Jaramillo}, Alfredo and Nielsen, {Jesper Kj{\ae}r} and Christensen, {Mads Gr{\ae}sb{\o}ll}",

year = "2021",

doi = "10.21437/Interspeech.2021-47",

language = "English",

pages = "1164--1168",

booktitle = "Proc. Interspeech 2021",

note = "22nd Annual Conference of the International Speech Communication Association, INTERSPEECH 2021 ; Conference date: 30-08-2021 Through 03-09-2021",

}

TY - GEN

T1 - Speech Decomposition Based on a Hybrid Speech Model and Optimal Segmentation

AU - Esquivel Jaramillo, Alfredo

AU - Nielsen, Jesper Kjær

AU - Christensen, Mads Græsbøll

PY - 2021

Y1 - 2021

N2 - In a hybrid speech model, both voiced and unvoiced components can coexist in a segment. Often, the voiced speech is regarded as the deterministic component, and the unvoiced speech and additive noise are the stochastic components. Typically, the speech signal is considered stationary within fixed segments of 20-40 ms, but the degree of stationarity varies over time. For decomposing noisy speech into its voiced and unvoiced components, a fixed segmentation may be too crude, and we here propose to adapt the segment length according to the signal local characteristics. The segmentation relies on parameter estimates of a hybrid speech model and the maximum a posteriori (MAP) and log-likelihood criteria as rules for model selection among the possible segment lengths, for voiced and unvoiced speech, respectively. Given the optimal segmentation markers and the estimated statistics, both components are estimated using linear filtering. A codebook-based approach differentiates between unvoiced speech and noise. A better extraction of the components is possible by taking into account the adaptive segmentation, compared to a fixed one. Also, a lower distortion for voiced speech and higher segSNR for both components is possible, as compared to other decomposition methods.

AB - In a hybrid speech model, both voiced and unvoiced components can coexist in a segment. Often, the voiced speech is regarded as the deterministic component, and the unvoiced speech and additive noise are the stochastic components. Typically, the speech signal is considered stationary within fixed segments of 20-40 ms, but the degree of stationarity varies over time. For decomposing noisy speech into its voiced and unvoiced components, a fixed segmentation may be too crude, and we here propose to adapt the segment length according to the signal local characteristics. The segmentation relies on parameter estimates of a hybrid speech model and the maximum a posteriori (MAP) and log-likelihood criteria as rules for model selection among the possible segment lengths, for voiced and unvoiced speech, respectively. Given the optimal segmentation markers and the estimated statistics, both components are estimated using linear filtering. A codebook-based approach differentiates between unvoiced speech and noise. A better extraction of the components is possible by taking into account the adaptive segmentation, compared to a fixed one. Also, a lower distortion for voiced speech and higher segSNR for both components is possible, as compared to other decomposition methods.

KW - speech decomposition

KW - hybrid speech model

KW - optimal segmentation

KW - autoregressive

U2 - 10.21437/Interspeech.2021-47

DO - 10.21437/Interspeech.2021-47

M3 - Article in proceeding

SP - 1164

EP - 1168

BT - Proc. Interspeech 2021

T2 - 22nd Annual Conference of the International Speech Communication Association, INTERSPEECH 2021

Y2 - 30 August 2021 through 3 September 2021

ER -

Speech Decomposition Based on a Hybrid Speech Model and Optimal Segmentation

Abstract

Konference

Adgang til dokumentet

AUB Link

Andre filer og links

Fingeraftryk

Citationsformater