Instantaneous Fundamental Frequency Estimation with Optimal Segmentation for Nonstationary Voiced Speech

Sidsel Marie Nørholm; Jesper Rindom Jensen; Mads Græsbøll Christensen

doi:10.1109/TASLP.2016.2608948

Instantaneous Fundamental Frequency Estimation with Optimal Segmentation for Nonstationary Voiced Speech

Sidsel Marie Nørholm, Jesper Rindom Jensen, Mads Græsbøll Christensen

Research output: Contribution to journal › Journal article › Research › peer-review

25 Citations (Scopus)

376 Downloads (Pure)

Abstract

In speech processing, the speech is often considered stationary within segments of 20–30 ms even though it is well known not to be true. In this paper, we take the non-stationarity of voiced speech into account by using a linear chirp model to describe the speech signal. We propose a maximum likelihood estimator of the fundamental frequency and chirp rate of this model, and show that it reaches the Cramer-Rao bound. Since the speech varies over time, a fixed segment length is not optimal, and we propose to make a segmentation of the signal based on the maximum a posteriori (MAP) criterion. Using this segmentation method, the segments are on average seen to be longer for the chirp model compared to the traditional harmonic model. For the signal under test, the average segment length is 24.4 ms and 17.1 ms for the chirp model and traditional harmonic model, respectively. This suggests a better fit of the chirp model than the harmonic model to the speech signal. The methods are based on an assumption of white Gaussian noise, and, therefore, two prewhitening filters are also proposed.

Original language	English
Article number	756754
Journal	I E E E Transactions on Audio, Speech and Language Processing
Volume	24
Issue number	12
Pages (from-to)	2354-2367
ISSN	1558-7916
DOIs	https://doi.org/10.1109/TASLP.2016.2608948
Publication status	Published - Dec 2016

Keywords

Harmonic chirp model
Parameter estimation
Segmentation
Prewhitening

Access to Document

10.1109/TASLP.2016.2608948

chirp_ultimateSubmitted manuscript, 623 KBLicence: Unspecified

AUB Link

Search for the material in Aalborg University Library's search engine

Localization and Tracking of Speech - a Joint Audio-Visual Approach
Jensen, J. R.
01/10/2013 → 30/09/2016
Project: Research
Spatio-Temporal Filtering Methods for Enhancement and Separation of Speech Signals
Christensen, M. G., Nørholm, S. M., Karimian-Azari, S. & Jensen, J. R.
01/08/2012 → 30/06/2015
Project: Research

Cite this

@article{044f42f77c2e47988688cbeb8949c444,

title = "Instantaneous Fundamental Frequency Estimation with Optimal Segmentation for Nonstationary Voiced Speech",

abstract = "In speech processing, the speech is often considered stationary within segments of 20–30 ms even though it is well known not to be true. In this paper, we take the non-stationarity of voiced speech into account by using a linear chirp model to describe the speech signal. We propose a maximum likelihood estimator of the fundamental frequency and chirp rate of this model, and show that it reaches the Cramer-Rao bound. Since the speech varies over time, a fixed segment length is not optimal, and we propose to make a segmentation of the signal based on the maximum a posteriori (MAP) criterion. Using this segmentation method, the segments are on average seen to be longer for the chirp model compared to the traditional harmonic model. For the signal under test, the average segment length is 24.4 ms and 17.1 ms for the chirp model and traditional harmonic model, respectively. This suggests a better fit of the chirp model than the harmonic model to the speech signal. The methods are based on an assumption of white Gaussian noise, and, therefore, two prewhitening filters are also proposed.",

keywords = "Harmonic chirp model, parameter estimation, segmentation, prewhitening., Harmonic chirp model, Parameter estimation, Segmentation, Prewhitening",

author = "N{\o}rholm, {Sidsel Marie} and Jensen, {Jesper Rindom} and Christensen, {Mads Gr{\ae}sb{\o}ll}",

year = "2016",

month = dec,

doi = "10.1109/TASLP.2016.2608948",

language = "English",

volume = "24",

pages = "2354--2367",

journal = "I E E E Transactions on Audio, Speech and Language Processing",

issn = "1558-7916",

publisher = "IEEE Signal Processing Society",

number = "12",

}

Instantaneous Fundamental Frequency Estimation with Optimal Segmentation for Nonstationary Voiced Speech. / Nørholm, Sidsel Marie; Jensen, Jesper Rindom ; Christensen, Mads Græsbøll.
In: I E E E Transactions on Audio, Speech and Language Processing, Vol. 24, No. 12, 756754, 12.2016, p. 2354-2367.

Research output: Contribution to journal › Journal article › Research › peer-review

TY - JOUR

T1 - Instantaneous Fundamental Frequency Estimation with Optimal Segmentation for Nonstationary Voiced Speech

AU - Nørholm, Sidsel Marie

AU - Jensen, Jesper Rindom

AU - Christensen, Mads Græsbøll

PY - 2016/12

Y1 - 2016/12

N2 - In speech processing, the speech is often considered stationary within segments of 20–30 ms even though it is well known not to be true. In this paper, we take the non-stationarity of voiced speech into account by using a linear chirp model to describe the speech signal. We propose a maximum likelihood estimator of the fundamental frequency and chirp rate of this model, and show that it reaches the Cramer-Rao bound. Since the speech varies over time, a fixed segment length is not optimal, and we propose to make a segmentation of the signal based on the maximum a posteriori (MAP) criterion. Using this segmentation method, the segments are on average seen to be longer for the chirp model compared to the traditional harmonic model. For the signal under test, the average segment length is 24.4 ms and 17.1 ms for the chirp model and traditional harmonic model, respectively. This suggests a better fit of the chirp model than the harmonic model to the speech signal. The methods are based on an assumption of white Gaussian noise, and, therefore, two prewhitening filters are also proposed.

AB - In speech processing, the speech is often considered stationary within segments of 20–30 ms even though it is well known not to be true. In this paper, we take the non-stationarity of voiced speech into account by using a linear chirp model to describe the speech signal. We propose a maximum likelihood estimator of the fundamental frequency and chirp rate of this model, and show that it reaches the Cramer-Rao bound. Since the speech varies over time, a fixed segment length is not optimal, and we propose to make a segmentation of the signal based on the maximum a posteriori (MAP) criterion. Using this segmentation method, the segments are on average seen to be longer for the chirp model compared to the traditional harmonic model. For the signal under test, the average segment length is 24.4 ms and 17.1 ms for the chirp model and traditional harmonic model, respectively. This suggests a better fit of the chirp model than the harmonic model to the speech signal. The methods are based on an assumption of white Gaussian noise, and, therefore, two prewhitening filters are also proposed.

KW - Harmonic chirp model, parameter estimation, segmentation, prewhitening.

KW - Harmonic chirp model

KW - Parameter estimation

KW - Segmentation

KW - Prewhitening

U2 - 10.1109/TASLP.2016.2608948

DO - 10.1109/TASLP.2016.2608948

M3 - Journal article

SN - 1558-7916

VL - 24

SP - 2354

EP - 2367

JO - I E E E Transactions on Audio, Speech and Language Processing

JF - I E E E Transactions on Audio, Speech and Language Processing

IS - 12

M1 - 756754

ER -

Instantaneous Fundamental Frequency Estimation with Optimal Segmentation for Nonstationary Voiced Speech

Abstract

Keywords

Access to Document

AUB Link

Fingerprint

Projects

Localization and Tracking of Speech - a Joint Audio-Visual Approach

Spatio-Temporal Filtering Methods for Enhancement and Separation of Speech Signals

Cite this