On testing the missing at random assumption

Manfred Jaeger

On testing the missing at random assumption

Department of Computer Science

Research output: Contribution to book/anthology/report/conference proceeding › Article in proceeding › Research › peer-review

15 Citations (Scopus)

Abstract

Most approaches to learning from incomplete data are based on the assumption
that unobserved values are missing at random (mar). While the mar assumption, as such, is not testable, it can become testable in the context of other distributional assumptions, e.g. the naive Bayes assumption. In this paper we investigate a method for testing the mar assumption in the presence of other distributional constraints. We present methods to (approximately) compute a test statistic consisting of the ratio of two profile likelihood functions. This requires the optimization of the likelihood under no assumptionson the missingness mechanism, for which we use our recently proposed
AI \& M algorithm. We present experimental results on synthetic data that show that our approximate test statistic is a good indicator for whether data is mar relative to the given distributional assumptions.

Original language	English
Title of host publication	Machine Learning: ECML 2006 : 17th European Conference on Machine Learning. Berlin, Germany, September 2006. Proceedings
Number of pages	8
Publication date	2006
Pages	671-678
Publication status	Published - 2006
Event	European Conference on Machine Learning - Berlin, Germany Duration: 18 Sept 2006 → 22 Sept 2006

Conference

Conference	European Conference on Machine Learning
Country/Territory	Germany
City	Berlin
Period	18/09/2006 → 22/09/2006

AUB Link

Search for the material in Aalborg University Library's search engine

Cite this

@inproceedings{c6234ae0a48c11db8ed6000ea68e967b,

title = "On testing the missing at random assumption",

abstract = "Most approaches to learning from incomplete data are based on the assumptionthat unobserved values are missing at random (mar). While the mar assumption, as such, is not testable, it can become testable in the context of other distributional assumptions, e.g. the naive Bayes assumption. In this paper we investigate a method for testing the mar assumption in the presence of other distributional constraints. We present methods to (approximately) compute a test statistic consisting of the ratio of two profile likelihood functions. This requires the optimization of the likelihood under no assumptionson the missingness mechanism, for which we use our recently proposedAI \& M algorithm. We present experimental results on synthetic data that show that our approximate test statistic is a good indicator for whether data is mar relative to the given distributional assumptions.",

author = "Manfred Jaeger",

note = "Serie: Lecture Notes in Artificial Intelligence, Springer-Verlag, 4212, 0302-9743 ; European Conference on Machine Learning ; Conference date: 18-09-2006 Through 22-09-2006",

year = "2006",

language = "English",

pages = "671--678",

booktitle = "Machine Learning: ECML 2006",

}

TY - GEN

T1 - On testing the missing at random assumption

AU - Jaeger, Manfred

N1 - Serie: Lecture Notes in Artificial Intelligence, Springer-Verlag, 4212, 0302-9743

PY - 2006

Y1 - 2006

N2 - Most approaches to learning from incomplete data are based on the assumptionthat unobserved values are missing at random (mar). While the mar assumption, as such, is not testable, it can become testable in the context of other distributional assumptions, e.g. the naive Bayes assumption. In this paper we investigate a method for testing the mar assumption in the presence of other distributional constraints. We present methods to (approximately) compute a test statistic consisting of the ratio of two profile likelihood functions. This requires the optimization of the likelihood under no assumptionson the missingness mechanism, for which we use our recently proposedAI \& M algorithm. We present experimental results on synthetic data that show that our approximate test statistic is a good indicator for whether data is mar relative to the given distributional assumptions.

AB - Most approaches to learning from incomplete data are based on the assumptionthat unobserved values are missing at random (mar). While the mar assumption, as such, is not testable, it can become testable in the context of other distributional assumptions, e.g. the naive Bayes assumption. In this paper we investigate a method for testing the mar assumption in the presence of other distributional constraints. We present methods to (approximately) compute a test statistic consisting of the ratio of two profile likelihood functions. This requires the optimization of the likelihood under no assumptionson the missingness mechanism, for which we use our recently proposedAI \& M algorithm. We present experimental results on synthetic data that show that our approximate test statistic is a good indicator for whether data is mar relative to the given distributional assumptions.

M3 - Article in proceeding

SP - 671

EP - 678

BT - Machine Learning: ECML 2006

T2 - European Conference on Machine Learning

Y2 - 18 September 2006 through 22 September 2006

ER -

On testing the missing at random assumption

Abstract

Conference

AUB Link

Fingerprint

Cite this