Résumé
Model-based unsupervised learning, as any learning task, stalls as soon as
missing data occurs. This is even more true when the missing data are
informative, or said missing not at random (MNAR). In this paper, we propose
model-based clustering algorithms designed to handle very general types of
missing data, including MNAR data. To do so, we introduce a mixture model for
different types of data (continuous, count, categorical and mixed) to jointly
model the data distribution and the MNAR mechanism, remaining vigilant to the
relative degrees of freedom of each. Several MNAR models are discussed, for
which the cause of the missingness can depend on both the values of the missing
variable themselves and on the class membership. However, we focus on a
specific MNAR model, called MNARz, for which the missingness only depends on
the class membership. We first underline its ease of estimation, by showing
that the statistical inference can be carried out on the data matrix
concatenated with the missing mask considering finally a standard MAR
mechanism. Consequently, we propose to perform clustering using the Expectation
Maximization algorithm, specially developed for this simplified
reinterpretation. Finally, we assess the numerical performances of the proposed
methods on synthetic data and on the real medical registry TraumaBase as well.