Home // DBKDA 2012, The Fourth International Conference on Advances in Databases, Knowledge, and Data Applications // View article
Constructing a Synthetic Longitudinal Health Dataset for Data Mining
Authors:
Shima Ghassem Pour
Anthony Maeder
Louisa Jorm
Keywords: cluster analysis; synthetic data
Abstract:
The traditional approach to epidemiological research is to analyse data in an explicit statistical fashion, attempting to answer a question or test a hypothesis. However, increasing experience in the application of data mining and exploratory data analysis methods suggests that valuable information can be obtained from large datasets using these less constrained approaches. Available data mining techniques, such as clustering, have mainly been applied to cross-sectional point-in-time data. However, health datasets often include repeated observations for individuals and so researchers are interested in following their health trajectories. This requires methods for analysis of multiple-points-over-time or longitudinal data. Here, we describe an approach to construct a synthetic longitudinal version of a major population health dataset in which clusters merge and split over time, to investigate the utility of clustering for discovering time sequence based patterns.
Pages: 86 to 90
Copyright: Copyright (c) IARIA, 2012
Publication date: February 29, 2012
Published in: conference
ISSN: 2308-4332
ISBN: 978-1-61208-185-4
Location: Saint Gilles, Reunion
Dates: from February 29, 2012 to March 5, 2012