High Dimensional data clustering using Partition-Constraint Algorithm for knowledge Discovery
Abstract
Computational efficiency and result quality both necessitate that huge data-bases have been clustered. The data mining experts believe that feature space clustering over the original data space is necessary in order to reach both of the aforementioned goals. As a consequence, we used COP-KMEANS (Con-straint-Partitioning K-Means) clustering on our high-dimensional dataset. This method did not successfully group the data into effective and efficient clus-ters, because of the inherent sparseness of our high-dimensional dataset, and therefore produced erroneous and indeterminate clusters. Dimensionality re-duction can apply on original dataset as a preparatory step to high dimen-sional data clustering. Once we have successfully clustered the dimensions reduced dataset with the COP-KMEANS method, we will do that task again, but this time on the resulting clusters. We use two artificial high-dimensional datasets to test the working of the proposed technique. According to the ex-perimental results, the suggested method is highly successful in built-up ac-curate and exact clusters.







