Department of Electrical Engineering and Computer Science

Significance Metrics for Clusters of Mixed Numerical and Categorical Yeast Data

Bill Andreopoulos, Aijun An and Xiaogang Wang

Technical Report CS-2003-12

York University

October 2003

Abstract

We designed, implemented and tested a clustering tool for numerical data sets derived from gene expression DNA microarray studies. Our tool incorporates into the clustering process the existing knowledge as categorical annotations (CAs) on the genes, as well as the uncertainties concerning the correctness of the CAs as confidence values (CVs) on the CAs. CVs are a measure of the certainty of correctness of the existing knowledge and are derived from GeneOntology Evidence Codes. This allowed us to apply new significance metrics to the resulting clusters to extract the most prominent CAs in the clusters. We applied the extracted CAs to the other genes in the cluster to predict their function and we validated the predictions.

Download paper in PDF format.

The documents distributed by this server have been provided by the contributing authors as a means to ensure timely dissemination of scholarly and technical work on a noncommercial basis. Copyright and all rights therein are maintained by the authors or by other copyright holders, notwithstanding that they have offered their works here electronically. It is understood that all persons copying this information will adhere to the terms and constraints invoked by each author's copyright. These works may not be reposted without the explicit permission of the copyright holder.

Department of Electrical Engineering & Computer Science

Significance Metrics for Clusters of Mixed Numerical and Categorical Yeast Data

Bill Andreopoulos, Aijun An and Xiaogang Wang