A Bayesian Feature Selection Method for Multiclass Classification with Application to Cancer Gene Expression Data

Read the full article

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Feature selection is essential for high-dimensional datasets, as it facilitates analysis and classification by finding the most informative features. This particularly important in gene expression data, where disease prediction and diagnosis are critical. Cancer microarray and sequencing datasets are characterized by high dimensionality and limited sample sizes, which pose significant challenges to large-scale data-driven classification. Since not all genes (features) contribute equally to predictive performance, effective feature selection is crucial in reducing dimensionality, improving model interpretability, and enhancing computational efficiency. In this study, we introduce a multiclass classification method that employs the relative belief ratio as a filtering criterion for high dimensional data. The proposed method generates a strength measure that serves as an importance score for each feature and enables systematic assessment of feature relevance across multiple classes. The resulting scores are used to rank features according to their discriminative power with respect to the target variable. The effectiveness of the proposed approach is evaluated using synthetic data and eight multiclass cancer gene expression datasets, and it is compared with several baseline filter-based feature selection methods. Four popular classifiers: logistic regression, random forest, K-nearest neighbors, and support vector machine were employed for classification. Experimental results demonstrate that the proposed approach is capable of identifying the relevant genes associated with various types of cancer (leukemia, colon tumor, lung cancer, breast cancer, lymphoma, and central nervous system tumor) and achieves competitive or superior classification accuracy while maintaining interpretability and low dimensional representations, highlighting its potential for biomedical research and large-scale machine learning and data analytics applications.

Article activity feed