Document Type
Article
Publication Date
2026
DOI
10.4236/cmb.2026.162002
Publication Title
Computational Molecular Bioscience
Volume
16
Issue
2
Pages
17-36
Abstract
Prostate cancer disproportionately impacts African American men, who experience significantly higher mortality rates and earlier disease onset than other populations. Current diagnostic approaches, including prostate-specific antigen testing and biopsy, lack sufficient specificity and sensitivity, underscoring the need for accurate, molecular-level classification tools. This paper presents a machine learning framework for binary classification of genomic DNA sequences as cancerous or healthy. A dataset of 1684 FASTA-formatted sequences obtained from the National Library of Medicine - GenBank was analyzed, with 1662 sequences retained after quality control filtering. Feature engineering yielded 67 attributes, including GC content, Shannon entropy, sequence length, and trinucleotide k-mer frequencies. To address class imbalance, we applied the Synthetic Minority Over-sampling Technique to the training data. Seven classification algorithms were evaluated using stratified train-test splits, cross-validation, and hyperparameter optimization. Among the models, the optimized Random Forest classifier achieved superior performance, with a cross-validation accuracy of 97.2% (±0.006), a weighted F1-score of 0.95, a cancer-class recall of 0.96, and an ROC-AUC of 0.974. Feature importance analysis identified sequence length and Shannon entropy as the most discriminative predictors, followed by specific trinucleotide motifs (TTC, AAC, ACC, and GGG). These results demonstrate the potential of interpretable machine learning approaches for genomic sequence-based PCa classification, offering a promising pathway toward improved, equitable diagnostic tools for high-risk populations.
Rights
© 2026 The Authors.
This work is licensed under the Creative Commons Attribution International (CC BY 4.0) License.
Original Publication Citation
Rawat, K., Banerjee, H. N., Noble, J., Deloatch, S. N., Banerjee, S., Shetty, S., & Banerjee, S. (2026). Machine learning classification of prostate cancer genomic sequences using k-mer and sequence-derived features. Computational Molecular Bioscience, 16(2), 17-36. https://doi.org/10.4236/cmb.2026.162002
ORCID
0000-0002-8789-0610 (Shetty)
Repository Citation
Rawat, Kuldeep; Banerjee, Hirendra Nath; Noble, Jamie; Deloatch, Saa Naudia; Banerjee, Satyendra; Shetty, Sachin; and Banerjee, Soumya, "Machine Learning Classification of Prostate Cancer Genomic Sequences Using K-Mer and Sequence-Derived Features" (2026). VMASC Publications. 161.
https://digitalcommons.odu.edu/vmasc_pubs/161
Included in
Artificial Intelligence and Robotics Commons, Oncology Commons, Theory and Algorithms Commons