Document Type

Article

Publication Date

2026

DOI

10.4236/cmb.2026.162002

Publication Title

Computational Molecular Bioscience

Volume

16

Issue

2

Pages

17-36

Abstract

Prostate cancer disproportionately impacts African American men, who experience significantly higher mortality rates and earlier disease onset than other populations. Current diagnostic approaches, including prostate-specific antigen testing and biopsy, lack sufficient specificity and sensitivity, underscoring the need for accurate, molecular-level classification tools. This paper presents a machine learning framework for binary classification of genomic DNA sequences as cancerous or healthy. A dataset of 1684 FASTA-formatted sequences obtained from the National Library of Medicine - GenBank was analyzed, with 1662 sequences retained after quality control filtering. Feature engineering yielded 67 attributes, including GC content, Shannon entropy, sequence length, and trinucleotide k-mer frequencies. To address class imbalance, we applied the Synthetic Minority Over-sampling Technique to the training data. Seven classification algorithms were evaluated using stratified train-test splits, cross-validation, and hyperparameter optimization. Among the models, the optimized Random Forest classifier achieved superior performance, with a cross-validation accuracy of 97.2% (±0.006), a weighted F1-score of 0.95, a cancer-class recall of 0.96, and an ROC-AUC of 0.974. Feature importance analysis identified sequence length and Shannon entropy as the most discriminative predictors, followed by specific trinucleotide motifs (TTC, AAC, ACC, and GGG). These results demonstrate the potential of interpretable machine learning approaches for genomic sequence-based PCa classification, offering a promising pathway toward improved, equitable diagnostic tools for high-risk populations.

Rights

© 2026 The Authors.

This work is licensed under the Creative Commons Attribution International (CC BY 4.0) License.

Original Publication Citation

Rawat, K., Banerjee, H. N., Noble, J., Deloatch, S. N., Banerjee, S., Shetty, S., & Banerjee, S. (2026). Machine learning classification of prostate cancer genomic sequences using k-mer and sequence-derived features. Computational Molecular Bioscience, 16(2), 17-36. https://doi.org/10.4236/cmb.2026.162002

ORCID

0000-0002-8789-0610 (Shetty)

Share

COinS