Date of Award

Summer 7-2026

Document Type

Dissertation

Degree Name

Doctor of Philosophy (PhD)

Department

Computer Science

Program/Concentration

Computer Science

Committee Director

Ravi Mukkamala

Committee Director

Fengjiao Wang

Committee Member

Michele C. Weigle

Committee Member

Jiangwen Sun

Abstract

Real-world datasets frequently exhibit severe class imbalance, where certain categories are significantly underrepresented relative to others, leading standard empirical risk minimization to bias learning toward majority classes and degrade performance on minority categories. This dissertation addresses imbalanced classification across both structured tabular data and long-tailed visual recognition by developing methods that improve representation learning and decision reliability under skewed distributions. For tabular data, it introduces Conditional Probability Representation (CPR), a target-based encoding framework that embeds feature–label relationships directly into the representation space, enhanced by a progressive feature upgrading mechanism for semi-supervised settings and a class-frequency–aware extension that improves robustness to label noise and limited supervision. For long-tailed visual recognition, it proposes LTBoost, a two-phase framework combining confusion-matrix-guided mixup augmentation with optimization-based logit adjustment to improve representation quality and reduce majority-class bias. Beyond representation learning, this dissertation addresses the challenge of reliable inference by incorporating conformal prediction and proposing Multiple-Prediction Inference (MPI), a rank-based prediction-set construction method that produces statistically valid prediction sets while explicitly controlling set size. By leveraging class-score rankings rather than probability thresholds, MPI improves the trade-off between coverage and usability, particularly for rare classes. Additionally, this work examines limitations of conventional aggregate evaluation metrics under extreme imbalance and emphasizes the importance of class-wise reliability analysis. Experimental results across multiple real-world datasets demonstrate that the proposed methods improve minority-class performance, enhance robustness to imbalance, and provide more reliable and interpretable predictions. Collectively, this dissertation advances imbalanced learning by integrating representation-aware training and uncertainty aware inference into a unified framework for improving both predictive performance and decision reliability.

Rights

In Copyright. URI: http://rightsstatements.org/vocab/InC/1.0/ This Item is protected by copyright and/or related rights. You are free to use this Item in any way that is permitted by the copyright and related rights legislation that applies to your use. For other uses you need to obtain permission from the rights-holder(s).

DOI

10.25777/xb6p-0j78

ISBN

9798193214144

ORCID

0000-0001-9157-5077

Share

COinS