Jalen Cai

USC · computational 2024

From Wet Lab to Machine Learning: Predicting Sepsis Mortality

Transcriptomic patient data, computationally

The first major move from wet-lab research into computational biology: building and iterating machine-learning workflows to estimate sepsis mortality risk from transcriptomic patient data.

Completed TranscriptomicsMachine learningClinical data
The USC School of Pharmacy poster, Benchmarking of machine learning algorithms to predict mortality in sepsis from transcriptomic data
The poster

Why this one mattered

Everything before this had been physical — animals, plates, a bench. This project was the first major move from wet-lab research into computational biology.

Working with transcriptomic patient data, I helped build and iterate machine-learning workflows for estimating sepsis mortality risk, while learning how choices such as feature selection, validation strategy, and cohort differences can change what a model appears to know.

Methods

Working with public microarray transcriptomic datasets from NCBI GEO and ArrayExpress, our team combined sepsis cohorts and benchmarked four classifiers — SVM, LightGBM, logistic regression, and random forest — for mortality prediction. We compared performance with ROC/AUC, accuracy, confusion matrices, and feature-importance analyses. In the poster analysis, LightGBM produced the highest reported accuracy, about 75%, with an AUC of 0.74.

Logistic regression came in around 72% accuracy (AUC 0.69) and SVM around 70% (AUC 0.74); random forest was close to LightGBM.

Figures

Open the full poster (PDF) to view the ROC curves, confusion matrices, and feature-importance plots for the four benchmarked classifiers.

Images, posters & documents

The full sepsis machine-learning poster, model comparison figures down the centre
Full poster