USC · computational 2024
From Wet Lab to Machine Learning: Predicting Sepsis Mortality
Transcriptomic patient data, computationally
The first major move from wet-lab research into computational biology: building and iterating machine-learning workflows to estimate sepsis mortality risk from transcriptomic patient data.
Why this one mattered
Everything before this had been physical — animals, plates, a bench. This project was the first major move from wet-lab research into computational biology.
Working with transcriptomic patient data, I helped build and iterate machine-learning workflows for estimating sepsis mortality risk, while learning how choices such as feature selection, validation strategy, and cohort differences can change what a model appears to know.
Methods
Working with public microarray transcriptomic datasets from NCBI GEO and ArrayExpress, our team combined sepsis cohorts and benchmarked four classifiers — SVM, LightGBM, logistic regression, and random forest — for mortality prediction. We compared performance with ROC/AUC, accuracy, confusion matrices, and feature-importance analyses. In the poster analysis, LightGBM produced the highest reported accuracy, about 75%, with an AUC of 0.74.
Logistic regression came in around 72% accuracy (AUC 0.69) and SVM around 70% (AUC 0.74); random forest was close to LightGBM.
Figures
Open the full poster (PDF) to view the ROC curves, confusion matrices, and feature-importance plots for the four benchmarked classifiers.
Images, posters & documents