Thanh
M. Brown
Data Analyst with an M.S. in Operations Research and 5+ years working with complex clinical and genomic data at 3 academic medical centers. I combine HPC-scale pipeline engineering, observational study design, and hands-on ML to turn raw data into reproducible, publication-ready results, including co-authored peer-reviewed research.
Capabilities
Technical Skills
Full-stack data science from raw data to deployed model.
Work
Portfolio Projects
Health data, clinical ML, large-scale EDA, and bioinformatics research.
01 — Data Pipeline Engineering - Clinical Analytics Platform
ISCVAM — Multi-Study Oncology Data Pipelines Cancer Research
Experimental scientists had no direct computational path to their own data, leaving finished experiments stuck in an analyst queue. I built the automated data layer that removed the delay: two containerized R pipelines handling end-to-end ingestion, quality control, and cell-type annotation across 198 studies (~6.3M cells). Orchestrated on Slurm job arrays with size-aware scheduling and per-step checkpointing, the pipelines convert raw sequencing output into standardized HDF5 datasets the platform reads directly. Platform accepted for presentation at AACR 2023.
02 — Clinical ML · Epidemiology · Python
Osteoporosis Risk — Recall Asymmetry in Matched Case-Control Data
Why do models miss cases on a perfectly balanced dataset? In a 1:1 matched cohort of 61,026 osteoporosis patients, classifiers from logistic regression and linear SVM to tree ensembles and a feed-forward neural network miss 40–49% of cases, despite exact class balance and ROC-AUCs up to 0.79. After validating the cohort by reproducing a published hyponatremia odds ratio (3.98 vs. 3.97), I use diagnostic analyses to test whether the recall gap stems from model limitations or from properties of the matched observational data.
03 — ML App · Streamlit · Scikit-learn
Hiring Intelligence — Recruitment Outcome Predictor
End-to-end ML web app predicting candidate hiring outcomes using a tuned Gradient Boosting classifier — 94.7% precision, 93.3% ROC-AUC on 1,500 recruitment records. The web app features an EDA explorer and live candidate predictor with probability gauge and radar chart. Key finding: recruitment strategy (SHAP = 3.052) dominates all candidate-level signals — stronger than interview score, skill score, or education combined. A useful reminder that process design shapes outcomes as much as candidate quality does.
04 — FDA Data · R · Interactive App
FDA Medical Device Harm Trends — RShiny Dashboard
Built an interactive visualization app over the 2016 MAUDE (FDA medical device passive surveillance) dataset. Users explore temporal harm trends across device categories and manufacturers. Demonstrates stakeholder-facing data product design.
05 — Public Health · Unsupervised Learning
COVID-19 Vaccine Adverse Symptoms — Association Rule Mining
Mined COVID-19 adverse-event reports from VAERS (CDC/FDA) using the Apriori algorithm to surface frequent adverse-symptom patterns and compare reported events between Moderna and Pfizer. Key insight: adverse-event patterns were broadly similar across both vaccines, suggesting perceived safety differences may be driven more by reporting frequency than by fundamentally different symptom profiles.
Background
About Me
I'm a data analyst with an M.S. in Operations Research and 5+ years at academic medical centers (MedStar Health, Moffitt Cancer Center, Huntsman Cancer Institute), working with clinical EHR data, FDA safety reports, and high-dimensional genomic profiles. My work spans statistical modeling, machine learning, and large-scale data pipelines. I focus on the questions that matter before modeling: what the data can support, where bias enters, and when an association isn't enough to act on. I own the full path from raw, messy sources to reproducible results and tools that non-technical stakeholders actually use.
Core competencies
- Scale & Engineering Rigor: Engineered parallelized R/Python ETL and analysis pipelines on HPC/SLURM clusters with run-tracking checkpointing (6.3M+ cells across 198 studies).
- Methodological Depth: Paired machine learning with advanced observational study design, including matched case-control pairing, conditional logistic regression, and Rubin’s rules for multiple imputation.
- Oncology Genomics: Built analysis pipelines for single-cell RNA-seq, single-cell ATAC-seq, and multiome data—covering QC, dimension reduction, clustering, and annotation with Seurat and Signac—with additional work in spatial transcriptomics and bulk sequencing.
- Real-World & Federal Data: Executed hands-on data retrieval, extraction, and modeling across structured and unstructured sources, including hospital EHR data, FDA MAUDE, and VAERS safety databases.
- Scientific Impact: Co-authored peer-reviewed research in top-tier journals (Clinical Cancer Research and Cancer Research) and presented first-author platform work at AACR 2023.
Contact