Thanh
M. Brown

Data Analyst with an M.S. in Operations Research and 5+ years working with complex clinical and genomic data at 3 academic medical centers. I combine HPC-scale pipeline engineering, observational study design, and hands-on ML to turn raw data into reproducible, publication-ready results, including co-authored peer-reviewed research.

Capabilities

Technical Skills

Full-stack data science from raw data to deployed model.

🐍
Languages
Python R SQL
⚙️
Data Science
Machine Learning Statistical Modeling Feature Engineering Hypothesis Testing
Big Data & Distributed Computing
PySpark HPC (SLURM) Parallel Processing
🐳
MLOps & Deployment
Docker Containerized Workflows Reproducible Pipelines
📈
Data Visualization & Applications
R Shiny Plotly Tableau
🛠️
Tools & Environment
Git Jupyter Notebook VS Code

Work

Portfolio Projects

Health data, clinical ML, large-scale EDA, and bioinformatics research.

01 — Data Pipeline Engineering - Clinical Analytics Platform

ISCVAM — Multi-Study Oncology Data Pipelines Cancer Research

Experimental scientists had no direct computational path to their own data, leaving finished experiments stuck in an analyst queue. I built the automated data layer that removed the delay: two containerized R pipelines handling end-to-end ingestion, quality control, and cell-type annotation across 198 studies (~6.3M cells). Orchestrated on Slurm job arrays with size-aware scheduling and per-step checkpointing, the pipelines convert raw sequencing output into standardized HDF5 datasets the platform reads directly. Platform accepted for presentation at AACR 2023.

R Pipeline design HPC-slurm Checkpointing Research

02 — Clinical ML · Epidemiology · Python

Osteoporosis Risk — Recall Asymmetry in Matched Case-Control Data

Why do models miss cases on a perfectly balanced dataset? In a 1:1 matched cohort of 61,026 osteoporosis patients, classifiers from logistic regression and linear SVM to tree ensembles and a feed-forward neural network miss 40–49% of cases, despite exact class balance and ROC-AUCs up to 0.79. After validating the cohort by reproducing a published hyponatremia odds ratio (3.98 vs. 3.97), I use diagnostic analyses to test whether the recall gap stems from model limitations or from properties of the matched observational data.

Classification Case-Control Design Conditional Logistic Regression Model Diagnostics XGBoost Neural Network Scikit-learn

04 — FDA Data · R · Interactive App

FDA Medical Device Harm Trends — RShiny Dashboard

Built an interactive visualization app over the 2016 MAUDE (FDA medical device passive surveillance) dataset. Users explore temporal harm trends across device categories and manufacturers. Demonstrates stakeholder-facing data product design.

RShiny FDA · MAUDE Time-series Dashboard R

05 — Public Health · Unsupervised Learning

COVID-19 Vaccine Adverse Symptoms — Association Rule Mining

Mined COVID-19 adverse-event reports from VAERS (CDC/FDA) using the Apriori algorithm to surface frequent adverse-symptom patterns and compare reported events between Moderna and Pfizer. Key insight: adverse-event patterns were broadly similar across both vaccines, suggesting perceived safety differences may be driven more by reporting frequency than by fundamentally different symptom profiles.

Association Rules VAERS · CDC/FDA Unsupervised Public Health Python

Background

About Me

I'm a data analyst with an M.S. in Operations Research and 5+ years at academic medical centers (MedStar Health, Moffitt Cancer Center, Huntsman Cancer Institute), working with clinical EHR data, FDA safety reports, and high-dimensional genomic profiles. My work spans statistical modeling, machine learning, and large-scale data pipelines. I focus on the questions that matter before modeling: what the data can support, where bias enters, and when an association isn't enough to act on. I own the full path from raw, messy sources to reproducible results and tools that non-technical stakeholders actually use.

Core competencies

  • Scale & Engineering Rigor: Engineered parallelized R/Python ETL and analysis pipelines on HPC/SLURM clusters with run-tracking checkpointing (6.3M+ cells across 198 studies).
  • Methodological Depth: Paired machine learning with advanced observational study design, including matched case-control pairing, conditional logistic regression, and Rubin’s rules for multiple imputation.
  • Oncology Genomics: Built analysis pipelines for single-cell RNA-seq, single-cell ATAC-seq, and multiome data—covering QC, dimension reduction, clustering, and annotation with Seurat and Signac—with additional work in spatial transcriptomics and bulk sequencing.
  • Real-World & Federal Data: Executed hands-on data retrieval, extraction, and modeling across structured and unstructured sources, including hospital EHR data, FDA MAUDE, and VAERS safety databases.
  • Scientific Impact: Co-authored peer-reviewed research in top-tier journals (Clinical Cancer Research and Cancer Research) and presented first-author platform work at AACR 2023.

Contact

Get in Touch