Building model systems grounded in reproducibility, computational speed, and reliability, with the model-risk depth to keep them honest in production.
Experience
Andrew Davidson & Company · Data Scientist, Model Risk Management
Jun 2021 — Present
Lead risk strategy across model performance, data integrity, and public-facing applications; present strategy decisions annually to the Board of Directors
Designed and built an independent monitoring framework that detects bias, error dispersion, and drift in credit models across the aggregation stack; normalized multi-output residual diagnostics that are empirically estimated, overdispersion-aware, and rolled into a Mahalanobis trigger consistent with SR 26-2
Surfaced silent data and labeling errors hidden by aggregate performance views, driving corrections to production applications and data ingestion pipelines
Developed a multimodal GraphRAG pipeline that utilizes Nemotron 3 Nano Omni (30B-A3B), supports hybrid search, and runs exclusively on-premises to support the review of a 1000+ PDF/HTML documentation inventory
Challenge production model assumptions using open-source machine learning challenger models (e.g., Temporal Fusion Transformers) to benchmark performance and expose the limitations of each
Reduced processing times of legacy data pipelines from weeks to minutes using Redshift, Python, and RAPIDS
Developed an MLOps framework using Prefect to better implement CI/CD and continuous-training (CT) practices and improve model reproducibility
Developed default and severity models for the NonAgency product
Developed NLP pipelines that flag potential FINRA compliance violations in real time, with suggested remediations
Data Scientist
Automated extraction of financial information from scanned images using YOLOv3 and a custom OCR network replacing Tesseract (sole inventor, U.S. Patent 11,790,215)
Pre-trained and fine-tuned a BERT model on a proprietary dataset using self-supervised pre-training techniques outlined in the BERT paper
Created an NLP pipeline for extracting pertinent language from legal documents with limited labeled training data using XGBoost and Shapley values
Peng Lab, NCSU College of Veterinary Medicine · Bioinformatics Analyst
Dec 2017 — Jul 2018
Performed differential expression analysis and consulted with researchers on findings
Analyzed and managed high-throughput RNA sequence data
Education
Master of Science in AnalyticsNorth Carolina State University
2019
Master of Science in Biomedical SciencesDuke University School of Medicine
2016
Bachelor of Science in BiochemistryNorth Carolina State University