About
CS junior at Harvey Mudd College ('27). I build full-stack products, statistical and ML models, and AI workflow tooling. I care about the gap between what a model claims and what the data actually shows.
Background
I study Computer Science at Harvey Mudd College, where the curriculum emphasizes first-principles reasoning and experimental validation, habits that carry directly into how I build and evaluate models. Mudd's required core also means I've worked seriously across physics, biology, and engineering, which sharpens how I think about domain context. That cross-disciplinary grounding carries into my work in baseball analytics, where the statistical methods and domain context reinforce each other. A mixed-effects model for batting skill makes more sense when you understand what a BABIP actually measures.
Projects
Pitcher Injury Risk+: A multi-model MLB pitcher health platform built over Statcast 2015–2024 (3,249 pitchers, 205,911 pitcher-game rows, 76 features). Spans baseline classifiers, survival models, and an ERA+-style Injury Risk+ composite (mean 100 per season, YoY r = 0.583). Walk-forward CV AUC 0.571. Findings: injury history and 90-day cumulative workload dominate; chronic velocity decline ≥2 mph doubles observed injury rate (6% → 12%).
Patio: Solo-built social betting app for backyard games (Caps, Beer Pong, Beerball) where friends wager virtual caps. Centerpiece is a scipy house-odds engine (~760 LOC): recency-weighted player means, harmonic-mean team strength, ~4% house edge, push-proof x.5 lines. React 19 + Flask + Supabase Postgres. ~6,150 LOC, 20 routes, 6 pages. Pre-launch MVP; currently pivoting to React Native/iOS.
Claude Code OS: A personal developer operating system built on top of Claude Code. 48 custom skills (slash commands), structured project memory with per-project README indexes, and a multi-agent dev team (Engineer → QA → Optimization Reviewer → Bug Fixer convergence loop). Includes an interactive Project Dashboard for tracking active work across repos.
Batting Average Ability (BAA): A same-season skill-isolation metric for MLB hitters. Mixed-effects model with player random intercepts (ICC: 24.7% of batting average variance is between-player); CLR transform for batted-ball composition; PA-weighted fitting. Test R² 0.457 / MAE 0.022. Random forest tried and rejected on evidence (BABIP Test R² 0.293 vs. linear 0.372). Face validity: Arraez 2023 = BAA 133.2, Gallo 2023 = BAA 61.2.
NBA Shot-Value Model (with Michael O'Brien): 4-model ML comparison (LR / DT / RF / XGBoost) on 122,203 cleaned shots from the 2014–15 season. XGBoost AUC 0.6365, Brier 0.2296. Key finding: hard predictability ceiling at ~0.63–0.64 AUC; defender distance ratio is the top SHAP feature; calibrated probabilities produce trustworthy expected-value estimates by zone.
What I'm looking for
Roles in software engineering (backend, full-stack), ML/AI engineering (model development, inference, pipelines), data science (statistical modeling, analysis), or baseball analytics. Available for summer 2026 internships and full-time roles starting spring/summer 2027.
Contact
nseluga@g.hmc.edu · github.com/nseluga ↗ · linkedin.com/in/nseluga ↗