Teams Data Platform
A medallion pipeline that answers seven business questions — and refuses to answer where the data can't support one.
- Source rows
- 703,448 · 7 feeds
- Tests
- 390 · 5 gates
- Scale
- 10× rows → 1.9× time
PySpark · Aurora PostgreSQL · S3 Parquet · AWS EKS · AWS Lambda · Terraform · React · Jupyter
Built for Citi's coding workshop, against a provided brief and repository scaffold. The pipeline, serving layer, API and interface are my work — 43 commits.