Data Engineer
AMEX
Apr 2021 - Apr 2023 (2 years 1 month)
Designed and implemented scalable data pipelines and robust ETL solutions using PySpark and T-SQL, aligning with modern data architecture best practices and enhancing analytical capabilities. Migrated a 10TB on-premise data warehouse to Snowflake and an AWS S3-based Data Lake, optimizing data transfer from on-prem host to cloud and improving query performance while reducing annual infrastructure costs by 35%. Developed and deployed 15+ high-volume ETL/ELT pipelines using Apache Airflow and PySpark on AWS EMR to ingest over 500 million records daily, implementing buffering, compression, and backpressure strategies to prevent data loss and ensure reliability for business analytics and operational reporting. Established a robust data governanc