Principal Engineer, Data Analytics Engineering
SanDisk
Jan 2019 - Current (7 years 8 months)
• Darwin: Designed and implemented low-latency, mission critical streaming data pipelines on GKE using Bitnami Kafka and Confluent Kafka Connect, processing approximately 500 GB of data daily to enable real-time prediction of die quality prior to wafer dicing in semiconductor assembly. Built cross-site Kafka replicator consumers to ensure high availability and data redundancy. Deployed and managed infrastructure components - including Kafka, ELK stack, and Redis on Kubernetes using Helm charts, ensuring scalable and reliable system operations. • Inspector: Designed and containerized a high-throughput, business-critical streaming data pipeline on Kubernetes (Google Kubernetes Engine), leveraging Docker to queue, process, and bundle over 50