S

Sai Yonas

About

Detail

Emmett, Michigan, United States

Timeline


work
Job

Résumé


Jobs verified_user 0% verified
  • Cortex
    Staff Data Engineer
    Cortex
    Apr 2021 - Current (5 years 7 months)
    • Designed and optimized large scale healthcare data platforms supporting batch, streaming, and hybrid workloads using Python, PySpark, Kafka, and Airflow.
    • Architected HIPAA compliant lakehouse and warehouse systems (Snowflake, BigQuery, Databricks, Delta Lake) enabling secure healthcare data integration and multi tenant isolation.
    • Built scalable ETL/ELT pipelines processing complex healthcare datasets including HL7, DICOM, EMR, lab, and clinical records.
    • Implemented end-to-end data governance, lineage, observability, and PII compliance frameworks across distributed cloud environments.
    • Developed healthcare focused data models and semantic layers powering BI dashboards, predictive analytics, and AI/ML clinical ins
  • C
    Senior Data Engineer
    Cambridge Analytica
    Nov 2018 - Mar 2021 (2 years 5 months)
    • Built scalable SaaS data pipelines for multi-tenant analytics platforms using Python, SQL, Airflow, dbt, Kafka, and cloud services (AWS, GCP, Snowflake).
    • Designed SaaS oriented data warehouses and lakehouse architectures supporting high velocity transactional and event driven data.
    • Developed analytics-ready datasets powering SaaS product analytics, BI tools (Looker, Tableau, Power BI), and customer insights.
    • Implemented CI/CD, data quality frameworks, and observability systems to ensure reliable SaaS data delivery.
    • Built real time and batch pipelines supporting usage tracking, product metrics, and customer behavior analytics.
    • Collaborated with product and engineering teams to convert SaaS requirements in
  • Jawbone
    Data Engineer
    Jawbone
    Apr 2016 - Sep 2018 (2 years 6 months)
    • Developed secure and scalable data pipelines for financial and transactional data using Python, SQL, Airflow, dbt, and cloud platforms.
    • Designed fintech grade data warehouses (BigQuery, Redshift, Snowflake) optimized for financial reporting, reconciliation, and analytics.
    • Built real time streaming pipelines using Kafka and Pub/Sub for transaction processing and event monitoring.
    • Created structured financial datasets supporting risk analytics, reporting dashboards, and business intelligence systems.
    • Implemented data validation, monitoring, and compliance focused data quality frameworks.
    • Worked closely with engineering and analytics teams to ensure accuracy and consistency of financial data flows.
    • C
Education verified_user 0% verified
  • not specified
    Bachelors In Computer Science
    not specified
  • Confluent
    Apache Kafka Certified Developer
    Confluent
  • dbt Labs
    dbt Fundamentals
    dbt Labs
  • HashiCorp
    Terraform Associate
    HashiCorp
  • Databricks
    Databricks Certified Data Engineer
    Databricks
  • Amazon Web Services
    AWS Certified Data Analytics
    Amazon Web Services
  • Google
    Google Professional Data Engineer (GCP)
    Google
Projects (professional or personal) verified_user 0% verified
  • P
    Personalized Recommendation Engine
    Developed batch and real-time pipelines to support AI-driven product recommendations using Spark, Kafka, Airflow, and BigQuery. Led 9 engineers in building feature stores and improving pipeline reliability by 35%.
  • M
    Multi-Tenant Data Platform
    Designed and implemented scalable pipelines for multi-tenant SaaS analytics, integrating API, event-stream, and log data into BI-ready datasets. Mentored 8 engineers on CI/CD pipelines, orchestration, and cloud architecture.
  • P
    Patient Outcomes Data Lake
    Created a centralized data lake for EMR, lab, and clinical data, applying NLP and OCR to extract structured insights for predictive modeling. Led 9 engineers on HIPAA-compliant pipelines and longitudinal patient-level datasets.
  • R
    Real-Time Transaction Analytics Platform
    Built and optimized real-time pipelines to monitor global transactions and prevent fraud using Python, PySpark, Kafka, Airflow, and BigQuery. Mentored 8 engineers while improving Spark job performance and pipeline reliability.