M

Mohammad Azmat Shaik

About

Detail

United States

Contact Mohammad regarding: 
work
Full-time jobs

Timeline


work
Job

Résumé


Jobs verified_user 0% verified
  • CommonSpirit Health
    Sr. Data Engineer
    CommonSpirit Health
    Jun 2024 - Current (2 years 3 months)
    • Designed secure and scalable ETL/ELT pipelines for data lakes and data marts using AWS Glue, SnapLogic, Talend, PySpark, ensuring alignment with conceptual and logical data models. • Configured AWS Lake Formation for fine-grained permissions and integrated Glue Crawlers, Catalog, and Registry to automate metadata management and enforce schema consistency. • Collaborated with data scientists to curate labeled datasets for AI/ML model training, implementing data quality checks and versioned transformations in Snowflake and PySpark, supporting advanced analytics and predictive healthcare use cases. • Orchestrated complex multi-step workflows with Apache Airflow, AWS Step Functions, Apache Flink, automating error handling, retries, and SL
  • Citibank
    Sr. Data Engineer
    Citibank
    Dec 2020 - Jul 2023 (2 years 8 months)
    • Designed and implemented Azure Data Factory and PySpark ETL/ELT pipelines to ingest, transform, and validate large-scale financial datasets into Azure Data Lake Storage and Synapse Analytics. • Built PySpark jobs leveraging DataFrame API and Spark SQL for secure, scalable risk reporting transformations, optimizing queries and storage formats (Parquet, Avro). • Orchestrated complex workflows using Apache Airflow, managing dependencies, error handling, retries, and SLA compliance for batch and streaming jobs. • Managed resource allocation and scheduling with Yarn on Azure Databricks and Hadoop clusters to ensure balanced utilization during peak loads. • Integrated Apache Kafka and Azure Event Hubs for real-time streaming ingestion of t
  • Ameriprise Financial
    Data Engineer
    Ameriprise Financial
    Jan 2019 - Dec 2020 (2 years)
    • Developed end-to-end 15+ ETL pipelines using Informatica PowerCenter and SSIS to extract, transform, and load financial data from legacy systems into relational databases, supporting enterprise reporting needs. • Built batch processing workflows on Cloudera Hadoop clusters with Apache Hive and Pig, transforming large-scale datasets for standardized reporting and analytics. • Created Spark-Scala jobs for distributed data processing, handling complex joins, filtering, and aggregations across multi-terabyte financial transaction datasets. • Migrated historical datasets from Oracle databases to AWS S3 as part of cloud modernization efforts, ensuring data integrity, security, and improved accessibility. • Integrated Kafka streaming ingest
Education verified_user 0% verified
  • University of Texas
    Master's in Computer Science
    University of Texas
  • M
    Microsoft Certified: Azure DevOps Engineer Expert (AZ-400)
  • M
    Microsoft Certified: Azure Developer Associate (AZ-204)