T

Taha Mazhar

About

Detail

Houston, Texas, United States

Timeline


work
Job
school
Education

Résumé


Jobs verified_user 0% verified
  • Aimpoint Digital
    Data Solution Architect
    Aimpoint Digital
    Sep 2023 - Current (3 years)
    • Architected and delivered enterprise-scale ELT data platforms across AWS and Azure using Airflow, AWS Glue, Azure Data Factory, and dbt Cloud, reducing development effort by 40% through reusable frameworks and standardized engineering practices. • Designed and implemented multi-petabyte Lakehouse architectures leveraging Delta Lake, Apache Iceberg, AWS S3, and ADLS Gen2, enabling ACID transactions, schema evolution, time-travel capabilities, and scalable analytics workloads. • Established real-time data integration frameworks using Debezium, Kafka Connect, and Apache Kafka, modernizing legacy batch processes and enabling sub-minute data availability across critical business domains. • Defined enterprise data governance standards utilizing
  • D
    Lead Data Engineer
    Datawise Data Engineering
    Oct 2020 - Aug 2023 (2 years 11 months)
    • Led the design and optimization of large-scale batch and real-time data platforms leveraging Apache Spark, Apache Flink, Kafka, and Amazon Kinesis, reducing processing latency by 40% across pipelines handling over 1 billion events daily. • Standardized enterprise analytics engineering practices by implementing dbt Core across Snowflake and BigQuery, delivering 60+ governed data models with automated testing, documentation, and version-controlled deployments. • Architected scalable data ingestion frameworks using Fivetran, Airbyte, PySpark, and Databricks, integrating 30+ enterprise data sources into a centralized Lakehouse platform for analytics and reporting. • Built and maintained streaming data solutions on Databricks and Spark Structu
  • D
    Senior Data Engineer
    Datacoral Inc.
    Feb 2018 - Sep 2020 (2 years 8 months)
    • Built and managed scalable ETL pipelines using Apache Airflow and PySpark for multi-terabyte datasets in Parquet, ORC, and Avro formats across structured and semi-structured data sources. • Led migration of on-premises Oracle and SQL Server data warehouses to Amazon Redshift and Snowflake using Terraform-based IaC provisioning, improving average query performance by 60% and cutting infrastructure costs by 35%. • Designed Lakehouse architecture on Delta Lake and Databricks to support both batch and streaming ingestion with ACID transactions, enabling reliable analytics across structured and unstructured data. • Implemented Apache Kafka and Amazon Kinesis real-time event streaming pipelines, processing 500M+ events daily for product analyti
  • T
    Data Engineer
    Tredence Inc.
    Mar 2015 - Feb 2018 (3 years)
    • Designed and maintained ETL pipelines for enterprise reporting and analytics using Informatica, Talend, and custom Python scripts running on Hadoop and Hive clusters. • Developed scalable batch data ingestion frameworks handling 50M+ daily records into PostgreSQL and MySQL data warehouses, with automated reconciliation and error handling. • Applied advanced SQL (window functions, CTEs, execution plan analysis) and Python (Pandas, SQLAlchemy) for complex transformations, data validation, and reconciliation. • Replaced legacy Hive batch jobs with Apache Spark and SparkSQL, reducing large-scale batch processing runtime by 35% and lowering cluster resource consumption. • Collaborated with data science teams to build clean, documented feature
Education verified_user 0% verified
  • A
    AWS Certified Solutions Architect
    Feb 2026 - Current (7 months)
  • D
    Databricks Certified Data Engineer Professional
    Apr 2018 - Current (8 years 5 months)
  • University of Sialkot
    Bachelors in Science
    University of Sialkot
    Feb 2011 - Mar 2015 (4 years 2 months)