T

Tariq Faro

About

Detail

Florida, United States

Timeline


work
Job
school
Education
folder
Project

Résumé


Jobs verified_user 0% verified
  • Carbon Health
    Lead Data Engineer
    Carbon Health
    Feb 2022 - Current (4 years 6 months)
    • Built for clinical, claims, and HL7/FHIR a multi-cloud lakehouse using AWS plus GCP and Databricks plus Snowflake, reducing time to understanding from months to days with a 40% reduction. • Built real time pipelines with Kafka, Flink, and/or Spark Structured Streaming (exactly-once and CDC via Debez ium) with 99%+ uptime. • Standardized ELT with dbt and Airflow that included tests, docs, and exposures, resulting in onboarding new data feeds 35% faster and over 50% fewer defects. • Designed and implemented privacy through tokenization, KMS-backed encryption, least privilege IAM, and audit trails to ensure compliance regarding HIPAA for PHI. • Tuned Snowflake/Databricks using partitioning, clustering/Z-order, and file compaction to imp
  • Chewy
    Senior Data Engineer
    Chewy
    Jul 2017 - Dec 2021 (4 years 6 months)
    • Scaled ingestion and transformation on Databricks, Glue, and EMR with Airflow orchestration, reducing failure rate by over 50% and cutting cost by approximately 18% through Parquet layout optimization. • Built Kafka streaming pipelines for near real time marts (millisecond-level SLAs), including schema evolution and replay/backfill mechanics. • Built RAG-ready datasets with embeddings and vector indexes in Pinecone and FAISS to enable semantic search and agent workflows across product analytics. • Hardened multi-cloud deployments (AWS, Azure, GCP) with Terraform modules, secrets management, and envi ronment parity. • Collaborated with data governance teams to trace column level lineage using Atlas and DataHub, set data quality gates,
  • A
    Data Engineer
    Anblicks
    Jan 2014 - Jun 2017 (3 years 6 months)
    • Consolidated multiple disparate APIs and databases into a governed data lake and data warehouse (Star Schema, Snowflake, and SCDs), enhancing key dashboard performance by over 50%. • Developedreusable Spark ETLframeworkswithSLAmonitoring, automatedalerts, retries, and idempotent writes, substantially reducing on call noise and improving overall pipeline reliability. • Implemented GDPR and CCPA compliance controls, including pseudonymization, data retention policies, and standardized PII discovery across data pipelines. • Optimized SQL performance on Snowflake, Redshift, and Synapse through advanced partitioning, distribution, and clustering strategies, reducing compute utilization by 15–20%. • Authored comprehensive runbooks, contribu
Education verified_user 0% verified
  • B
    Bachelor's in Computer Science
    Apr 2009 - Aug 2013 (4 years 5 months)
Projects (professional or personal) verified_user 0% verified
  • D
    Dockerized CI/CD for data pipelines
    May 2022 - Dec 2025 (3 years 8 months)
    Implemented Dockerized CI/CD for data pipelines using GitHub Actions, enabling consistent testing, packaging, and release workflows across environments. Accelerated deployment cycles and minimized environment drift through containerized orchestration.
  • d
    dbt consumption models for Power BI
    May 2018 - Oct 2022 (4 years 6 months)
    Built dbt consumption models powering Power BI across finance and healthcare domains, standardizing metrics and reducing dashboard build time. Improved stakeholder trust and analytical consistency through version-controlled data transformations.
  • P
    Python + dbt ingestion framework into Snowflake
    May 2017 - Aug 2018 (1 year 4 months)
    Designed a Python + dbt ingestion framework into Snowflake replacing fragmented jobs, delivering BI-ready datasets with automated quality checks and lineage. Enhanced scalability and reduced manual maintenance through modular design and parameterized configurations.