Arif Mahmood

Arif Mahmood

About

Detail

Lead Data Engineer| AI/ML Engineer
New York, United States

Contact Arif regarding: 
work
Full-time jobs
Starting at USD175k/year

Timeline


work
Job
school
Education

Résumé


Jobs verified_user 0% verified
  • Meta
    Lead Data Scientist
    Meta
    May 2021 - Current (5 years 5 months)
    Spearheaded advanced data analysis and modeling initiatives, including fine-tuning LLaMA2 with SFT and DPO for a clinic management use case.
    Implemented advanced NLP algorithms and machine learning models powering multiple chatbot and conversational AI products.
    Led development of a diffusion model text-to-image (txt2img) project, applying cutting-edge data visualization and image processing techniques.
    Provided technical direction across development initiatives, enforcing coding standards, best practices, and delivery timelines.
    Led architectural discussions, coordinated cross-functional teams, and defined project requirements, milestones, and delivery plans.
    Monitored project progress continuously and im
  • Amazon
    Senior AI/ML Engineer and Data Scientist
    Amazon
    Apr 2019 - Apr 2021 (2 years 1 month)
    Built a demand forecasting platform using XGBoost and Prophet, improving inventory accuracy by 31% and reducing waste by $4.2M annually.
    Engineered a real-time fraud detection system using Isolation Forest and Deep Autoencoders, reducing false positives by 26%.
    Developed an NLP review analysis pipeline with spaCy and BERT, increasing unstructured data coverage by 45%.
    Modernized MLOps workflows using Docker, Amazon EKS, and SageMaker Pipelines, reducing deployment time from weeks to hours.
    Built automated testing for ML training and inference pipelines to ensure reproducibility and reliability.
    Created an active learning pipeline for customer intent classification, reducing manual labeling costs by 50%.
  • Perforce Software
    AI/ML Engineer & Data Scientist
    Perforce Software
    Mar 2016 - Mar 2019 (3 years 1 month)
    Engineered real-time fraud and risk prediction models using transaction and behavioral data, improving fraud detection recall by 37%.
    Built distributed data pipelines with Apache Spark, Apache Kafka, and Cassandra to process 2TB+ of financial data daily.
    Developed ensemble credit scoring models using XGBoost and LightGBM, exposed through scalable REST APIs.
    Architected cloud-based ML infrastructure on AWS and Google Cloud for automated model retraining and deployment.
    Implemented GitLab CI/CD, Docker, and model versioning, reducing deployment time from 2 weeks to 2 days.
    Built production data validation pipelines using Great Expectations to ensure data quality and model reliability.
    Deployed real-time in
Education verified_user 0% verified
  • Purdue University
    Master of Business Administration - MBA, Business Administration and Management, General
    Purdue University
    Jan 2017 - Jan 2019 (2 years 1 month)