D

David Kidwell

About

Detail

United States

Timeline


work
Job
school
Education

Résumé


Jobs verified_user 0% verified
  • Canoe Intelligence
    Senior Data Specialist
    Canoe Intelligence
    Apr 2024 - Current (2 years 5 months)
    • Engineered an automated real-time data pipeline for portfolio rebalancing using Python, PySpark, and AWS Glue, enabling seamless data integration, transformation, and reduced operational costs by 50% through optimized workflows. • Developed data ingestion and processing workflows to measure investor risk profiles, leveraging ETL pipelines and AWS Lambda. Used Pandas and NumPy to transform and clean data to increase processing speed and drive a 30% boost in customer conversion through accurate and standardized risk analysis. • Built a data integration solution for arbitrage analysis across public and private markets using PySpark and AWS Redshift, enabling seamless data flow, automated calculations, and improving annual returns by 2%, w
  • MasterClass
    AI/ML Data Engineer
    MasterClass
    Aug 2021 - Mar 2024 (2 years 8 months)
    • Launched an AI driven recommendation system using TensorFlow, PyTorch and OpenAI API, increasing user engagement with suggested lectures by 70% through behavior based insights and collaborative filtering powered by LLM. • Developed a high performance FastAPI service for retrieving high-K similar vectors with batch querying capabilities and optimizing RAG models for advanced Q/A system • Established robust data quality checks and observability using Great Expectations and Datadog, reducing pipeline failures by 40% and increasing data trust across teams. • Architected and deployed scalable ETL pipelines using Apache Airflow, Azure Data Factory and Spark, processing over 10TB of user interaction data monthly to enable personalized conten
  • D
    ETL Data Engineer
    Data.ai (Formerly App Annie)
    Sep 2020 - Aug 2021 (1 year)
    • Rollout and optimized core data processing logic using Java, R and MyBatis for objects mapping to SQL statements using XML, enabling seamless data integration of over 500TB of market • Migrated legacy big data infrastructure from Hadoop(MapReduce, HDFS) to Snowflake data warehouse. • Developed analysis services in Databricks to model mobile app performance, market trends and user behavior using ML algorithms such as ARIMA for time-series forecasting and XGBoost for regression. • Utilized PySpark and MLflow for distributed feature engineering and model lifecycle management, with Snowflake as the central data platform powering strategic insights for app owners. • Utilized Snowflake for automatic clustering and micro partitioning for fa
  • P
    Data Engineer
    PaloAlto Networks
    Oct 2015 - Aug 2020 (4 years 11 months)
    • Worked for conversational AI agents utilizing NLP to handle user requests raised 50k every day and performed tasks and FAQs based on business needs in the cybersecurity industry. • Built a data pipeline for a SaaS platform to handle 1 M+ daily events utilizing Apache Kafka and Apache Spark for frequent user profile updates. • Accomplished a robust CI/CD pipeline leveraging AWS EKS, GitHub Actions, AWS CodeBuild and AWS CodePipeline to automate an deployment and integration processes
Education verified_user 0% verified
  • Binghamton University
    M.S in Computer Science
    Binghamton University
    Oct 2013 - Sep 2015 (2 years)
  • Hartwick College
    B.S. in Data
    Hartwick College
    Aug 2009 - Jul 2013 (4 years)