L

Lakshmi Venkata Sailesh Vemula

About

Detail

Senior Data Engineer
Austin, Texas, United States

Timeline


work
Job
school
Education

Résumé


Jobs verified_user 0% verified
  • Amazon
    Data Engineer II
    Amazon
    Mar 2022 - Current (4 years 7 months)
    • Automated roster audits on 10M+ records daily by integrating Airflow, DynamoDB, Redshift, SNS, SQS, and APIs. Generated multilingual alerts that reduced compliance resolution time by 60% and served as a real-world telemetry system for global operations.
    • A 1TB/day ETL framework using AWS Glue, PySpark, and S3 brought data from RDS, Redshift, Lake Formation tables, and DynamoDB into an S3 Data Lake, powering a metrics data platform for 10K+ monthly users and removing 15,000+ hours of manual reporting work each month.
    • Introduced an attendance anomaly detection service with Python and Airflow, which automatically raised alerts (Emails, Calls, Messages), saving 5,260+ manual hours yearly and enabling HR to respond faster to staff
  • Root Insurance
    Data Engineer
    Root Insurance
    Jun 2021 - Mar 2022 (10 months)
    • Delivered a distributed AWS Redshift data warehouse with dimensional modeling, consolidating policy, claims, lifetime value, and telematics data into a single analytics source that accelerated enterprise reporting cycles.
    • Integrated APIs, SQL Server, and CSV feed into Redshift and Snowflake using Databricks while embedding Python- and SQL-driven data validation, maintaining 99.9% dataset accuracy for downstream analytics.
    • Reduced monthly infrastructure costs by $6K through migrating six high-volume workloads from SQL Server, DynamoDB, and S3 into Redshift with optimized distribution strategies.
    • Equipped executives with 20+ Tableau dashboards that refreshed hourly from Redshift, providing timely insight into underwriti
  • I
    Data Engineer - Intern
    Indiana Business Research Center
    Jan 2020 - Dec 2020 (1 year)
    • Developed a Microsoft SQL Server Master Data Management system with automated SSIS and Python ETL jobs, unifying pandemic-related data from multiple international health agencies for centralized reporting.
    • Migrated terabytes of healthcare records into an Azure Data Lake with optimized partitioning, cutting query times by 40% while ensuring HIPAA compliance with sensitive patient information.
  • Tata Consultancy Services
    Data Engineer
    Tata Consultancy Services
    Sep 2017 - Jul 2019 (1 year 11 months)
    • Automated data ingestion processes by building Python-based SSIS jobs, which reduced a 2+ hour manual workload to minutes and ensured on-time availability of operational datasets in the SQL Server and Power BI Dashboards.
    • Cut SQL Server query execution times by 60% through careful tuning of indexes, stored procedures, and partition strategies, enabling faster data retrieval for analytics users.
    • Consolidated multiple data sources, including flat files, APIs, and relational systems, into a centralized SQL Server data mart, streamlining reporting across diverse business units.
    • Introduced reusable Python modules for validation and transformation, improving dataset reliability by 25%, and reducing reprocessing caused by da
Education verified_user 0% verified
  • Indiana University Indianapolis
    MS in Applied Data Science
    Indiana University Indianapolis
    Aug 2019 - Dec 2020 (1 year 5 months)
  • Jawaharlal Nehru Technological University
    Bachelor of Technology in Electronics and Communication
    Jawaharlal Nehru Technological University
    Aug 2013 - May 2017 (3 years 10 months)
Projects (professional or personal) verified_user 0% verified
  • D
    Data Visualization of US Foreign Trade Statistics
    • Ingested and standardized US Census trade datasets across all states, ensuring consistency in commodity codes, units, and valuation of metrics for downstream analytics and economic research.
    • Built interactive dashboards using Python and Pandas to visualize trade volumes, top commodities, and value trends across geographies, implementing drilldowns and time-series analysis that accelerate stakeholder insights.
  • D
    Dashboard for World Population Statistics
    • Consolidated global demographic datasets from multiple open data sources into a structured SQL data warehouse and developed an interactive Tableau dashboard displaying population growth, density, and demographic distribution trends across continents and countries.
    • Incorporated time-series analytics and calculated KPIs to track population changes over decades, enabling comparative regional insights and data-driven demographic analysis.
  • P
    Predictive Analysis on Zomato Restaurant Data Using Machine Learning
    • Processed 1GB+ of restaurant data using Spark RDDs and Spark SQL, applying feature engineering to derive predictors for ratings and costs, and developed TensorFlow models with hyperparameter tuning for improved rating prediction accuracy.
    • Applied NLP sentiment analysis to customer reviews using GPU-accelerated training on SageMaker, reducing model runtime, and delivering actionable insights for business optimization.
This is a community-created genome.