Vikas Sevak

Vikas Sevak

About

Detail

Data Engineer | Databricks | AWS Glue | PySpark | Python | SQL | Apache Spark | Amazon Redshift | ETL | Data Lake | Data Warehouse
Mumbai, Maharashtra, India

Contact Vikas regarding: 
work
Full-time jobs

Timeline


work
Job
school
Education
folder
Project

Résumé


Jobs verified_user 0% verified
  • Likedin Services Private Ltd
    Data Engineer
    Likedin Services Private Ltd
    Jun 2025 - Current (1 year 3 months)
    Data Engineer | AWS Glue | PySpark | Python | SQL | Apache Spark | Amazon Redshift | ETL | Data Lake | Data Warehouse | Data Pipeline Designed and delivered scalable Amazon S3 Data Lake and Amazon Redshift data warehouse solutions for enterprise analytics. Built production-grade ETL pipelines using Python, PySpark, SQL, AWS Glue, and Apache Spark for high-volume data ingestion, transformation, and processing. Improved SQL performance by 20% through query optimization, indexing, partitioning, and data model optimization. Developed scalable Spark processing frameworks to efficiently transform large datasets and improve pipeline throughput. Implemented data quality, validation, monitoring, and logging frameworks to ensure reliable and consis
  • IDFC FIRST Bank
    Data Engineer
    IDFC FIRST Bank
    Jun 2023 - May 2025 (2 years)
    Designed and optimized scalable Data Warehouse architectures and Star Schema data models for enterprise analytics. Developed production-grade ETL pipelines using Python, PySpark, Apache Spark, SQL, and Pandas to process large-scale datasets. Improved SQL query performance by 30–45% through query optimization, indexing, partitioning, and performance tuning. Built distributed PySpark and Apache Spark data processing frameworks for high-volume batch processing and data transformation. Optimized data ingestion, ETL workflows, and batch processing pipelines to improve throughput and reduce processing latency. Refactored SQL queries and Spark jobs to enhance scalability, maintainability, and overall processing efficiency. Implemented data validat
  • Nivoda
    Data Engineer
    Nivoda
    Mar 2022 - Feb 2023 (1 year)
    Developed and maintained scalable ETL pipelines using Python, PySpark, AWS Glue, and Apache Spark to process high-volume customer transaction data. Designed end-to-end data ingestion frameworks to extract data from Amazon S3, perform cleansing, transformation, validation, and load curated datasets into Amazon RDS (MySQL). Optimized PySpark applications through partitioning, caching, efficient transformations, and SQL tuning, improving pipeline performance and scalability. Implemented robust data quality frameworks including validation, reconciliation, exception handling, and logging to ensure reliable and accurate data delivery. Developed optimized SQL queries for data extraction, transformation, validation, and performance optimization acr
Education verified_user 0% verified
  • SIES School of Packaging  Packaging Technology Centre
    Postgraduate Degree, Packaging Science
    SIES School of Packaging Packaging Technology Centre
    Jul 2017 - Jun 2019 (2 years)
  • K
    Bachelor of Science, Chemistry
    KJ Somaiya Science Commerce Vidyavihar Mumbai
    Jun 2014 - May 2017 (3 years)
  • K
    Bachelor of Science, Chemistry
    K J Somaiya College of Arts Commerce
Projects (professional or personal) verified_user 0% verified
  • D
    Data Lake Implementation and Analytics
    Jun 2025 - Current (1 year 3 months)
    Developed a scalable AWS-based Data Lake solution to support enterprise analytics by designing and implementing Amazon S3 data lake architecture. Built and maintained ETL pipelines using AWS Glue and PySpark to ingest, transform, and load large volumes of data from multiple structured sources into Amazon Redshift and Amazon RDS. Designed and optimized database schemas to improve query performance and storage efficiency. Implemented distributed data processing using Apache Spark on AWS Glue jobs to process large datasets efficiently. Collaborated with business stakeholders and analytics teams to deliver reliable datasets for reporting and dashboarding through Amazon Athena. Ensured high data quality and pipeline reliability by implementing
  • D
    Data Warehouse Optimization with PySpark and SQL
    Jun 2023 - May 2025 (2 years)
    Worked on optimizing a large-scale enterprise data warehouse by designing scalable schemas and improving data processing performance. Developed and maintained ETL pipelines using PySpark, Apache Spark, Python, and SQL to process millions of records from multiple data sources. Implemented partitioning, indexing, and SQL query optimization techniques, reducing complex query execution time by 30–45%. Refactored ETL workflows and optimized batch ingestion pipelines, improving data processing efficiency, reducing latency, and increasing overall system throughput. Collaborated with Data Analysts, BI teams, and business stakeholders to identify performance bottlenecks and deliver scalable data solutions for analytics and reporting.
  • C
    Customer Transaction Data Pipeline
    Mar 2022 - Feb 2023 (1 year)
    Developed a scalable ETL pipeline using PySpark and AWS to process large volumes of customer transaction data. The pipeline ingested raw data from Amazon S3, performed data cleansing and transformations, and loaded the processed data into AWS RDS (MySQL) for reporting and analytics. Focused on performance optimization, data quality, and reliable batch processing.
This is a community-created genome.