H
Hareesh K
Hareesh K
About
Detail
Data Engineer
United States
● More than 5 years of experience as a Senior GCP Data Engineer with demonstrated professional working Experience in Big Data technologies like Ingestion, Data Modelling, Querying, Processing, Analysis, and Implementing Enterprise level Systems Spanning Big Data and Data Integration. ● Hands-on Experience on Hadoop Distribution Platforms Namely IBM Big Insights, Hortonworks, and Cloudera, and Cloud platforms GCP and AWS ● Snowflake is one such SaaS-based platform that powers the data cloud and offers an intelligent infrastructure, optimized storage, and an elastic performance engine. ● Played key role in Migrating Teradata objects into Snowflake environment. ● Experience tuning spark jobs for efficiency in terms of storage and processing. ● Hands-on experience working in GCP services like Big Query, Cloud Storage (GCS), Cloud Function, Cloud dataflow, Pub/sub, Cloud Shell, GSUTIL, Big Query, Data Proc, and Operations Suite. ● Worked with GCP services like Cloud Storage, Compute Engine, App engine, Cloud SQL, Cloud Functions, ● Cloud Run,Cloud Composer, Cloud Bigtable and Pub/Sub to process data for the downstream customers. ● Expertise in Big Data Technologies and Hadoop Ecosystems such as Pyspark, Spark-Scala, HDFS, GPFS, Hive, Sqoop, PIG, Spark SQL, Kafka, Hue, Yarn, Trifacta, and EPIC data sources. ● Solid Experience and understanding of Implementing large scale Data warehousing Programs and E2E Data Integration Solutions on Snowflake Cloud, AWS Redshift, Informatica Intelligent Cloud Services (IICS - CDI) & Informatica PowerCenter integrated with multiple Relational databases (MySQL, Teradata, Oracle, Sybase, SQL server, DBT) ● Good knowledge of Amazon AWS concepts like EMR and EC2 web services which provide fast and efficient processing of Big Data and Machine Learning Concepts. ● Hands-on experience in Building Data pipelines and Data marts using the Hadoop stack. ● Hands-on experience in Apache Spark creating RDDs and Data Frames applying Operations Transformation and Actions and concerting RDDs to Data Frames. ● Experienced in data processing like collecting, aggregating, and moving from various sources using Apache Flume and Kafka. ● Experience in writing REST APIs in Python for large-scale applications. ● Optimizing and tuning Snowflake's performance for efficient query execution. ● Extensive experience working with AWS Cloud services and AWS SDKs to work with services like AWS API Gateway, Lambda, S3, IAM, and EC2. ● Developed a data pipeline using Kafka and Spark Streaming to store data in HDFS and performed real-time analytics on the incoming data. ● In-depth understanding of Apache Spark job execution components like DAG, Executors, Task Scheduler, Stages, and Spark Steaming. ● Experience in Creating and executing Data Pipelines in GCP and AWS platforms. ● Hands-on Experience in GCP, Big query, cloud functions, and data proc. ● Strong Experience in Control-M Job Scheduler Tool, Apache Airflow, ESP, and D-series and monitored the jobs on a call base to close Incident tickets. ● Hands-on experience with Amazon EC2, S3, RDS, IAM, Auto Scaling, CloudWatch, SNS, Athena, Glue, Kinesis, Lambda, EMR, Redshift, DynamoDB, and other services of the AWS family. ● Expertise in using the CI/CD JENKINS pipeline to deploy the codes into production. ● Designed and developed the programming paradigm to support data collection and filtering processes in the data warehouse and Hadoop data mart. ● Worked on Dimensional Data modelling in Star and Snowflake schemas and Slowly Changing Dimensions (SCD). ● Deep understanding of cybersecurity, pen testing, and working with them to get approvals to deploy code into production. ● Expertise in working with Ab Initio for data integration and ETL software. ● Expertise with the Big-data database HBase and NoSQL databases MongoDB and Cassandra. ● Developed Python scripts for data modeling and import/export, and extensive experience in deploying, managing, and developing MongoDB clusters. ● Proficient in designing and developing ETL processes using Control-M and other ETL tools, such as Informatica to integrate data from multiple sources and support downstream analytics and reporting.
Contact Hareesh regarding:
work
Full-time jobs