s
suraj kumar
suraj kumar
About
Detail
.
United States
• Over 5+ years of experience in Data Engineer, including profound expertise and experience on statistical data analysis such as transforming business requirements into analytical models, designing algorithms, and strategic solutions that scales across massive volumes of data. • IT experience on Big Data technologies, Spark, database development. • Good experience in Amazon Web Service (AWS) concepts like EMR and EC2 webservices which provides fast and efficient processing of Teradata Big Data Analytics. • Experience and domain knowledge in various industries such as healthcare, insurance, retail, banking, media, and technology. Moreover, working closely with customers, cross-functional teams, research scientists, software developers, and business teams in an Agile/Scrum work environment to drive data model implementations and algorithms into practice. • Hands-on experience with AWS services, such as EMR, Kinesis stream, kinesis firehose, IAM, S3, AWS sage maker, EC2, route S3, RDS, Elastic Load balancer ELB, DynamoDB, Glue, SNS, SQS, Cloud formation and also have pretty good experience on Redshift spectrum and AWS Athena query services for reading the data from S3. • Hands on experience in MS SQL Server with Business Intelligence in SQL Server Integration Services (SSIS), SQL Server Analysis Services (SSAS), SQL Server Reporting Services (SSRS), Azure Cloud Technologies including Azure Database, Azure SQL, Azure Datawarehouse, Azure Data Factory (ADF), Azure Data Lake (ADL), Azure Databricks (ADB). • Integrated Kafka with Spark Streaming for real time data processing. • Strong experience with spark real time streaming data using Kafka and Spring boot API. • Experience in usage of Hadoop distribution like Cloudera and Hortonworks. • Excellent Experience in Designing, Developing, Documenting, Testing of ETL jobs and mappings in Server and Parallel jobs using Data Stage to populate tables in Data Warehouse and Data marts. • Have experience in Apache Spark, Spark Streaming, Spark SQL and NoSQL databases like HBase, Cassandra, and MongoDB. • Understanding of structured data sets, data pipelines, ETL tools, data reduction, transformation and aggregation technique, Knowledge of tools such as DBT, DataStage. • Experience in designing & developing applications using Big Data technologies HDFS, Map Reduce, Sqoop, Hive, PySpark & Spark SQL, HBase, Python, Snowflake, S3 storage, Airflow. • Expert in Migrating SQL database to Azure data Lake storage, Azure Data Factory (ADF), Azure data lake Analytics, Azure SQL Database, Data Bricks and Azure SQL Data warehouse and controlling and granting database access and migrating on premise databases to Azure Data Lake store using Azure Data factory. • Establishes and executes the Data Quality Governance Framework, which includes end - to-end process and data quality framework for assessing decisions that ensure the suitability of data for its intended purpose. • Proficiency in Big Data Practices and Technologies like HDFS, MapReduce, Hive, Pig, HBase, Sqoop, Oozie, Flume, Spark, Kafka. • Hands-on experience in setting up the Azure Data factory and creating the ingestion Pipelines to pull data to Azure Data Lake Store and Azure Blob Storage. • Experience in Data Ingestion projects to inject data into Data Lake using multiple source systems using Talend Big Data. • Excellent knowledge in Performance tuning in SQL Server and Azure SQL DB. • Data Engineering with a strong background in GCP Data post, GCS, Cloud functions, Big Table and Big Query. • Extensive experience in Relational Data Modeling, Dimensional Data Modeling, Logical data model/Physical data models Designs, ER Diagrams, Forward and Reverse Engineering, Publishing ERWIN diagrams, analyzing data sources and creating interface documents. • Designed and developed Data Marts by following Star Schema and Snowflake Schema Methodology, using industry leading Data Modeling tools like Erwin.
Contact suraj regarding:
work
Full-time jobs