S
Sahithi Siripuram
Sahithi Siripuram
About
Detail
United States
• 5+ years of experience designing, building, and optimizing scalable data pipelines with Apache Spark, Databricks, and Azure Data Factory. • Integrated structured, semi-structured, and unstructured data from relational databases, REST APIs, SFTP, and streaming platforms into cloud data lakes. • Developed robust ETL/ELT workflows for Azure and AWS, transforming raw data into analytics-ready datasets using PySpark, SQL, and Delta Lake. • Specialized in data modeling with Star and Snowflake schemas for business analytics, reporting, and machine learning. • Built and maintained complex Azure Data Factory pipelines with parameterization, dependency management, and dynamic datasets. • Optimized large-scale data processing in Apache Spark with partitioning, broadcast joins, and caching to reduce runtimes. • Created event-driven and real-time pipelines with Kafka and Event Hubs for healthcare and transactional data streams. • Managed and optimized Delta Lake tables for version control, schema enforcement, and ACID transactions in Azure Databricks. • Automated CI/CD for data pipelines using Azure DevOps and GitHub Actions. • Implemented data security: RBAC, data masking, HIPAA compliance with Azure Purview and Key Vault integrations. • Managed Databricks Workflows for scheduling, chaining notebooks, monitoring job runs, and auditing. • Developed metadata-driven frameworks to dynamically configure ingestion and transformation, improving scalability and reducing code duplication. • Built and curated datasets for Power BI and Tableau dashboards, supporting executive and operational reporting. • Collaborated closely with data scientists to deliver clean, feature-enriched datasets for ML use cases (fraud, risk, recommendations). • Led code reviews, sprint demos, and mentored junior engineers on data engineering and cloud best practices. • Engineered SCD Type 1 & 2 logic for robust enterprise data warehousing. • Established data quality monitoring with Great Expectations and rule-based validation, reducing downstream data issues. • Wrote efficient T-SQL, PL/SQL, and PostgreSQL queries for large-scale data extraction and profiling. • Designed ingestion frameworks for CSV, JSON,
Contact Sahithi regarding:
work
Full-time jobs