Data Engineer at npv labs | Torre

Data Engineer

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Full-time

Legal agreement: To be defined

Compensation
USD90 - 150/year
location_on
Remote (for United States residents)
Remote (for United Kingdom residents)
Shared by
Emma of Torre.ai
10 days ago

Responsibilities


tl;dr: Data Engineer; post-acquisition profitable health/adtech; processing 20TB+ daily at near-realtime speed; Python, Spark; remote, up to 150k USD base, we can talk higher figures and EU/UK/US employmentPulsePoint sits at the intersection of healthcare and adtech. We help brands and agencies interpret the hard-to-read signals across the health journey and unify these digital determinants of health with real-world data to produce the most dimensional view of the customer. We are 300+ and growing, post-acquisition business and one of the leading players in the US healthcare ad market.We're looking for a Data Engineer to join a team, playing a key role at a technology company experiencing exponential growth. The data pipeline processes over 80 billion impressions a day (20TB+ of data, 220TB uncompressed), powering reports, budget updates, and optimization engines against extremely tight SLAs, with stats and reports delivered as close to real-time as possible.Team responsibilitiesInstall, maintain, and monitor Kafka, Hadoop, Presto, and RDBMS systems.Ingest, validate, and process internal and third-party data.Create, maintain, and monitor data flows in Hive, SQL, and Presto for consistency, accuracy, and lag time.Maintain and enhance the framework for jobs (primarily aggregate jobs in Hive).Build Kafka consumers using Spark Streaming for near-real-time aggregation.Train developers and analysts on data tools.Evaluate, select, and implement new tools.Handle backups, retention, high availability, and capacity planning.Review and approve DDL for databases, Hive framework jobs, and Spark Streaming to ensure standards are met.Participate in 24×7 on-call rotation for production support.StackAirflow, Docker, Graphite/Beacon, Hive, Impala, Kafka, Kubernetes, Presto, Spark Streaming, SQL Server, Sqoop.RequirementsBA/BS degree in Computer Science or a related field.5+ years of software engineering experience.Strong Spark expertise, Spark Streaming is highly desirable.Proficiency in Linux.Fluency in Python; experience with Scala/Java is a strong plus.Strong understanding of RDBMS and SQL.Willingness to participate in 24×7 on-call rotation.Nice to haveKnowledge and exposure to distributed production systems like Hadoop.Knowledge and exposure to cloud migration.We offerRemote work, high engineering bar and comfortable culture.Flat hierarchy with easy access to business, product, and operations.Enormous scale (80B+ impressions/day) with real growth potential.Ownership and direct impact, you have room to shift focus as your interests evolve.Up to 150k USD salary, higher figures and EU/UK/US employment are negotiable.