Site Reliability Engineer (Baremetal) at npv labs | Torre

Site Reliability Engineer (Baremetal)

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Full-time

Legal agreement: To be defined

Compensation
USD90k - 150k/year
location_on
Remote (for United States residents)
Remote (for United Kingdom residents)
Shared by
Emma of Torre.ai
10 days ago

Responsibilities


tl;dr: SRE; post-acquisition profitable health/adtech; baremetal 10M+ rps k8s production; kubernetes internals, expansion and tuning plus some baremetal infra experience are required; remote, up to 150k USD base, we can talk higher figures and EU/UK/US employmentPulsePoint sits at the intersection of healthcare and adtech. We help brands and agencies interpret the hard-to-read signals across the health journey and unify these digital determinants of health with real-world data to produce the most dimensional view of the customer. We are 300+ and growing, post-acquisition business and one of the leading players in the US healthcare ad market.We're looking for a Site Reliability Engineer to join PulsePoint and ensure the reliability of the platform behind large-scale data and advertising systems that power business-critical services used every day across the company. The Platform Engineering team owns the foundation that lets engineering teams move quickly and safely – infrastructure supporting Kubernetes workloads, data platforms, developer tooling, observability, and production operations at scale. Unlike many cloud-only environments, the team owns the full lifecycle, from bare-metal hardware and networking to Kubernetes, observability, and developer experience, across multi-petabyte data systems and business-critical services used by multiple engineering organizations.You willHelp design, build, and operate the Kubernetes platform used across PulsePoint – architecture and lifecycle management.Own reliability, observability, and incident response across platform services.Build infrastructure automation and GitOps workflows to reduce operational toil.Work on networking, service connectivity, and platform security.Improve developer experience through self-service, reliable platform capabilities.Operate large-scale distributed systems running on bare-metal infrastructure.StackKubernetes, ArgoCD, Puppet, Terraform, OpenTelemetry, Prometheus, Alertmanager, Kafka, Redis, Ceph. Bare-metal and hybrid cloud/on-prem infrastructure. Experience with every technology is not required.RequirementsExperience operating production infrastructure at meaningful scale.Understanding of how distributed systems fail and recover.A preference for automation over repetitive operational work, and for simplifying systems rather than adding complexity.Ownership extending beyond the boundaries of a single component.Willingness to work 9am–6pm ET US hours (fully remote).We offerRemote work, high engineering bar and comfortable culture.Flat hierarchy with easy access to business, product, and operations.Enormous scale (10M+ peak rps) with real growth potential.Ownership and direct impact, you have room to shift focus as your interests evolve.Up to 150k USD salary, higher figures and EU/UK/US employment are negotiable.