Data Center Facilities Operations Lead at Gimlet | Torre

Data Center Facilities Operations Lead

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Full-time

Legal agreement: To be defined

Provide your expected compensation while applying
location_on
Remote (for United States residents)
Shared by
Emma of Torre.ai
about 1 month ago

Responsibilities


About usGimlet is building the first multi-silicon neocloud designed for fast, efficient AI inference.We combine large-scale compute infrastructure with an execution platform that partitions AI workloads and maps each stage to the hardware best suited to run it.We work with foundation labs, hyperscalers, and AI-native companies, giving our team access to technical problems spanning frontier models, production infrastructure, and emerging hardware.About the roleGimlet Labs is seeking a Data Center Facilities Operations Lead to own the critical facilities operating model for Gimlet data centers and high-density AI infrastructure deployments. In this role, you will make sure the facility-side systems that support Gimlet's compute capacity are ready, monitored, maintained, and operating inside the required envelope.You will focus on the infrastructure that keeps liquid-cooled AI systems healthy: facility water loops, CDUs, supply and return temperatures, flow, pressure, water quality, leak detection, alarms, heat rejection, power and cooling coordination, BMS/DCIM telemetry, maintenance procedures, and vendor repair workflows.This role is well-suited for a critical facilities operator who understands data center MEP systems, liquid cooling, operational monitoring, and the discipline required to keep high-density compute environments stable as Gimlet scales.What success looks likeIn the first 12-18 months, you will:Build the facilities operations model for current and future Gimlet sites, including operating standards, escalation paths, maintenance routines, acceptance criteria, and facility readiness gates.Translate OEM and engineering requirements for liquid-cooled platforms into practical site operating envelopes for temperature, flow, pressure, water quality, alarms, and heat rejection.Own monitoring and response for facility-side telemetry, including supply and return water temperatures, delta-T, flow, pressure, leak detection, CDU status, cooling capacity margins, and BMS/DCIM alarms.Partner with colocation providers, facility vendors, OEMs, Site Managers, Data Center Technicians, Deployment Leads, and TPMs to ensure facilities are ready before new compute capacity is deployed.Create and maintain MOPs, SOPs, EOPs, maintenance windows, runbooks, inspection routines, and incident response procedures for critical facilities and liquid cooling operations.Coordinate preventive maintenance, repairs, and vendor response for CDUs, facility water loops, filters, valves, pumps, sensors, leak detection systems, chillers, dry coolers, CRAHs, and related infrastructure.Lead facility-side root cause analysis for thermal, leak, power, cooling, monitoring, and environmental events, then drive durable corrective actions.Build reporting that shows facility health, risk, readiness, capacity margin, recurring issues, open repairs, and operational trends across Gimlet sites.You may be a good fit if you haveExperience in data center facilities operations, critical facilities engineering, MEP operations, commissioning, or facilities maintenanceExperience operating liquid-cooled, high-density compute infrastructureFamiliarity with facility water systems, CDUs, heat rejection, leak detection, and water-quality controlsExperience using BMS, DCIM, EPMS, or similar systems to monitor and respond to facility conditionsThe ability to create and execute operational procedures with strong attention to safety and reliabilityExperience coordinating across site teams, colocation providers, OEMs, and facilities vendorsThe ability to work in active data center environments and support urgent facilities escalationsStrong candidates may also haveExperience supporting GPU, HPC, or rack-scale liquid-cooled infrastructureExperience with commissioning, integrated systems testing, site acceptance, or facility turnoverFamiliarity with power distribution, UPS and generator systems, chilled water, dry coolers, CRAH/CRAC systems, or rear-door heat exchangersExperience managing colocation obligations, service levels, maintenance windows, and vendor escalationsA track record of improving facility reliability through monitoring, preventive maintenance, and incident analysisWhy join now?Gimlet is expanding from its core technology into a production neocloud spanning new hardware, customers, and data centers.Solve hard problems.Own meaningful work.Build for production.Help define what’s next.