Lead the facilities operations for Gimlet Labs' data centers and high-density AI infrastructure deployments.
Posted by employer 1 month ago
First seen on Joblaze 1 hour ago
Last verified on the company career page 1 hour ago
Skills & Technologies
What you'll build
Must have
Nice to have
Not disclosed in this posting: compensation, years of experience, work arrangement, visa sponsorship.
Joblaze summary
The Data Center Facilities Operations Lead at Gimlet Labs is responsible for managing the operational integrity of facilities that support high-density AI infrastructure, ensuring systems like liquid cooling and telemetry are effectively monitored and maintained. Key skills include expertise in data center MEP systems, liquid cooling operations, and familiarity with monitoring systems such as BMS and DCIM. This role is ideal for someone with a strong background in critical facilities engineering or operations, particularly in environments with liquid-cooled compute infrastructure. Gimlet Labs is in a growth phase, expanding its technology and infrastructure to meet increasing demands.
Joblaze insights
Quick facts
From the original posting
About the role
Gimlet Labs is seeking a Data Center Facilities Operations Lead to own the critical facilities operating model for Gimlet data centers and high-density AI infrastructure deployments. In this role, you will make sure the facility-side systems that support Gimlet's compute capacity are ready, monitored, maintained, and operating inside the required envelope.
You will focus on the infrastructure that keeps liquid-cooled AI systems healthy: facility water loops, CDUs, supply and return temperatures, flow, pressure, water quality, leak detection, alarms, heat rejection, power and cooling coordination, BMS/DCIM telemetry, maintenance procedures, and vendor repair workflows.
This role is well-suited for a critical facilities operator who understands data center MEP systems, liquid cooling, operational monitoring, and the discipline required to keep high-density compute environments stable as Gimlet scales.
What success looks like
In the first 12-18 months, you will:
Build the facilities operations model for current and future Gimlet sites, including operating standards, escalation paths, maintenance routines, acceptance criteria, and facility readiness gates.
Translate OEM and engineering requirements for liquid-cooled platforms into practical site operating envelopes for temperature, flow, pressure, water quality, alarms, and heat rejection.
Own monitoring and response for facility-side telemetry, including supply and return water temperatures, delta-T, flow, pressure, leak detection, CDU status, cooling capacity margins, and BMS/DCIM alarms.
Partner with colocation providers, facility vendors, OEMs, Site Managers, Data Center Technicians, Deployment Leads, and TPMs to ensure facilities are ready before new compute capacity is deployed.
Create and maintain MOPs, SOPs, EOPs, maintenance windows, runbooks, inspection routines, and incident response procedures for critical facilities and liquid cooling operations.
Coordinate preventive maintenance, repairs, and vendor response for CDUs, facility water loops, filters, valves, pumps, sensors, leak detection systems, chillers, dry coolers, CRAHs, and related infrastructure.
Lead facility-side root cause analysis for thermal, leak, power, cooling, monitoring, and environmental events, then drive durable corrective actions.
Build reporting that shows facility health, risk, readiness, capacity margin, recurring issues, open repairs, and operational trends across Gimlet sites.
Experience in data center facilities operations, critical facilities engineering, MEP operations, commissioning, or facilities maintenance
Experience operating liquid-cooled, high-density compute infrastructure
Familiarity with facility water systems, CDUs, heat rejection, leak detection, and water-quality controls
Experience using BMS, DCIM, EPMS, or similar systems to monitor and respond to facility conditions
The ability to create and execute operational procedures with strong attention to safety and reliability
Experience coordinating across site teams, colocation providers, OEMs, and facilities vendors
The ability to work in active data center environments and support urgent facilities escalations
Strong candidates may also have
Experience supporting GPU, HPC, or rack-scale liquid-cooled infrastructure
Experience with commissioning, integrated systems testing, site acceptance, or facility turnover
Familiarity with power distribution, UPS and generator systems, chilled water, dry coolers, CRAH/CRAC systems, or rear-door heat exchangers
Experience managing colocation obligations, service levels, maintenance windows, and vendor escalations
A track record of improving facility reliability through monitoring, preventive maintenance, and incident analysis
Solve hard problems.
Own meaningful work.
Build for production.
Help define what’s next.
Standard company text repeated across Gimlet Labs's postings is omitted here.