Senior Manager, Facilities Remote Operations Center
1 нед. назад
USASeniorOnsite
Lead management of remote operations center facilities for AI infrastructure company.
Требования
- 5+ years of experience in Data Center Operations, with direct responsibility for critical facility uptime.
- Strong working knowledge of MEP systems within data center environments, including electrical distribution, backup power (generators/UPS), and mechanical cooling systems (air- and liquid-cooled), specifically as applied in high-density, AI/HPC data center environments.
- Prior experience in a shift-based, 24/7 operations environment, including staffing and managing rotating shift schedules.
- Demonstrated experience with incident management, escalation procedures, and root cause analysis for critical facility events.
- Experience with DCIM, BMS/EPMS, or similar monitoring and alarm management platforms.
- Proven people leadership experience, including hiring, coaching, and performance management of operations teams.
- Strong communication skills, with the ability to translate technical facility issues into clear, actionable information for both technical and non-technical stakeholders.
- Ability to work on-site in Dallas, Texas, and to support a 24/7 operation, including availability for off-hours escalations.
Будет плюсом
- Experience specifically within hyperscale, colocation, or AI/GPU-cluster data center operations.
- Industry certifications such as CDCP, CDCS, CDCE, DCPRO, or equivalent.
- Experience standing up or scaling a remote/centralized operations function (NOC, GOC, or ROC) from the ground up.
- Familiarity with liquid cooling infrastructure (CDUs, manifolds, rear-door heat exchangers) common in AI-optimized data centers.
- Experience supporting multi-site, geographically distributed critical infrastructure portfolios.
- Bachelor's degree in Electrical Engineering, Mechanical Engineering, Facilities Management, or a related technical field, or equivalent practical experience.
Условия
- Competitive compensation and equity packages
- Restricted Stock Units
- Paid time off, paid holidays & leave of absence programs
- Comprehensive health, dental & vision insurance
- Employer contributions to HSA account
- Paid parental leave
- Paid life insurance, short-term and long-term disability
- Professional development & tuition reimbursement
- Mental health & wellness support
- Commuter benefits (parking & transit)
- Cell phone stipend
- 401(k) Retirement plan with company match up to 4% of salary
- Volunteer time off
- Global travel insurance & emergency assistance
- Daily meals allowance
- Additional perks & programs specific to location
- Compensation will be paid in the range of up to $175,000 -$200,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicants knowledge, education, and abilities, as well as internal equity and alignment with market data.
- is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.
Другое
- is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join , you join a team that is building the future, faster.
- We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.
- We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.
- If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at .
- is building the energy and infrastructure foundation for the AI era. We design, build, and operate purpose-built, AI-optimized data centers — from gigawatt-scale campuses like Stargate in Abilene, Texas, to modular and distributed deployments — powered by an energy-first approach that pairs compute with reliable, low-cost power. Our Facility Remote Operations Center is the nerve center for this footprint: a 24/7 command function that provides centralized visibility, monitoring, and rapid response across 's critical infrastructure nationwide.
- is seeking a Senior Manager to lead the Facility Remote Operations Center (ROC) based in Dallas, Texas. This role owns the people, processes, and technology behind 24/7/365 remote monitoring of critical facility infrastructure — electrical, mechanical, fire/life safety, and building automation systems — across 's national data center portfolio.
- The Senior Manager will build and lead a team of shift supervisors and operations specialists who serve as the first line of detection, triage, and escalation for facility events, working in close coordination with on-site Data Center Operations, Engineering, and Critical Facilities teams. This is a highly visible role that blends operational leadership, technical fluency in MEP (Mechanical, Electrical, Plumbing) systems, and program-building responsibility, as the ROC scales alongside 's rapidly growing AI infrastructure footprint.
- Lead 24/7/365 ROC operations, including staffing, scheduling, shift coverage, and performance management for a team of remote operations supervisors and specialists.
- Establish and continuously improve monitoring protocols, escalation procedures, and standard operating procedures (SOPs) for facility alarms, events, and anomalies across the fleet.
- Serve as the senior escalation point for critical facility events outside of normal parameters, coordinating real-time response with on-site teams, vendors, and leadership.
- Drive a culture of accountability, urgency, and continuous improvement within the ROC team.
- Maintain deep working knowledge of MEP systems within AI data centers, including electrical distribution (utility feeds, switchgear, generators, UPS, PDUs, busway), mechanical/cooling systems (CRAH/CRAC units, chillers, cooling towers, liquid cooling/CDUs, air handling), fire detection/suppression, and building management systems (BMS/DCIM/EPMS).
- Partner with Facility/Critical Infrastructure Engineering to understand system design intent, sequences of operation, and normal vs. abnormal operating conditions for each site.
- Ensure the ROC's monitoring platforms (DCIM, EPMS/BMS, ticketing, and alarm management tools) are correctly configured, integrated, and actionable for new and existing sites.
- Own incident management processes for facility-impacting events detected remotely, including detection, notification, escalation, and post-incident documentation.
- Lead or support root cause analysis (RCA) and after-action reviews for significant events, driving corrective and preventive actions.
- Develop and maintain emergency response runbooks and escalation matrices in partnership with site teams, ensuring readiness for weather events, utility disturbances, and equipment failures.
- Act as the connective tissue between remote monitoring and on-site Data Center Operations, Critical Facilities Engineering, Construction/Commissioning, and Security teams.
- Support commissioning and turnover of new sites into the ROC monitoring scope as 's footprint expands.
- Report on ROC performance, uptime-impacting events, and trends to senior leadership.
- Recruit, train, and develop ROC staff, building technical competency in MEP systems and monitoring tools.
- Build training programs and certification paths for remote operations specialists.
- Identify opportunities for automation, tooling improvements, and process standardization to improve detection speed and reduce false-positive alarm fatigue.
- A high-performing, technically capable ROC team that reliably detects and escalates facility events within defined SLAs.
- Reduced mean-time-to-detect (MTTD) and mean-time-to-escalate (MTTE) for critical facility events across the fleet.
- Well-documented, continuously improving SOPs and runbooks that scale as new sites come online.
- Strong, trusted working relationships between the ROC and on-site/engineering teams.