Staff Software Engineer (Cloud Infrastructure)
1 мес. назад
USALeadOnsite
gpucloudinfrastructure
Responsible for advanced diagnosis, maintenance, and repair of high-performance GPU compute clusters to ensure uptime, reliability, and performance.
Обязанности
- We are seeking a highly skilled and motivated GPU Fleet Operations Engineer to join ’s Fleet Operations team. This role is focused on the advanced diagnosis, maintenance, and repair of high-performance GPU compute clusters, ensuring maximum uptime, reliability, and performance across our fleet.
- The ideal candidate will be hands-on with GPU rack-level troubleshooting and work closely with data center operations, engineering, and vendors to support cutting-edge infrastructure featuring the latest NVIDIA and AMD GPUs. This position plays a critical role in maintaining the health and scalability of ’s rapidly growing GPU fleet.
Будет плюсом
- Technical certification or Associate’s/Bachelor’s degree in Electrical Engineering, Computer Science, or a related field or demonstrated experience.
- Experience working directly with hardware vendors and escalations.
- Background in large-scale GPU fleet operations or hyperscale data center environments.
Условия
- Hybrid work schedule
- Industry competitive pay
- Restricted Stock Units in a fast growing, well-funded technology company
- Health insurance package options that include HDHP and PPO, vision, and dental for you and your dependents
- Employer contributions to HSA accounts
- Paid Parental Leave
- Paid life insurance, short-term and long-term disability
- Teladoc
- 401(k) with a 100% match up to 4% of salary
- Generous paid time off and holiday schedule
- Cell phone reimbursement
- Tuition reimbursement
- Subscription to the Calm app
- MetLife Legal
- Company paid commuter benefit; $300 per pay period
- Compensation will be paid in the range of $215,000 - $260,000. Restricted Stock Units are included in all offers. Compensation to be determined by the applicants knowledge, education, and abilities, as well as internal equity and alignment with market data.
- is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.
Другое
- is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join , you join a team that is building the future, faster.
- We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.
- We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.
- If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at .
- Automate deep-level diagnosis and troubleshooting of hardware faults within GPU racks and high-density compute systems.
- Develop software to troubleshoot and support GPU platforms including NVIDIA A100, H200, GB200, B200, B300 and AMD 350X / 355X.
- Execute component-level diagnosis and remediation for failed or degraded hardware.
- Partner with data center operations to manage and perform field-replaceable unit (FRU) repairs for GPUs, power supplies, cooling systems, interconnects, and networking hardware.
- Conduct post-repair validation, burn-in testing, torch testing, and NVIDIA NCCL testing to ensure system stability and performance.
- Implement and execute preventative maintenance procedures to improve fleet reliability and extend hardware lifespan.
- Perform firmware and BIOS upgrades across the GPU fleet.
- Maintain detailed documentation of maintenance activities, failures, and resolutions in ticketing and asset management systems.
- Develop and update standard operating procedures (SOPs) for troubleshooting, repair, and validation workflows.
- Collaborate with engineering, software, and data center operations teams to identify root causes of systemic failures and implement preventative solutions.
- Participate in a rotating infrastructure on-call schedule (about one week every 4–6 weeks) with daytime coverage and handoff to the Europe team.
- Ability to code in Golang
- Proven experience diagnosing and repairing high-density, rack-mounted compute hardware in production environments.
- Deep understanding of GPU architectures and hands-on experience with GPU-based systems.
- Experience supporting NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X series platforms.
- Familiarity with high-speed interconnects such as InfiniBand, NVLink, and RDMA over Converged Ethernet (RoCE).
- Strong Linux experience (Ubuntu, Rocky Linux, CentOS) using the command line for diagnostics and testing.
- Proficiency with GPU and system diagnostic tools such as NVIDIA DCGM and NVIDIA field diagnostic utilities.
- Experience working with enterprise server hardware, power delivery, and cooling systems.
- Strong analytical and problem-solving skills.
- Excellent communication and collaboration skills.
- Ability to work independently in a fast-paced data center or operations environment.