Data Center Infrastructure Software Engineer
2 нед. назад
USASeniorHybrid
software deploymentnetworkingautomation
Develop and automate AI infrastructure clusters in a hybrid role based in Bellevue WA.
Обязанности
- A well-funded, rapidly growing AI infrastructure company is building a next-generation cloud platform designed to power the full lifecycle of artificial intelligence. The organization is developing a comprehensive AI infrastructure, platform, and services portfolio that supports the full spectrum of AI workloads including large-scale compute, model training, fine-tuning, inference, and emerging agentic AI applications.
- Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization. Engineering teams are intentionally lean, highly collaborative, and AI-native, leveraging modern tooling and automation to build infrastructure capable of supporting the industry's most demanding AI workloads.
- We're seeking Data Center Software Engineers to lead the design, development, configuration, and automation of AI infrastructure clusters.
- Your responsibility begins once servers and racks are installed in the data center and extends through software deployment, networking, configuration, cluster bring-up, and automation, ensuring the platform is fully operational and ready for customer workloads.
- Develop infrastructure-as-code, automation, and provisioning systems for compute, networking, and storage.
- Deploy and optimize Kubernetes, container, and distributed computing platforms.
- Optimize GPU, networking, storage, and system performance for large-scale AI workloads.
- Troubleshoot complex issues across hardware, operating systems, networking, storage, and software stacks.
- Build reliability, observability, and operational excellence practices for mission-critical infrastructure
Требования
- 5+ years of experience designing, building, or operating large-scale Linux-based infrastructure.
- Hands-on experience with Kubernetes, containerization, and distributed systems in production environments.
- Experience with infrastructure-as-code and automation tools such as Terraform, Ansible, or similar frameworks.
- Strong experience operating cloud or datacenter-scale infrastructure.
Будет плюсом
- Experience with bare-metal provisioning and hardware lifecycle management platforms (e.g., MAAS, Ironic, xCAT, Foreman, or similar).
- Experience with IPMI, Redfish, PXE boot, and automated operating system deployment at scale.
- Experience managing GPU clusters in datacenter or cloud environments.
Условия
- Competitive base pay for Bellevue market
- Certain roles are eligible for additional rewards, including merit increases, annual bonus, and long term incentives. These awards are allocated based on individual performance
- U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays.
- Build the inference platform powering the next generation of AI applications.
- Work directly on large-scale model serving, GPU optimization, and production AI systems.
- Solve complex challenges around latency, throughput, reliability, and cost efficiency.
- Join early enough to influence architecture, tooling, and engineering practices.
- Collaborate with a highly experienced team building critical AI infrastructure from the ground up.
- Enjoy the ownership and technical impact of a startup environment backed by significant long-term investment.
Другое
- Location: Hybrid | Bellevue, WA Area Titles: Senior | Staff | Principal (multiple roles available)
- Hybrid role based in the Bellevue, WA area.
- Approximately three days per week in the office.
- Candidates elsewhere in the U.S. who are open to relocation are encouraged to apply.
- U.S. work authorization is required. Visa sponsorship is not currently available.