AI Training Infrastructure Engineer
1 мес. назад
USASeniorRemote
gpu
Engineer responsible for building and scaling distributed systems to support large-scale AI model training infrastructure.
Обязанности
- A well-funded, rapidly growing AI infrastructure company is building a next-generation cloud platform designed to power the full lifecycle of artificial intelligence. The organization is developing a comprehensive AI infrastructure, platform, and services portfolio that supports the full spectrum of AI workloads—including large-scale compute, model training, fine-tuning, inference, and emerging agentic AI applications.
- Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization. Engineering teams are intentionally lean, highly collaborative, and AI-native, leveraging modern tooling and automation to build infrastructure capable of supporting the industry's most demanding AI workloads.
- We're seeking AI Training Infrastructure Engineers to build and scale the distributed systems that power large-scale AI model training. This team focuses on reliability, efficiency, and operational excellence across GPU clusters, enabling researchers and engineers to train and deploy advanced AI models at scale.
- This is a foundational engineering role focused on building the infrastructure layer behind large-scale AI training workloads. You'll work on distributed training systems, GPU clusters, model pipelines, and the tooling required to make AI development more reliable, efficient, and scalable.
- You'll collaborate closely with infrastructure, orchestration, performance, and machine learning teams to solve complex challenges around distributed computing, fault tolerance, training efficiency, and production readiness.
- This opportunity is ideal for engineers who enjoy building highly scalable systems and working at the intersection of AI research, infrastructure engineering, and distributed computing.
- Build and scale distributed training infrastructure supporting large AI models across large GPU clusters.
- Design and improve systems that increase training reliability, efficiency, and resource utilization.
- Develop solutions for fault tolerance, checkpointing, recovery, and large-scale training operations.
- Integrate AI models into production training pipelines in partnership with platform, orchestration, and performance engineering teams.
- Diagnose and resolve issues impacting training throughput, stability, reliability, and cost efficiency.
- Build tools and automation that improve the developer experience for AI researchers and engineers.
- Establish best practices for training infrastructure, operational processes, and platform reliability.
- Contribute to the evolution of the AI infrastructure platform as an early member of the engineering team.
Требования
- Hands-on experience building and operating distributed training systems or large-scale machine learning infrastructure.
- Experience supporting large AI models, foundation models, post-training workflows, or similar ML systems.
- Strong understanding of the reliability, scalability, and efficiency challenges associated with multi-node GPU training.
- Experience integrating training systems with production machine learning pipelines.
- Strong programming skills and experience working with complex distributed systems.
- Ability to independently own technically challenging projects in a fast-moving engineering environment.
- Comfortable operating with high ownership and limited process overhead.
Будет плюсом
- Experience with distributed training frameworks such as PyTorch Distributed, DeepSpeed, Megatron-LM, Ray, or similar technologies.
- Experience with supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), or other post-training workflows.
- Background operating AI training infrastructure at scale within a hyperscaler, AI research organization, cloud provider, or GPU cloud environment.
- Experience optimizing GPU utilization, training performance, or distributed system reliability.
- Familiarity with Kubernetes, containerized AI workloads, and large-scale infrastructure platforms.
Условия
- Competitive base pay for Bellevue market
- Certain roles are eligible for additional rewards, including merit increases, annual bonus, and stock. These awards are allocated based on individual performance
- U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays.
- Build the infrastructure powering the next generation of AI models and applications.
- Work directly on distributed training systems, GPU clusters, and large-scale AI platforms.
- Solve some of the industry's most challenging problems around AI scalability, reliability, and efficiency.
- Join early enough to influence architecture, tooling, and engineering practices.
- Collaborate with a highly experienced team building critical AI infrastructure from the ground up.
- Enjoy the ownership and technical impact of a startup environment backed by significant long-term investment.
Другое
- Location: Hybrid | Bellevue, WA Area Titles: Senior and Staff (multiple roles available)
- Hybrid role based in the Bellevue, WA area.
- Approximately three days per week in the office.
- Candidates elsewhere in the U.S. who are open to relocation are encouraged to apply.
- U.S. work authorization is required. Visa sponsorship is not currently available.