Hardware / Machine Design Engineer
1 мес. назад
USASeniorRemote
hardware designmachine designgpu servers
Lead the design of next-generation GPU server and rack infrastructure for AI infrastructure company.
Обязанности
- A well-funded, rapidly growing AI infrastructure company is building a next-generation cloud platform designed to power the full lifecycle of artificial intelligence. The organization is developing a comprehensive AI infrastructure, platform, and services portfolio that supports the full spectrum of AI workloads—including large-scale compute, model training, fine-tuning, inference, and emerging agentic AI applications.
- Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization. Engineering teams are intentionally lean, highly collaborative, and AI-native, leveraging modern automation and tooling to build world-class infrastructure capable of supporting the most demanding AI workloads.
- We're seeking a Principal Hardware / Machine Design Engineer to lead the design of next-generation GPU server and rack infrastructure. This role is ideal for an engineer who enjoys translating vendor reference architectures into highly optimized, production-ready systems that maximize performance, reliability, serviceability, and cost efficiency.
- As one of the earliest infrastructure engineers on the team, you'll own the physical design of GPU servers and rack-scale systems that form the foundation of a rapidly expanding AI platform. Working closely with GPU vendors, ODMs, data center engineering, networking, and operations teams, you'll drive hardware architecture from concept through deployment while influencing future infrastructure strategy.
- This is a highly visible individual contributor role with broad ownership and the opportunity to shape hardware standards for a rapidly growing AI infrastructure organization.
- Lead the design and configuration of GPU servers and rack-scale infrastructure using vendor reference architectures from NVIDIA, AMD, and emerging AI hardware providers.
- Partner closely with ODMs and OEMs to develop, validate, and optimize custom hardware platforms for production deployment.
- Own rack-level architecture, including mechanical layout, power distribution, thermal design, cable management, serviceability, and deployment standards.
- Evaluate new GPU platforms, server technologies, and hardware innovations to determine their suitability for next-generation AI workloads.
- Collaborate with data center engineering, networking, infrastructure, and operations teams to ensure seamless integration across facilities and production environments.
- Balance performance, reliability, manufacturability, scalability, and total cost of ownership when designing infrastructure solutions.
- Support hardware validation, qualification, and production readiness for large-scale deployments.
- Help establish engineering standards, hardware specifications, and design best practices as the infrastructure organization grows.
Требования
- Extensive experience designing GPU-based servers, rack infrastructure, or large-scale compute platforms.
- Deep knowledge of GPU server architecture and vendor reference designs, particularly NVIDIA and AMD ecosystems.
- Experience working directly with ODMs and OEM partners to design, customize, qualify, and manufacture production hardware.
- Strong understanding of rack-level infrastructure, including power, cooling, airflow, mechanical integration, cable management, and serviceability.
- Experience supporting hyperscale cloud, AI infrastructure, HPC, or large-scale data center environments.
- Ability to evaluate architectural trade-offs across performance, scalability, reliability, and cost.
- Comfortable operating as a senior technical leader within a fast-moving, high-growth engineering organization.
Будет плюсом
- Experience designing infrastructure supporting AI training clusters or large-scale inference platforms.
- Experience with liquid cooling technologies, high-density rack deployments, or advanced thermal management.
- Familiarity with multiple GPU generations and evolving accelerator technologies.
- Background designing hardware platforms for cloud providers, hyperscalers, or AI infrastructure companies.
- Experience influencing long-term hardware architecture and infrastructure strategy.
Условия
- Competitive base pay for Bellevue market
- Certain roles are eligible for additional rewards, including merit increases, annual bonus, and stock. These awards are allocated based on individual performance
- U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays.
- Design the hardware foundation powering one of the industry's most ambitious AI infrastructure platforms.
- Own architecture decisions spanning GPU servers, rack systems, and next-generation compute infrastructure.
- Work directly with leading hardware vendors and manufacturing partners to build production-scale AI systems.
- Join early enough to influence technical direction, engineering standards, and infrastructure strategy.
- Collaborate with an exceptional team solving some of the most challenging problems in AI infrastructure, high-performance computing, and hyperscale systems.
- Enjoy the autonomy and impact of a startup environment backed by substantial long-term investment and a bold vision for the future of AI.
Другое
- Location: Hybrid | Bellevue, WA Area Titles: Principal (multiple roles available) ; Staff level may be considered
- Hybrid role based in the Bellevue, WA area.
- Approximately three days per week in the office.
- Candidates elsewhere in the U.S. who are open to relocation are encouraged to apply.
- U.S. work authorization is required. Visa sponsorship is not currently available.