Head of Engineering - GPU Cloud
1 нед. назад
FranceEuropeHeadHybrid
Lead engineering teams for GPU Cloud business focusing on deployment and operation of GPU clusters.
Условия
- ✔ A rich and diverse product offering: offers over 100 public cloud products in IaaS, PaaS, and AI.
- ✔ A cutting-edge technical environment: provides modern infrastructures, including high-performance bare metal servers, to tackle exciting technical challenges.
- ✔ Commitment to responsible cloud: is dedicated to a more responsible cloud, with data centers powered solely by renewable energy since 2017, minimizing our ecological footprint and holding top-level certification.
Как откликнуться
- HR discovery call (30 min)
- Interview with to understand your technical skills and approach to the role (45 min)
- Technical interview with CTO to validate your expertise (1h)
- Interview with Management to deepen discussions and assess your fit with the team (45 min)
- HR interview to tour our offices and meet your future colleagues
Другое
- As our GPU Cloud business continues to scale, we are strengthening our engineering leadership to support the deployment and operation of increasingly large and complex GPU clusters.
- Your mission will be to lead our Support Engineering, HPC, and SRE teams , own key technology and architecture decisions, and ensure that our most strategic GPU infrastructure projects are successfully designed, delivered, and operated.
- You will also play a critical role in validating the technical feasibility of commercial proposals and ensuring that the commitments we make to customers can be delivered reliably on our sovereign cloud infrastructure.
-
- We work in a collaborative and international environment where the diversity of Scalers, combined with a strong culture of knowledge sharing, helps us bring ambitious projects to life.
- You will lead an organization of 14 engineers across two squads , each managed by an Engineering Manager reporting directly to you.
- As part of the broader GPU Cloud organization , reporting to the SVP GPU Cloud, you will work closely with GTM, Operations, Service Management, Product, and other engineering teams to build and operate large-scale AI and HPC infrastructure.
-
- Tasks
- Lead the Support Engineering, HPC, and SRE organizations, directly managing two Engineering Managers responsible for 14 engineers
- Own the technical strategy, architecture, and key technology choices for GPU Cloud infrastructure
- Review and validate the technical and service dimensions of strategic commercial proposals, ensuring commitments are realistic and deliverable
- Oversee the design, deployment, and operational readiness of new GPU clusters
- Drive the evolution of our cluster management, capacity management, automation, and operational capabilities
- Ensure the reliability, scalability, performance, and maintainability of our GPU infrastructure
- Provide technical leadership on complex AI and HPC infrastructure projects
- Build strong alignment between Engineering, GTM, Product, and Operations
- Develop the engineering organization through clear direction, effective delegation, coaching, and long-term team development
- Establish and maintain high standards of engineering rigor, operational excellence, and technical decision-making
- 10+ years of experience in infrastructure engineering, including significant experience leading senior technical teams and managers
- Proven experience designing, deploying, or operating large-scale infrastructure and compute clusters
- Expertise in GPU and/or HPC environments , ideally involving NVIDIA and AMD technologies
- Strong understanding of distributed infrastructure, cluster architecture, reliability, and production operations
- Experience with orchestration, provisioning, and observability technologies such as Kubernetes, Proxmox, Warewulf, Prometheus, and Grafana
- Strong knowledge of high-performance networking technologies such as InfiniBand, NVIDIA Spectrum-X, and Broadcom Tomahawk
- Experience with high-performance and distributed storage technologies such as Lustre, DDN, and VAST Data
- Ability to assess architectural trade-offs and translate complex technical constraints into clear engineering and business decisions
- Strong leadership skills with the ability to lead experienced engineers and Engineering Managers
- Ability to navigate complex technical, organizational, and business situations
- High level of rigor and a strong sense of ownership
- Excellent organizational, prioritization, and planning skills
- Strong communication and stakeholder-management abilities
- Ability to synthesize complex engineering topics and communicate them clearly to both technical and non-technical stakeholders
- Comfortable making decisions in fast-moving environments with high technical and operational stakes
- Hybrid work: We offer up to 3 days of remote work per week.
- Offices: Our offices are spacious, dynamic workspaces with bold design, conveniently located near public transport. Most of our offices feature outdoor spaces (terraces) and bike parking facilities.
- Dining : Our chef provides a healthy meal service at the headquarters, and breakfast is available across all our sites year-round. Scalers working from regional sites enjoy a Swile card for lunches.
- Well-being commitments: Whether it’s access to a gym, daycare places, or discounted services for caring services, is committed to supporting Scalers in maintaining a balanced life.
- International environment: With dozens of nationalities, offers a stimulating environment where English is as widely spoken as French.
- Career & Mobility: Our managers value internal mobility, and opportunities to transition to other entities within the Iliad Group are accessible to all Scalers.