SRE Engineering Manager - GPU Cloud
1 мес. назад
FranceEuropeLeadHybrid
gpucloudsite reliability engineeringautomation
Lead the Site Reliability Engineering team focused on GPU Cloud infrastructure and AI initiatives.
Как откликнуться
- Discovery call with HR
- Technical interview with the HPC team to understand your technical skills and approach to the role
- Manager interview to validate your expertise
- Interview with an Engineering Manager / Head of Engineering to deepen discussions and assess your fit with the team
- HR interview and office visit to tour our offices and meet your future colleagues
Другое
- Our growth is driving us to strengthen our GPU Cloud team to support our expanding infrastructure and key AI roadmap initiatives.
- Your mission will be leading the Site Reliability Engineering (SRE) team in order to build, automate, and maintain a highly reliable, production-grade GPU cluster infrastructure powering our sovereign cloud.
- We work in a collaborative and international environment where the diversity of Scalers, combined with a spirit of sharing, helps bring new projects to life every day, advancing our ambitions together.
- You will be part of a team of 6 SREs within the GPU Cloud organization. The team focuses on critical AI and HPC infrastructure challenges, including automating key components of our stack and implementing support for modern GPU technologies.
- Lead and manage a team of 6 Site Reliability Engineers, supporting their career growth and technical execution
- Design and implement automated solutions for server lifecycle management across GPU clusters
- Design and implement observability, logging, and monitoring solutions for large-scale GPU clusters
- Plan, prioritize, and manage the technical development roadmap for the SRE team
- Collaborate and coordinate closely with software engineering, product, and cross-functional teams across
- Handle recruitment and career management for team members
- Maintain, scale, and optimize high-availability production systems under heavy load
- Participate in on-call rotations to ensure production reliability and fast incident resolution
- Strong experience managing engineering teams in high-constraint production environments
- Proven expertise with Kubernetes container orchestration
- Direct experience with cluster management and virtualization tools (Proxmox, Warewulf)
- Experience with monitoring, metrics, and observability stacks (Prometheus, Grafana)
- Exposure to modern GPU hardware ecosystems (Nvidia, AMD) and high-speed networking fabric (InfiniBand, Spectrum-X, Tomahawk)
- Knowledge of distributed and high-performance storage solutions (Lustre DDN, VAST)
- Strong engineering leadership and team management capabilities
- Technical rigor and high attention to detail in production-critical environments
- Ability to handle high-pressure operational situations and manage incident stress pragmatically
- Excellent communication skills with the ability to convey challenging messages effectively
- Collaborative mindset with a focus on empowering engineers rather than micromanaging
- Hybrid work: We offer up to 3 days of remote work per week.
- Offices: Our offices are spacious, dynamic workspaces with bold design, conveniently located near public transport. Most of our offices feature outdoor spaces (terraces) and bike parking facilities.
- Dining: Our chef provides a healthy meal service at the headquarters, and breakfast is available across all our sites year-round. Scalers working from regional sites enjoy a Swile card for lunches.
- Well-being commitments: Whether it’s access to a gym, daycare places, or discounted services for caring services, is committed to supporting Scalers in maintaining a balanced life.
- International environment: With dozens of nationalities, offers a stimulating environment where English is as widely spoken as French.
- Career & Mobility: Our managers value internal mobility, and opportunities to transition to other entities within the Iliad Group are accessible to all Scalers.
- 🚀 Why join the adventure?
- ✔ A rich and diverse product offering: offers over 100 public cloud products in IaaS, PaaS, and AI.
- ✔ A cutting-edge technical environment: provides modern infrastructures, including high-performance bare metal servers, to tackle exciting technical challenges.
- ✔ Commitment to responsible cloud: is dedicated to a more responsible cloud, with data centers powered solely by renewable energy since 2017, minimizing our ecological footprint and holding top-level certification.