Principal Site Reliability Engineer
2 мес. назад
260k–275k / yearCanadaWorldwideLeadHybrid
Lead site reliability engineering efforts for a mission-critical SaaS platform.
О компании
- Why Join • Work on a mission-critical SaaS platform used by global enterprises • Solve complex reliability challenges at scale • Influence architecture and engineering culture at a company level • Competitive compensation, benefits, and growth opportunities Security & Compliance This role requires compliance with ’s information security and privacy policies, including annual security training
Обязанности
- In this pivotal role, you will be instrumental in designing, building, and maintaining the shared infrastructure services and platforms that our product and application teams will depend on You will focus on creating reusable, reliable, and scalable solutions that abstract away complexity, enabling other teams to focus on their core business logic and deliver features faster in a multi-cloud environment Design and build core platform components and shared infrastructure services that other development teams will integrate with and leverage to deploy and operate their applications Architect, implement, and manage highly available and scalable Kubernetes platforms as a service for internal consumers Develop robust, internal-facing tools and automation for infrastr
Требования
- 1+ years of experience as a Principal SRE with a strong focus on building tools and services for other engineers Deep expertise with Kubernetes in production environments, particularly in providing it as a platform(i.e single tenant and multi-tenant deployment architectures) Strong programming skills in Go (Golang) and Python, with experience building robust, maintainable backend services and automation Extensive hands-on experience with at least one major Cloud Provider (AWS, GCP, or Azure); multi-cloud experience is a strong plus, especially in building abstractions over them Proven experience designing and implementing Event-Driven Architecture and message queuing systems (e.g., Kafka, RMQ, NATS) as shared services Solid understanding and practical exp