Principal Site Reliability Engineer, Google Cloud
2 мес. назад
240k–250k USD / yearUSALeadHybrid
site reliability engineeringcloudinfrastructureplatform
Lead and define the reliability strategy for a SaaS platform, shaping architecture and operations to ensure performance and availability at scale.
О компании
- 's AI-powered identity platform manages and governs human and non-human access to all of an organization's applications, data, and business processes. Customers trust to safeguard their digital assets, drive operational efficiency, and reduce compliance costs. Built for the AI age, is today helping organizations safely accelerate their deployment and usage of AI. is recognized as the leader in identity security, with solutions that protect and empower the world’s leading brands, Fortune 500 companies and government institutions. For more information, please visit . Why This Role Matters ’s platform is mission-critical for our customers. As we scale globally, reliability, availability, and performance are not optional—they are core
Обязанности
- In this pivotal role, you will be instrumental in designing, building, and maintaining the shared infrastructure services and platforms that our product and application teams will depend on • You will focus on creating reusable, reliable, and scalable solutions that abstract away complexity, enabling other teams to focus on their core business logic and deliver features faster in a multi-cloud environment • Design and build core platform components and shared infrastructure services that other development teams will integrate with and leverage to deploy and operate their applications • Architect, implement, and manage highly available and scalable Kubernetes platforms as a service for internal consumers • Develop robust, internal-facing tools and automation for infr
Другое
- • 1+ years of experience as a Principal SRE with a strong focus on building tools and services for other engineers • Deep expertise with Kubernetes in production environments, particularly in providing it as a platform(i.e single tenant and multi-tenant deployment architectures) • Strong programming skills in Go (Golang) and Python, with experience building robust, maintainable backend services and automation • Extensive hands-on experience with at least one major Cloud Provider (GCP is a must); multi-cloud experience is a strong plus, especially in building abstractions over them. • Proven experience designing and implementing Event-Driven Architecture and message queuing systems (e.g., Kafka, RMQ, NATS) as shared services • Solid understanding and practical experience with CI/CD pi