Staff Site Reliability Engineer
2 нед. назад
USALead
site reliability engineeringinfrastructure managementcloud computingmonitoringautomation
Responsible for maintaining and improving platform infrastructure to handle billions of customer events daily with resilience and cost-efficiency.
Другое
- About the Role Our Platform Infrastructure team is the backbone of everything we do at , providing a resilient and cost-effective platform that seamlessly handles billions of events from over 100 million customers daily. We own everything from compute, persistence, and networking to observability and deployments. Joining our team offers a high-growth career opportunity to collaborate with some of the world’s most talented engineers in a high-performance, high-impact culture. As part of the Infrastructure and Platform organization, the Production Engineering Team is focused on delivering a fast and reliable platform that empowers engineers to deliver solutions quickly and safely. We build scalable systems that automate routine tasks so we can focus on other impactful effo
- Design and Deliver High-Impact Solutions: Design and implement systems that enhance reliability, observability, traceability, and incident management, ensuring the platform scales effectively
- Lead Strategic Initiatives: Take ownership of cross-team collaborations and drive impactful projects by providing technical leadership and guidance
- Partner Across Teams: Collaborate with engineers from AI/ML, Data, Platform, and Product teams to develop best-in-class services
- Partner with engineers from AI/ML, Data, Platform, Product, and other groups to deliver best-in-class services
- Establish Standards and Best Practices: Define and enforce production standards, processes, and tools to ensure operational excellence
- Champion Reliability Goals: Advocate for and implement SLIs, SLOs, and other reliability-focused metrics across the engineering organization
- Mentorship and Knowledge Sharing: Guide and mentor team members, fostering technical growth and helping to develop the next generation of engineering leaders
- Innovate and Inspire: Drive continuous improvement by bringing creative ideas and challenging the status quo
- 7+ years of experience in Production Engineering, Backend Engineering, SRE, DevOps or similar role
- Strategic visionary: Your strong technical background enables you to look beyond solving the immediate problem, planning for the future.
- Proficient Problem-Solver: Strong coding ability in at least one language (e.g., Golang, Python, Java, Typescript) with the capability to solve complex issues through code
- Track Record of Success: Demonstrated experience delivering medium to large-scale projects that drive meaningful improvements in platform reliability and scalability
- Reliability Expertise: Deep understanding of production reliability concepts, including SLIs, SLOs, and incident management
- Strong Communicator: Excellent verbal and written communication skills with the ability to influence and collaborate across technical and non-technical teams
- Fast-Paced Experience: Familiarity with working in dynamic, reliability-focused production environments (preferred)
- You'll get competitive perks and benefits , from health & wellness to equity, to help you bring your best self to work.
- The US base salary range for this full-time position is $180,000 - $240,000 annually + equity + benefits
- Our salary ranges are determined by role, level and location
- #LI-HB1
- By applying for this position, your data will be processed as per 's Privacy Policy .