Site Reliability Engineer (Monetization)
1 мес. назад
120k–160k / yearCanadaWorldwideRemote
Site Reliability Engineer role focused on monetization systems.
О компании
- is a global commerce company with robust tools and services to help developers solve the inherent challenges of the video game industry. From indie to AAA, companies partner with to help them fund, distribute, market, and monetize their games. Grounded in the belief in the future of video games, is resolute in the mission to bring opportunities together, and continually make new resources available to creators. Headquartered and incorporated in Los Angeles, California, operates as the merchant of record and has helped over 1,500+ game developers to reach more players and grow their businesses around the world. With more paths to profits and ways to win, developers have all the things needed to enjoy the game. For more information, visit .com .
Обязанности
- Own the application-level infrastructure of the Monetization domain: Helm charts, Terraform configurations, Kubernetes deployments, runtime configuration, and service-level networking and integrations
- Own the domain's observability: design and implement SLOs/SLIs, monitors, alerts, and dashboards for critical services on Datadog and OpenTelemetry-based tooling
- Help to set up and evolve CI/CD pipelines for domain services (GitLab CI, GitHub Actions), including deploy and rollback automation
- Perform capacity planning and performance tuning ahead of expected load - product launches, sales events, and regional rollouts - including load testing and performance regression investigation
- Run Production Readiness Reviews for new services and major changes; define and enforce what "production-ready" means for the domain
- Support domain incident response: assist with deep investigation of complex incidents, contribute to post-mortems, drive follow-up reliability improvements, and maintain runbooks
- Build domain-specific automation that reduces operational toil: runbook automation, deploy helpers, recurring operational scripts
- Maintain and drive a forward-looking reliability roadmap for the domain together with product engineering leads
- Participate in product team planning, refinements, and architecture reviews, bringing the reliability perspective before design decisions become expensive to change
- Co-author company-wide SLO/SLI, capacity, and operational standards together with the broader SRE team; contribute improvements directly to shared SRE-operated subsystems
- Participate in the SRE duty rotation, supporting developers across the company
Требования
- We are looking for a Site Reliability Engineer (Monetization) who is pragmatic, product-minded, and equally comfortable writing code and running production systems to join our Infrastructure department's SRE team . The best candidate will be someone who thrives in a fast-paced, highly collaborative, and exceptionally dynamic setting and is excited to own the application-level infrastructure and reliability of a high-traffic commerce domain end to end - from deploy pipelines and Kubernetes manifests to SLOs, capacity planning, and production readiness . Strong Kubernetes, observability, and software engineering skills are essential, along with experience in operating production services in a cloud environment (GCP/GKE or comparable) and partnering closely with product development teams . Th
- 3+ years of proven SRE, DevOps, or platform engineering experience: on-call or incident response duty, SLO/monitoring ownership, deploy pipeline and infrastructure work for production services
- Software development background: you have built and shipped backend services, not only operated them - comfortable reading application code during an investigation and writing production-quality automation in at least one language (e.g., Go, PHP)
- Hands-on Kubernetes experience: Helm, manifests, deploy strategies, debugging application-level performance and networking issues (GKE or another managed Kubernetes)
- Solid observability practice: building monitors, dashboards, and SLOs/SLIs on a modern platform (Datadog preferred; Prometheus/Grafana experience also relevant), familiarity with OpenTelemetry
- Infrastructure as Code exposure (Terraform/Terragrunt) for collaboration with platform teams
- GCP experience (IAM, networking, managed services)
- Experience building and maintaining CI/CD pipelines (GitLab CI and/or GitHub Actions)
- Programming/scripting proficiency sufficient to build automation and tooling (e.g., Python, Go, or Bash)
- Practical experience with incident response, post-mortems, and driving reliability improvements from incidents
- Strong collaboration and communication skills — this role works embedded with product development teams daily
- Experience in payments, fintech, e-commerce, or gaming — high-traffic transactional systems
- Kubernetes certifications
- Google Cloud Platform certifications
- HashiCorp certifications