Site Reliability Engineer
3 дн. назад
AzerbaijanCISHybrid
kubernetesgcpgkecapacity planningslo
Responsible for owning application-level infrastructure reliability in a high-traffic commerce domain, focusing on Kubernetes, observability, and cloud production systems.
О компании
- ABOUT YOU We are looking for a Site Reliability Engineer who is pragmatic, product-minded, and equally comfortable writing code and running production systems to join our Infrastructure department's SRE team . The best candidate will be someone who thrives in a fast-paced, highly collaborative, and exceptionally dynamic setting and is excited to own the application-level infrastructure and reliability of a high-traffic commerce domain end to end - from deploy pipelines and Kubernetes manifests to SLOs, capacity planning, and production readiness . Strong Kubernetes, observability, and software engineering skills are essential, along with experience in operating production services in a cloud environment (GCP/GKE or comparable) and partnering closely with product developme
Обязанности
- Own the application-level infrastructure: Helm charts, Terraform configurations, Kubernetes deployments, runtime configuration, and service-level networking and integrations
- Own the domain's observability: design and implement SLOs/SLIs, monitors, alerts, and dashboards for critical services on Datadog and OpenTelemetry-based tooling
- Help to set up and evolve CI/CD pipelines for domain services (GitLab CI, GitHub Actions), including deploy and rollback automation
- Perform capacity planning and performance tuning ahead of expected load - product launches, sales events, and regional rollouts - including load testing and performance regression investigation
- Run Production Readiness Reviews for new services and major changes; define and enforce what "production-ready" means for the domain
- Support domain incident response: assist with deep investigation of complex incidents, contribute to post-mortems, drive follow-up reliability improvements, and maintain runbooks
- Build domain-specific automation that reduces operational toil: runbook automation, deploy helpers, recurring operational scripts
- Maintain and drive a forward-looking reliability roadmap for the domain together with product engineering leads
- Participate in product team planning, refinements, and architecture reviews, bringing the reliability perspective before design decisions become expensive to change
- Co-author company-wide SLO/SLI, capacity, and operational standards together with the broader SRE team; contribute improvements directly to shared SRE-operated subsystems
- Participate in the SRE duty rotation, supporting developers across the company
Требования
- 3+ years of proven SRE, DevOps, or platform engineering experience: on-call or incident response duty, SLO/monitoring ownership, deploy pipeline and infrastructure work for production services
- Software development background: you have built and shipped backend services, not only operated them - comfortable reading application code during an investigation and writing production-quality automation in at least one language (e.g., Go, PHP)
- Hands-on Kubernetes experience: Helm, manifests, deploy strategies, debugging application-level performance and networking issues (GKE or another managed Kubernetes)
- Solid observability practice: building monitors, dashboards, and SLOs/SLIs on a modern platform (Datadog preferred; Prometheus/Grafana experience also relevant), familiarity with OpenTelemetry
- Infrastructure as Code exposure (Terraform/Terragrunt) for collaboration with platform teams
- GCP experience (IAM, networking, managed services)
- Experience building and maintaining CI/CD pipelines (GitLab CI and/or GitHub Actions)
- Programming/scripting proficiency sufficient to build automation and tooling (e.g., Python, Go, or Bash)
- Practical experience with incident response, post-mortems, and driving reliability improvements from incidents
- Strong collaboration and communication skills — this role works embedded with product development teams daily
- Experience in payments, fintech, e-commerce, or gaming — high-traffic transactional systems
- Kubernetes certifications
- Google Cloud Platform certifications
- HashiCorp certifications