Site Reliability Engineer – AI-first Platform
1 мес. назад
SerbiaEuropeSeniorHybrid
kubernetesawsautomation
Responsible for making AI agent platform production-ready, secure, observable, and scalable by operating, automating, and hardening systems.
Другое
- Operate and scale Kubernetes (EKS) clusters running AI and cloud-native workloads
- Support productionization of AWS AgentCore-based systems
- Design, build, and maintain CI/CD pipelines using GitLab CI/CD for applications and platform services
- Build and improve automated deployment workflows across environments
- Implement and maintain observability systems (metrics, logs, traces)
- Define and enforce SRE practices: SLIs, SLOs, alerting, incident response, postmortems
- Partner with development teams to ensure services are production-ready and operable
- Participate in on-call rotation and support incident resolution and root cause analysis
- Build infrastructure using Terraform for repeatable deployments
- Improve cost efficiency, performance, and reliability of the platform
- 5+ years of experience in SRE, DevOps, or Platform Engineering
- Strong hands-on experience with AWS and Kubernetes (EKS)
- Experience operating or supporting AWS AgentCore or similar AI/agent platforms
- Strong proficiency in Python for automation
- Solid understanding of AWS identity and access management (IAM), including SSO/Identity Center, roles, policies, trust relationships, SCPs, and cross-account access patterns
- Experience with Infrastructure as Code (Terraform preferred)
- Strong understanding of CI/CD using GitLab CI/CD
- Ability to quickly onboard into existing AWS environments and operate independently
- Experience working with on-call and incident management in production systems
- Experience with AI/LLM or agent-based systems
- Running stateful workloads on Kubernetes (PostgreSQL, Redis, etc.)
- Knowledge of RAG, vector databases, or knowledge systems
- Multi-region AWS architecture experience
- Cost optimization and capacity planning experience