Site Reliability Engineer (SRE) (m/w/d)
3 нед. назад
GermanyEuropeOnsite
pythonmongodbci/cdgithub actionsmonitoringautomationscalable architectures
Responsible for building and maintaining reliable, scalable cloud infrastructure to support enterprise-scale AI workloads.
Обязанности
- Build and operate reliable cloud infrastructure for production workloads
- Improve performance, scalability, and reliability of backend services
- Define SLOs, monitoring, alerting, and observability across the platform
- Drive incident response, root cause analysis, and postmortems
- Optimize databases, deployments, and CI/CD pipelines
- Automate infrastructure and operational processes
- Partner closely with engineering teams to improve developer experience and system reliability
Требования
- Experience operating production workloads on a major cloud platform
- Good Python skills and experience optimizing backend services
- Strong understanding of monitoring, observability, and incident management
- Knowledge of distributed systems, asynchronous processing, and scalable architectures
- Experience with databases at scale (MongoDB is a plus)
- Familiarity with CI/CD pipelines (GitHub Actions preferred)
- A passion for automation, reliability, and building systems that scale
Другое
- At , we're building the AI workforce for procurement. As a Site Reliability Engineer , you'll ensure our platform remains fast, scalable, and reliable as we grow. You'll work closely with our product engineering teams to improve infrastructure, automate operations, and build systems that support enterprise-scale AI workloads.
- Build the infrastructure powering one of Europe's fastest-growing AI startups. Work on high-scale, production-critical systems. Own reliability, performance, and developer tooling. Competitive compensation, meaningful equity, and exceptional teammates. 100% on-site in our Munich office, where we build together every day.