Sr Platform/ Infrastructure Engineer
1 мес. назад
ArgentinaLATAMSeniorRemote
kubernetespythonprometheuscephawsazure
Senior platform/infrastructure engineer role focusing on Kubernetes, cloud-native infrastructure, monitoring, and distributed systems operations.
Обязанности
- Design, deploy, and maintain production Kubernetes clusters and related services.
- Build and maintain automation and tooling using Python to support platform operations.
- Integrate and operate Prometheus for monitoring, alerting, and observability.
- Deploy and manage Ceph storage solutions for distributed workloads.
- Support platform modernization initiatives and migrate services to cloud-native patterns.
- Troubleshoot and resolve issues in distributed systems across compute, storage, and network layers.
- Collaborate with development, SRE, and operations teams to define platform requirements and SLAs.
- Document platform designs, runbooks, and operational procedures.
- Participate in on-call rotations and incident response to maintain platform availability.
Требования
- 5+ years of experience in platform, infrastructure, or site reliability engineering roles.
- Proven experience deploying and operating Kubernetes in production.
- Strong Python skills for automation, tooling, and operational scripts.
- Experience implementing and operating Prometheus-based monitoring and alerting.
- Hands-on experience with Ceph or similar distributed storage systems.
- Cloud experience with AWS and Azure (designing, deploying, and operating services).
- Demonstrated ability to troubleshoot distributed systems and resolve production incidents.
- Experience collaborating across teams to deliver platform improvements and migrations.
Будет плюсом
- Experience with OpenSearch.
- Proficiency with Bash scripting.
- Familiarity with Java-based services.
- Experience with Fluent Bit for log collection.
- Experience working with PostgreSQL.
Другое
- We are seeking a senior Sr Platform/Infrastructure Engineer to strengthen our platform team and drive cloud-native infrastructure initiatives. This role focuses on deploying and maintaining Kubernetes services, integrating monitoring and storage platforms, and troubleshooting distributed systems to ensure resilient, scalable operations.
- You will work with Python-driven tooling, Prometheus-based monitoring, Ceph-backed storage, and public cloud environments (AWS and Azure) to modernize and operate our platform. This is an opportunity to shape platform reliability and performance in a hands-on engineering role.