Support Team Lead
1 мес. назад
USALead
splunkopentelemetryappdynamicsdatadogincident managementproblem managementchange management
Lead and manage a Site Reliability Engineering team focused on application support, ensuring performance, availability, and reliability of customer-facing applications.
Будет плюсом
- Bachelor's or Master's degree in Computer Science or related field.
- Certifications in ITIL, AWS, Azure, or GCP.
- Experience with Mulesoft, Postman, and API testing.
- Proficiency with Kubernetes
- Strong cloud-native networking knowledge.
Другое
- Coforge is seeking an experienced Site Reliability Engineering (SRE) Team Lead to guide our Application Support SRE function and manage a high‑performing team responsible for ensuring the performance, availability, and reliability of mission‑critical customer‑facing applications. This role combines hands‑on technical leadership with people management, process ownership, and operational excellence.
- Manage, mentor, and coach a team of Application Support SREs; support career progression and skills development.
- Oversee team performance, capacity planning, and staffing for a 24x7 support model.
- Serve as the senior escalation point during major incidents and high-severity events.
- Foster a culture of accountability, blameless postmortems, continuous learning, and operational excellence.
- Establish team OKRs, KPIs, and reliability goals aligned with business objectives.
- Own and mature SRE processes including incident management, problem management, change management, and service readiness.
- Lead major incident response, coordinate cross-functional teams, ensure communication excellence, and drive root cause analysis.
- Define and enforce SLOs, SLIs, and error budgets for supported applications.
- Implement preventative solutions and systemic fixes that reduce incident recurrence.
- Enhance observability practices across Splunk, OpenTelemetry, AppDynamics, Datadog, and similar tools.
- Improve dashboards, alerting strategies, and telemetry coverage.
- Provide insights and recommendations for reliability, scalability, and performance across AWS-hosted applications, Mulesoft APIs, and Kubernetes-based services.
- Collaborate with development and architecture teams to integrate SRE principles early in the lifecycle.
- Champion automation to reduce toil—CI/CD optimization, deployment improvements, self-healing mechanisms, and runbooks.
- Analyze logs, performance issues, and code behavior to support Tier 2/Tier 3 escalations.
- Recommend initiatives to expand Splunk automation and AI-driven insights.
- 5–8+ years in SRE, DevOps, or production engineering roles.
- 2–4+ years in a technical lead or people management capacity.
- Strong experience supporting AWS-based applications, microservices, or API-driven environments.
- Advanced troubleshooting skills.
- Hands-on experience with observability stacks (either opensource or splunk)
- Familiarity with ITIL and incident frameworks.
- A proactive, ownership-driven mindset.
- Strong communication and stakeholder management.
- Ability to lead during major incidents.
- Passion for continuous improvement and operational rigor.
-