Systems Reliability Engineer
8 г. назад
New ZealandWorldwideSeniorHybrid
automationscripting
Experienced AI-native Systems Reliability Engineer to support and improve engineering operations and infrastructure reliability.
О компании
- The Role: We're looking for an experienced, AI-native Systems Reliability Engineer to join our Engineering Team. This team works alongside our core development teams to support existing operations, with a mandate for continuous improvement and the goal to create systems that are as "hands-off" as possible. You'll work with our delivery teams throughout the development lifecycle to ensure systems delivered are scalable, robust, and supportable. You'll also work on optimising our software and infrastructure to improve reliability and decrease cost. You are agentic-first. You reach for AI to diagnose, script, automate and document as a matter of course, and you are at home operating within an AI-assisted delivery pipeline. There is an element of working on support issues arising in real time
Обязанности
- Working alongside core development teams to design cloud deployment architectures on Azure.
- Leading the provisioning of deployment environments - including all aspects of servers, services and networking - and communications with technical client stakeholders.
- Ensuring aspects of security, scalability, supportability and performance are implemented in solution architectures.
- Assisting development teams to build continuous integration and deployment pipelines using toolsets selected for the particular technology stack.
- Carrying out scheduled software releases in partnership with development teams where full automation has not yet been achieved.
- Configuring monitoring and alerting tools, and integrate into support desk software.
- Providing technical assistance to support ticket triaging and troubleshooting, while verifying the symptoms, impact and responsibility for further investigation.
- Actively participating in the diagnosis and remedy of reactive support issues during business hours and on a shared on-call after hours basis.
- Completing incident reports, including recommendations for ongoing mitigation. Implement mitigations in partnership with core development teams.
- Monitoring incident progress, and communicating status updates to clients.
- Designing and implementing proactive continuous improvement programmes for our solutions to minimise the occurrence of support issues.
- Automating first. When a task recurs, you build the automation rather than repeat it.
- Documenting, with AI-assistance, system architecture and configuration/ deployment processes; proactively sharing technical knowledge with support developers across multiple teams.
Требования
- Previous experience as a Site/Systems Reliability Engineer, DevOps Engineer or equivalent role.
- Agentic-first: you are very comfortable using AI to diagnose, script, automate and document, and to work within an AI-assisted delivery pipeline. This is central to the role.
- Expertise developing scalable web applications.
- Some proficiency in Rust, Java, .NET or Go. Experience in JavaScript, HTML, CSS, Node.js, ReactJS, AngularJS or similar front-end frameworks.
- Proven scripting experience using PowerShell, Ruby, Python, shell or other common scripting languages, including the use of AI tools to help construct these.
- Strong experience in Microsoft Azure deployment architectures, Microsoft DevOps deployment pipelines, GitHub Actions, and Terraform.
- System administration of Linux cloud-based server environments.
- Troubleshooting and configuration of virtual networks.
- Familiarity with configuration of New Relic, Dynatrace or similar monitoring framework, including building dashboards and alerting integration.
- Some experience with AWS and Bamboo/Bitbucket Pipelines tools would be an advantage.
- Ability to easily switch context across different technologies and complex environments.
- Ability to remain calm under the pressure of high-priority issue resolution scenarios.
- Good judgement about when to trust an AI-generated diagnosis and when to verify by hand.
- A strong desire to learn and up-skill.
- You must be Auckland based for this role. We offer a hybrid working model (working from home and working on-site), however office attendance is currently required three days per week on our dedicated teams in-office days, and occasionally as needed due to client and team requirements.