Sr. Manager, System Engeineering
3 нед. назад
Senior
systems architectureinfrastructure reliabilitydata center operationsincident response
Lead and manage a systems engineering team responsible for data center operations, infrastructure reliability, and systems orchestration at scale.
Обязанности
- Lead day-to-day engineering operations for data center and infrastructure systems, with accountability for reliability, maintainability, and execution quality.
- Manage and mentor a small systems engineering team while remaining technically engaged in architecture reviews, troubleshooting, and high-priority operational work.
- Own core infrastructure domains across compute, storage, networking, and orchestration, ensuring they work together as a coherent operating environment.
- Drive operational improvements in data center hardware reliability, installation standards, maintenance planning, and lifecycle management.
- Lead design and continuous improvement of systems and workflows related to provisioning, configuration management, and infrastructure automation.
- Coordinate infrastructure activities with internal partners, contractors, and external vendors to ensure design, deployment, and testing meet operational standards.
- Oversee technical incident response and follow-through, including root cause analysis, repair planning, and systemic reliability improvements.
- Establish and enforce strong engineering discipline around task tracking, documentation, and categorization of work across systems domains.
- Partner with adjacent teams to integrate shared infrastructure, improve visibility into system health, and support broader engineering goals.
- Contribute to hiring, interview processes, and team-building as part of the ongoing development of the systems engineering organization.
Будет плюсом
- Experience with low-level monitoring, profiling, or performance analysis tools.
- Experience with large-scale storage systems, hardware reliability programs, or infrastructure that supports high-throughput engineering workloads.
- Familiarity with modern infrastructure documentation practices and the use of automation to keep technical specifications and operational knowledge current.
- Experience supporting environments where infrastructure software and physical data center operations are tightly coupled.
- Familiarity with or other distributed database systems is a plus.
Другое
- The Infrastructure Systems Engineering team maintains and evolves ’s private infrastructure and data center environment, including systems that primarily support the company’s database test platform. That environment operates at a meaningful scale, running millions of tests per month and writing petabytes of storage data over time, while continuing to expand capacity and reliability requirements.
- We are seeking a Sr Manager, Systems Engineering to lead a team responsible for data center operations, infrastructure reliability, and systems orchestration. This role is intended to be a mostly technical manager with hands-on involvement: someone who can guide priorities, coordinate execution, and still dive into architecture, troubleshooting, incident response, and key implementation details when needed.
- The right candidate will help shape and operate core infrastructure spanning compute, storage, and networking, while improving the tooling, documentation, and operational discipline needed to scale the environment safely.
- 7+ years of experience in systems engineering, infrastructure software, platform engineering, or a closely related domain, including meaningful hands-on technical depth.
- 2+ years of experience leading engineers or technical teams in a manager, team lead, or equivalent player-coach capacity.
- Strong experience operating and improving infrastructure in at least several of the following areas: data center systems, storage platforms, network infrastructure, virtualization, and distributed systems.
- Demonstrated ability to troubleshoot complex failures that cross hardware, operating system, networking, and orchestration boundaries.
- Experience designing, implementing, or materially improving infrastructure automation, configuration management, or systems orchestration workflows.
- Hands-on experience with Terraform or comparable infrastructure-as-code tooling for provisioning, change management, and lifecycle management of infrastructure.
- Experience planning or executing the migration of data center services or infrastructure platforms onto Kubernetes, including consideration of networking, storage, deployment, and operational reliability.
- Strong working knowledge of data center networking, structured cabling, device installation, and infrastructure maintenance practices.
- Experience managing technical work through project plans, prioritization, and execution against timelines.
- Comfort working at both system level and implementation detail in at least one deep technical area such as storage, virtualization, kernel-level systems, high-performance computing, or service/platform architecture.
- Proficiency with at least one systems programming language such as C, C++, or Go, plus scripting in Python, Bash, or similar languages.
- Strong command-line fluency and comfort navigating and debugging Unix-like systems.
- Clear communication skills, sound technical judgment, and the ability to lead through collaboration rather than hierarchy alone.