DevOps / Platform Engineer
1 мес. назад
United KingdomEuropeOnsite
gpuinfrastructureplatform
DevOps / Platform Engineer responsible for defining and maintaining infrastructure to support large-scale GPU compute and developer environments for a superlearning AI platform.
Другое
- We are recruiting a DevOps / Platform Engineer to join us in our London office.
- Our mission is to make first contact with superintelligence.
- We are creating a superlearner that discovers all knowledge from its own experience, from elementary motor skills through to profound intellectual breakthroughs. This superlearning capability - the ability to endlessly discover knowledge and skills, without relying on human data - will be driven by the world’s most powerful reinforcement learning algorithms.
- The superlearner is expected to rediscover and then transcend the greatest inventions in human history, such as language, science, mathematics and technology. If successful, this will represent a scientific breakthrough of comparable magnitude to Darwin: where his law explained all life, our law will explain and build all intelligence. Role brief As a foundational hire on the platform team, you'll be joining at a point where individual decisions shape how the platform is designed, not just maintained. This isn't a role where you inherit someone else's architecture, you'll be one of the people defining it.
- We're pushing these learning methods to a scale that hasn’t been tried before and we care deeply about the infrastructure that gets us there. You'll own the infrastructure that the mission depends on; keeping large-scale GPU compute reliable, developer environments fast and the whole platform humming.
- This is an opportunity for someone who wants to be a part of defining a new paradigm in AI and wants their work to directly speed up the experimentation and progress of our ambitious research.
- This is a broad role, and the list below is non-exhaustive. You’ll be empowered to shape the role how you’d like in order to best enable our mission.
- Kubernetes and containers: manage clusters and deploy internal tools
- GPU scheduling at scale : hands-on with tools like KAI scheduler and Kueue to keep thousands of GPUs busy and productive
- Making large-scale infrastructure resilient: get to the root of hardware failures and cluster-scale chaos, building the systems that prevent them from recurring
- Cloud infrastructure: experience on Google Cloud and other major providers
- Observability that actually helps: log management and monitoring with tools like QuickWit, Grafana or Datadog
- Developer environments people love: tooling like Tailscale, Workbrew and dev containers that make everyone's day-to-day faster
- Clean, well-crafted code: e.g. Python and Rust for tools and infrastructure that are a joy to build on
- We move fast, give people real ownership and trust engineers to make good judgment calls. We are a team that cares as much about doing this well as doing it fast.
- If you're interested in joining and think your skills would be better suited to a different position please express your interest to the Member of Technical Staff position.