Machine Learning Platform Engineer
2 нед. назад
SwitzerlandEuropeHybrid
machine learningmodel deploymentmodel traininginferenceobservability
Build and operate ML infrastructure and platforms powering AI products, focusing on scalable, reliable systems for model training, evaluation, deployment, and inference.
Обязанности
- As an ML Platform Engineer, you will build the infrastructure and systems that power A1's AI capabilities.
- You will design and operate the systems behind the AI stack, from model training and evaluation to deployment, inference, observability, and continuous improvement.
- You will work closely with AI engineers, researchers, and product engineers to turn models into reliable, scalable, and cost-efficient production systems. You will build the platforms, tooling, and infrastructure that enable the team to experiment quickly and bring AI capabilities to production with confidence.
Другое
- There are over 5 billion users using basic applications today such email, notes, tasks that are not AI-native. Our mission is to build a proactive smart assistant for everyday users to bring intelligence to conversations, errands, organising and workflows, with minimal prompting.
- Our product focuses on achieving high reliability for long-running workflows, persistent context, and real-world task completion. The system must handle multi-step reasoning, interact with external tools, and remain reliable despite non-deterministic model behavior. Our objective is to help users complete tasks daily enjoyable with over ~90%* reduced time.
- Build and operate the ML infrastructure and platforms powering A1’s AI products
- Design systems for model training, evaluation, deployment, inference, and experimentation
- Build and optimise model serving and inference infrastructure for high-throughput and low-latency workloads
- Improve reliability, scalability, latency, and cost efficiency of AI systems
- Develop reliable pipelines for data preparation, training, evaluation, model release, and continuous improvement
- Build platforms and tooling that enable AI engineers and researchers to experiment, evaluate, and ship models faster
- Develop evaluation and benchmarking infrastructure to measure model quality, performance, and regressions
- Build production observability, monitoring, tracing, and alerting for AI/ML workloads
- Improve AI systems across reliability, scalability, latency, throughput, and cost
- Identify bottlenecks across the ML stack and continuously improve system performance
- Work closely with AI engineers, researchers, and product teams to turn evolving model requirements into production-ready infrastructure
- Python
- PyTorch / JAX
- LLM and ML serving infrastructure such as vLLM, SGLang, or TensorRT-LLM
- Cloud infrastructure
- Distributed systems
- ML/data pipelines and workflow orchestration
- GPU infrastructure and performance tooling
- Vector databases and retrieval infrastructure
- Strong software engineering fundamentals and experience building production systems
- Experience building ML infrastructure, platforms, or production machine learning systems
- Experience with model deployment, inference, evaluation, or data pipelines
- Strong understanding of distributed systems and system reliability
- Ability to write clean, maintainable, production-quality code
- Comfortable working in ambiguous, fast-moving environments
- Bias toward ownership, experimentation, and continuous improvement
- AI infrastructure reliably supports production workloads at scale
- Models can be trained, evaluated, deployed, and improved efficiently
- Inference systems deliver strong latency, throughput, reliability, and cost efficiency
- ML pipelines are reproducible, observable, maintainable, and robust
- Model and infrastructure regressions are detected quickly and diagnosed efficiently
- Common ML infrastructure capabilities become reusable platform primitives rather than being rebuilt for every AI product
- The AI stack can evolve rapidly as new models, architectures, and inference techniques emerge