Release Engineer
1 мес. назад
SeniorRemote
sredeploymentobservabilitymonitoring
Maintain and improve the reliability of deployment and release systems with SRE practices.
О компании
- was born-remote and open-source-first. We believe our globally distributed team is our secret weapon in building tools developers love.
- ~400 team members
- 60+ countries
- 20+ languages spoken
- Over $1B raised (including our $500M Series F)
- 540,000+ community members
- We move fast, build in public, and use what we ship. If it’s in your project, we probably use it in ours too. We believe deeply in the open-source ecosystem and strive to support—not replace—existing tools and communities.
Обязанности
- We're looking for a Release Engineer (SRE) to join our Release Engineering team (part of EngOps) — a production-operations expert who brings an SRE mindset to how ships and runs, making deploys safe, observable, and recoverable at scale.
- Release Engineering's scope has grown well beyond build-and-ship: we increasingly own the operational reliability of the systems that deploy and run . In this role you'll treat our deployment pipelines, pre-production signal, and the control plane itself as production systems — with SLOs, error budgets, and on-call ownership — and you'll be the person teams lean on when reliability is on the line.
- This is not a "gatekeeper" role. You'll make the reliable path the easy path: standardising how we deploy, instrumenting what we ship, and ensuring that when something breaks, we detect it quickly and recover quickly.
Условия
- Fully Remote We hire globally. We believe you can do your best work from anywhere. There are no offices, but we provide a WeWork membership or co-working allowance you can use anywhere in the world.
- ESOP Every team member receives ESOP (equity ownership) in the company. We want everyone to share in the upside of what we’re building together.
- Tech Allowance Use this budget to set up your ideal work environment—laptop, monitor, headphones, or whatever helps you do your best work.
- Health Benefits covers 100% of health insurance for employees and 80% for dependents, wherever you are. Your wellbeing and your family’s health are important to us.
- Annual Off-Sites Once a year, the entire company gathers in a new city for a week of connection, collaboration, and fun. It’s a highlight of our year.
- Flexible Work We operate asynchronously and trust you to manage your own time. You know what needs to be done and when.
- Professional Development Every team member receives an annual education allowance to spend on learning—courses, books, conferences, or anything that supports your growth.
Другое
- is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth.
- In this role, you'll:
- Own the reliability of 's deployment and release systems, and the control plane they run on, against clear SLOs and error budgets
- Turn pre-production into a trustworthy signal — standardizing and instrumenting today's fragmented, ad-hoc deployment workflows
- Drive disaster-recovery readiness, including making environments reproducibly deployable from scratch (untangling undocumented secrets, unclear configuration ownership, and circular service dependencies)
- Build and operate health and SLO monitoring for critical user flows, using synthetic testing to catch regressions before customers do
- Reduce mean-time-to-detect and mean-time-to-recover for deploy-related incidents — which account for a large share of our incident load
- Participate in on-call, lead blameless postmortems, and turn findings into runbooks, alerting, and automation that remove toil
- Improve deployment observability and auditability — a clear record of what shipped where, when, and by whom
- Document operational procedures — break-glass paths, access models, and runbooks — so reliability knowledge isn't tribal
- Define and track SLAs, SLOs, error budgets, and DORA delivery metrics — with meaningful alerting over noise
- Ensure deployments fail fast and safely when health checks degrade
- Harden access and break-glass workflows (e.g. scoped self-service) so the right people can act in an incident without unsafe workarounds
- Partner with product engineering and platform teams to align release practices with reliability and availability targets
- Have 5+ years in SRE, production operations, platform engineering, or release engineering
- Have operated production systems at scale and carried on-call for them
- Are fluent in SLAs, SLOs, error budgets, DORA metrics, and operational KPIs — and the observability tooling behind them (Prometheus, Grafana, Alertmanager, or similar)
- Have led incident response with tooling like (or PagerDuty / Opsgenie), run blameless postmortems, and driven down MTTD/MTTR
- Operate confidently on AWS (multiple accounts, IAM, VPC) in production
- Are comfortable with infrastructure-as-code (Pulumi, Terraform) and Kubernetes
- Script and automate to eliminate toil rather than absorb it
- Communicate clearly with both infrastructure specialists and product engineers
- Thrive in async, globally distributed teams
- Are comfortable navigating ambiguity and iterating toward better systems over time