Site Reliability Engineer
Apply
Apply for this job directly on SHORTList.
Referral
Share your custom referral link for this job with qualified candidates. Earn the referral you lead to a hire.
Supabase manages millions of Postgres instances and is growing. We have strong teams across observability, release engineering, and incident management — and we're concentrating our reliability efforts into a dedicated SRE practice that ties the discipline together across the platform.
You'll be embedded within Service Operations, and your primary job is to make every engineering team more reliable — not by owning their infrastructure, but by establishing the practices, frameworks, and feedback loops that let them own reliability themselves. You'll work across the org: sometimes setting the standard, sometimes pair-programming a fix, sometimes helping a team define their error budget, sometimes telling them it's exhausted.
This role is ideal for someone who has a strong vision for how SRE should work and thrives in async, fast-paced environments where influence matters more than authority.
What You'll Own:
- Partner with service teams to define meaningful SLIs and SLOs grounded in customer experience, and build the error budget policies that turn them into engineering decisions.
- Own and evolve the Operational Readiness Review (ORR) process — conducting reviews for new services and major changes across observability, alerting, runbooks, capacity, and graceful degradation.
- Strengthen the incident-to-improvement pipeline: connecting postmortem findings to operational readiness gaps, identifying repeat failure patterns, and driving systemic fixes.
- Act as the reliability expert teams pull in for architecture reviews, failure mode analysis, dependency mapping, and resilience design.
- Identify and quantify operational toil across the org, and build or advocate for automation that eliminates it.
- Help teams design sustainable on-call practices: alert quality, escalation paths, runbook coverage, and noise reduction.
- Track and report on org-wide operational maturity, surfacing systemic gaps and driving remediation.
You Might Be a Good Fit If You:
- Have 7+ years of experience in SRE, production engineering, or reliability-focused roles, including experience shaping SRE practices and driving adoption across engineering teams.
- Have a software engineering mindset — you write code and build tools, not just configure them.
- Have hands-on experience defining and operationalizing SLOs/SLIs at scale, including error budget policies that actually influenced engineering decisions.
- Have deep experience with incident response, postmortem facilitation, and turning incident learnings into systemic improvements.
- Have worked with large-scale multi-tenant systems (bonus: managed database platforms or Postgres).
- Are proficient with cloud infrastructure (AWS preferred) and infrastructure-as-code (Pulumi preferred, Terraform/CDK also acceptable).
- Communicate clearly and persuasively — this role requires influencing without authority across a distributed org.
- Have experience in async or globally distributed teams.
- Are energized by making other teams more effective rather than being the one who fixes everything.
Nice to Have:
- Experience with Kubernetes-based platform operations.
- Familiarity with OpenTelemetry, VictoriaMetrics, Grafana, or similar observability tooling.
- Experience building developer-facing reliability tooling (SLO dashboards, ORR frameworks, toil tracking, DORA metrics).
What We Offer:
- Fully Remote: We hire globally. We believe you can do your best work from anywhere. There are no Supabase offices, but we provide a WeWork membership or co-working allowance you can use anywhere in the world.
- ESOP: Every team member receives ESOP (equity ownership) in the company. We want everyone to share in the upside of what we’re building together.
- Tech Allowance: Use this budget to set up your ideal work environment—laptop, monitor, headphones, or whatever helps you do your best work.
- Health Benefits: Supabase covers 100% of health insurance for employees and 80% for dependents, wherever you are. Your wellbeing and your family’s health are important to us.
- Annual Off-Sites: Once a year,
Apply
Apply for this job directly on SHORTList.
Referral
Share your custom referral link for this job with qualified candidates. Earn the referral you lead to a hire.