Site Reliability Engineer
WFA Digital Insight
Supabase's commitment to reliability is evident in its creation of a dedicated SRE practice, and this role is central to that effort. As a Site Reliability Engineer, you will be tasked with making every engineering team more reliable, not by owning their infrastructure, but by establishing practices, frameworks, and feedback loops that let them own reliability themselves. This approach requires a unique blend of technical expertise, collaboration, and influence, making it an exciting opportunity for those who thrive in fast-paced, async environments.
Job Description
ABOUT SUPABASE Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. ABOUT THE ROLE Supabase manages millions of Postgres instances and is growing. We have strong teams across observability, release engineering, and incident management — and we're concentrating our reliability efforts into a dedicated SRE practice that ties the discipline together across the platform. You'll be embedded within Service Operations, and your primary job is to make every engineering team more reliable — not by owning their infrastructure, but by establishing the practices, frameworks, and feedback loops that let them own reliability themselves. You'll work across the org: sometimes setting the standard, sometimes pair-programming a fix, sometimes helping a team define their error budget, sometimes telling them it's exhausted. This role is ideal for someone who has a strong vision for how SRE should work and thrives in async, fast-paced environments where influence matters more than authority. WHAT YOU'LL OWN - Partner with service teams to define meaningful SLIs and SLOs grounded in customer experience, and build the error budget policies that turn them into engineering decisions - Own and evolve the Operational Readiness Review (ORR) process — conducting reviews for new services and major changes across observability, alerting, runbooks, capacity, and graceful degradation - Strengthen the incident-to-improvement pipeline: connecting postmortem findings to operational readiness gaps, identifying repeat failure patterns, and driving systemic fixes - Act as the reliability expert teams pull in for architecture reviews, failure mode analysis, dependency mapping, and resilience design - Identify and quantify operational toil across the org, and build or advocate for automation that eliminates it - Help teams design sustainable on-call practices: alert quality, escalation paths, runbook coverage, and noise reduction - Track and report on org-wide operational maturity, surfacing systemic gaps and driving remediation YOU MIGHT BE A GOOD FIT IF YOU - Have 7+ years of experience in SRE, production engineering, or reliability-focused roles, including experience shaping SRE practices and driving adoption across engineering teams - Have a software engineering mindset — you write code and build tools, not just configure them - Have hands-on experience defining and operationalizing SLOs/SLIs at scale, including error budget policies that actually influenced engineering decisions - Have deep experience with incident response, postmortem facilitation, and turning incident learnings into systemic improvements - Have worked with large-scale multi-tenant systems (bonus: managed database platforms or Postgres) - Are proficient with cloud infrastructure (AWS preferred) and infrastructure-as-code (Pulumi preferred, Terraform/CDK also acceptable) - Communicate clearly and persuasively — this role requires influencing without authority across a distributed org - Have experience in async or globally distributed teams - Are energized by making other teams more effective rather than being the one who fixes everything NICE TO HAVE - Experience with Kubernetes-based platform operations - Familiarity with OpenTelemetry, VictoriaMetrics, Grafana, or similar observability tooling - Experience building developer-facing reliability tooling (SLO dashboards, ORR frameworks, toil tracking, DORA metrics) WHAT WE OFFER - Fully Remote We hire globally. We believe you can do your best work from anywhere. There are no Supabase offices, but we provide a WeWork membership or co-working allowance you can use anywhere in the world.
- ESOP Every team member receives ESOP (equity ownership) in the company. We want everyone to share in the upside of what we’re building together.
- Tech Allowance Use this budget to set up your ideal work environment—laptop, monitor, headphones, or whatever helps you do your best work.
- Health Benefits Supabase covers 100% of health insurance for employees and 80% for dependents, wherever you are.
- Annual Off-Sites Once a year, the entire company gathers in a new city for a week of connection, collaboration, and fun. It’s a highlight of our year.
- Flexible Work We operate asynchronously and trust you to manage your own time. You know what needs to be done and when.
- Professional Development Every team member receives an annual education allowance to spend on learning—courses, books, conferences, or anything that supports your growth.
- ~400 team members - 60+ countries - 20+ languages spoken - Over $1B raised (including our $500M Series F) - 540,000+ community members We move fast, build in public, and use what we ship.
How to Stand Out
- Focus on demonstrating your ability to influence teams without authority, as this is a key requirement for the role.
- Be prepared to provide specific examples of how you have driven reliability efforts in previous roles.
- Highlight your experience with cloud infrastructure and infrastructure-as-code, as these are preferred skills.
- Emphasize your software engineering mindset and ability to write code and build tools.
- Showcase your experience with incident response and postmortem facilitation, and be prepared to discuss how you have turned incident learnings into systemic improvements.
- Consider creating a portfolio that demonstrates your reliability engineering skills and experience.
- Research the company culture and values to understand how you can contribute to and thrive in the organization.
This is a remote position listed on WFA Digital, the platform for professionals who work from anywhere. Browse more remote jobs across all categories.