AI Platform Engineer
WFA Digital Insight
Supabase is tackling a bold ambition: an AI‑native operating system that does more than assist—it carries a measurable share of the company’s workload. The Staff AI Platform Engineer sits at the centre of that vision, designing the execution layer that runs autonomous agents, logs every prompt, and enforces strict governance. What sets this role apart is the level of control built into the platform—risk‑tiered agents, immutable audit logs, and a human‑review gate that make dangerous actions structurally impossible. Candidates will own end‑to‑end architecture, from the event‑triggered queue to the evaluation suite that validates every run. It’s a rare chance to shape AI operations from the ground up in a fully remote, developer‑first environment.
Job Description
ABOUT SUPABASE Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. ABOUT THE ROLE We are hiring a AI Platform Engineer to build the execution layer for Supabase's internal AI systems. Supabase is building an AI-native internal operating system: a common way of working across the company where AI carries a meaningful share of the operational load rather than sitting alongside it as an assistant. We are standing up a new central team to build those systems, enable the teams, and embed AI operations throughout the organization. You are the engineer on that team. The execution layer is yours. You will build the platform that actually runs agents: an event-triggered queue, a headless model-agnostic runtime, durable state so work survives a restart, a human review gate, atomic rollback, and full logging of every prompt, tool call and decision so any run can be reconstructed. You will build the evaluation layer that makes any of it trustworthy, because an agent that cannot be measured cannot be trusted with anything beyond reading. And you will build the agents themselves, across everything from executive reporting down to a layer of agents that watch the platform and improve it. This is a governance-heavy environment by design, and that is the interesting part of the problem. Agents are risk-tiered from read-only internal data through to external-facing output, with review depth, evaluation requirements and human approval scaling by tier. Some capabilities are permanently off limits: an agent may read and report, it may carry a human-authored update into a system of record once a human consents, and it may initiate contact within a strict budget, but it may never autonomously write a commitment (an owner, a due date, a status) into a shared work system. Your job is to make that structurally impossible rather than merely forbidden. You will be the only engineer on this platform. You will close open architecture decisions yourself, own the infrastructure end to end, and instrument the system so it reports its own return. WHAT YOU'LL BE RESPONSIBLE FOR In this role, you'll: - Ship the agent platform to production. An event-triggered queue, a headless model-agnostic runtime (choosing the runtime is an open decision you will close), durable state that survives a failed run, a human review gate, atomic rollback, and complete run logging in the warehouse.
- Own the evaluation layer, and switch on the gate that depends on it. Golden suites with behavioral assertions rather than intuition, judge criteria with a written rubric, safety cases that must pass on every run, and a CI gate that blocks a regression from merging.
- Build and register the agent portfolio. Reporting, drafting, linting, triage and question-answering agents across the executive, team-lead and individual-contributor layers, plus a meta layer that observes the platform and improves it.
- Enforce governance in code.
- Design how the system contacts people.
- Own the platform tooling. The compiler and validator, inventory integrity, and the paths that distribute context and capabilities into the repositories and chat surfaces where work happens.
- Compute the operating measures from production data.
- Instrument the platform's own return. A ledger that logs the work each agent absorbs and computes the monthly figure, so the value of the system is a measurement rather than a claim.
- Design evaluations, not spot checks. You build golden sets, write behavioral assertions, define judge rubrics, set pass thresholds and gate CI on the result.
- Have done deep API work against the systems work actually lives in, and have authored MCP servers. You know the specific failure modes of those APIs, not just that they exist.
- Own infrastructure end to end in Python on GCP, with a cloud warehouse and infrastructure as code.
- Have taste about how software contacts humans. You treat every notification as spending a limited amount of trust. STRONG SIGNAL - Public work in this space. An open-source agent framework, an MCP server, an evaluation harness, or writing on agent reliability that other practitioners cite.
- LLM observability and cost instrumentation.
- You have built an internal platform that non-engineers adopted voluntarily, and can describe what you changed after watching them use it. WHAT SUCCESS LOOKS LIKE - The platform is boring. Agents run on a schedule and on events, state survives restarts, failed runs roll back cleanly, and every run can be reconstructed from its log.
- Nothing ships unevaluated.
- The dangerous action is impossible, not discouraged. An audit of any agent's credentials shows it cannot perform the writes it is not allowed to perform, and the audit log makes every consequential decision traceable.
- Teams pull the platform instead of being pushed.
- The platform reports its own value. The work absorbed is measured and published, so the case for expanding it is made with data rather than enthusiasm. WHAT WE OFFER - Fully Remote We hire globally.
- ESOP Every team member receives ESOP (equity ownership) in the company. We want everyone to share in the upside of what we’re building together.
- Tech Allowance Use this budget to set up your ideal work environment—laptop, monitor, headphones, or whatever helps you do your best work.
- Health Benefits Supabase covers 100% of health insurance for employees and 80% for dependents, wherever you are.
- Annual Off-Sites Once a year, the entire company gathers in a new city for a week of connection, collaboration, and fun. It’s a highlight of our year.
- Flexible Work We operate asynchronously and trust you to manage your own time. You know what needs to be done and when.
- Professional Development Every team member receives an annual education allowance to spend on learning—courses, books, conferences, or anything that supports your growth.
- ~400 team members - 60+ countries - 20+ languages spoken - Over $1B raised (including our $500M Series F) - 540,000+ community members We move fast, build in public, and use what we ship.
How to Stand Out
- Showcase concrete examples of event‑driven systems you built, including architecture diagrams or open‑source repos.
- Highlight any experience integrating AI models in production, especially model‑agnostic runtimes.
- Prepare to discuss how you have enforced security or governance controls in past projects.
- Bring a short demo or walkthrough of a monitoring dashboard you created for system observability.
- During interviews, ask about the team’s definition of “risk tiers” to demonstrate your understanding of governance.
- Emphasize remote‑work habits: async communication, documentation practices, and time‑zone coordination.
- When negotiating, consider equity and remote stipend as part of the total compensation package.
This is a remote position listed on WFA Digital, the platform for professionals who work from anywhere. Browse more remote jobs across all categories.