Platform Operations Engineer
WFA Digital Insight
Turquoise Health is building its Platform Engineering function from scratch, and the first Platform Operations Engineer will set the tone for everything that follows. The role sits at the intersection of infrastructure, reliability, and observability, giving the incumbent ownership of the Amazon EKS container fleet, GitOps pipelines with ArgoCD, and the monitoring stack that feeds data to engineering teams. Because the team is brand‑new, there is room to define standards for SLIs, on‑call rotations, and disaster‑recovery processes rather than inherit legacy practices. Candidates who enjoy greenfield projects, automation before they’re asked, and translating complex cloud networking into repeatable code will find a natural fit. The remote‑first culture means you’ll need discipline and clear communication, but you’ll also benefit from a flexible schedule and a budget to set up a home office.
Job Description
This is a fully remote role within the United States. Turquoise Health is building its Platform Engineering function from the ground up and this is your chance to shape it. As our first Platform Operations Engineer, you'll help lay the foundation for how we build, deploy, and operate infrastructure that supports real-world healthcare outcomes. This is a high-ownership role on a new team. You'll help establish the practices, tooling, and standards that the rest of engineering will rely on, from observability and reliability to deployment workflows and scalability. If you're energized by greenfield problems, comfortable navigating ambiguity, and motivated to automate and simplify before being asked, we'd love to build something great together.
Responsibilities
- Support the Platform Infrastructure - Help manage and scale our container environment on Amazon EKS, implement GitOps workflows using ArgoCD, and maintain CI/CD pipelines through GitHub Actions to ensure that deployments are fast, consistent, and automated - Build for Reliability - Define and track SLIs and SLOs, lead incident response including on-call rotations, root cause analysis, and post-mortems, and contribute to disaster recovery planning to keep our systems highly available - Drive Observability - Design and maintain our monitoring and logging stack using Datadog, Sentry, and CloudWatch — giving engineering teams clear visibility into system health and performance before problems reach users - Shape the Platform's Future - Collaborate on architectural decisions, build internal tooling and self-service workflows that make the platform easier to operate, and contribute meaningfully to how we scale and evolve our infrastructure WHAT YOU’LL BRING: - 3+ years in SRE, DevOps, or Cloud Infrastructure - Confident working with core AWS services (VPC, IAM, EKS, RDS) and a strong understanding of cloud networking and security best practices - Expert in using Infrastructure as code with Terraform, CloudFormation, or Crossplane - Proficient with GitHub and GitHub Actions as a core component of your CI/CD and automation pipelines- not just for source control - Experienced with running Kubernetes clusters in production and managing application deployments through GitOps workflows (ArgoCD/Flux) and Helm Charts - Proficient with observability tooling such as Datadog, Sentry, CloudWatch, Grafana to include building alerts, dashboards, and log pipelines - Experience writing solid Python scripts to glue systems together, automate infrastructure tasks, or handle custom workflows - Comfortable working independently in a remote setup, asking questions when needed, and keeping momentum without being micromanaged - Bachelor’s degree in Computer Science, Engineering, or equivalent experience.
Benefits
- Competitive pay with equity options - Stellar health care plan options (Medical, Dental & Vision), with FSA, DCFSA, & HSA options - Company-sponsored disability & life insurance - Unlimited PTO - 401(k) + 4% Matching - Fully remote work + flexible working hours - $750 work-from-home setup budget - Paid biannual in-person company summits - Quarterly $150 co-hanging stipend to meet up with coworkers - Monthly $100 health and wellness benefit - Generous paid family leave - Annual $1,200 learning & development stipend ABOUT TURQUOISE HEALTH Turquoise Health is a Series C price transparency platform for finance leaders across healthcare.
How to Stand Out
- Highlight concrete examples of Terraform modules or CloudFormation stacks you built; show the impact on reliability or cost.
- Prepare a short demo or walkthrough of a GitOps pipeline you set up with ArgoCD or Flux; interviewers love seeing the workflow live.
- Emphasize any on‑call experience, especially how you handled incidents, performed root‑cause analysis, and wrote post‑mortems.
- Include a brief portfolio of Python scripts that automate cloud tasks; attach snippets or a GitHub repo link.
- When discussing AWS expertise, reference specific services you secured (VPC, IAM policies, RDS) rather than generic cloud experience.
- Ask thoughtful questions about Turquoise’s disaster‑recovery strategy and how the platform team measures SLO success.
- Negotiate remote‑work allowances early; reference the $750 home‑office budget and ensure it covers your ergonomic needs.
This is a remote position listed on WFA Digital, the platform for professionals who work from anywhere. Browse more remote jobs across all categories.