Data Reliability Engineer
WFA Digital Insight
Vytalize Health’s Data Reliability Engineer sits at the crossroads of data engineering and site reliability, a niche that few companies expose so explicitly. The role isn’t just about keeping pipelines running; it demands a proactive stance on observability, AI‑driven anomaly detection, and strict compliance with healthcare regulations. Candidates will be expected to define service level targets, build monitoring frameworks, and lead incident response for data quality issues that could affect patient outcomes. What makes this posting distinct is its blend of traditional SRE rigor with cutting‑edge AI tools for root‑cause analysis, positioning the engineer as a guardian of both reliability and regulatory integrity.
Job Description
Description of the Role The Data Reliability Engineer (DRE) at Vytalize Health is responsible for ensuring the end-to-end reliability, quality, and operational health of data across the full data lifecycle — from ingestion through downstream delivery and consumption. This role sits at the intersection of Data Engineering and Data Services, with a primary focus on building confidence that data is accurate, timely, observable, and dependable for both internal and external consumers. The DRE role applies Site Reliability Engineering (SRE) principles to data systems, emphasizing proactive monitoring, automation, failure prevention, and rapid recovery. This individual partners closely with Data Engineering, Data Services, DevOps, Product, and Analytics teams to define and enforce reliability standards, service levels, and operational practices for mission-critical healthcare data pipelines and data products. Given the sensitive and regulated nature of healthcare data, this role plays a key part in ensuring data reliability while maintaining strict compliance with security, privacy, and regulatory requirements. You will be metrics-driven — establishing clear reliability targets (SLIs, SLOs, SLAs) and measuring success through data quality, freshness, and delivery timeliness. Essential Functions of the Role Data Pipeline Reliability & Operations - Own and continuously improve the reliability of data pipelines across ingestion, transformation, and delivery layers, ensuring data is accurate, complete, and delivered on schedule.
- Establish and maintain data reliability standards, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs) for both upstream ingestion and downstream data delivery.
- Design, implement, and maintain comprehensive monitoring, logging, and observability frameworks for data pipelines, datasets, and data services with clear visibility into freshness, volume, schema changes, and data quality.
- Design and implement data quality testing and validation frameworks — establishing test cases, golden datasets, and regression tests to detect quality issues early.
- Establish data quality metrics and KPIs; measure and track data accuracy, completeness, timeliness, and consistency across pipelines.
- Lead incident response for data reliability issues, including detection, triage, communication, root cause analysis, and post-incident remediation with documented corrective actions.
- Drive improvements in pipeline resiliency through retry strategies, backfills, idempotency, schema enforcement, and safe deployment practices.
- Implement and optimize AI-powered root cause analysis tools and LLM-assisted incident investigation workflows to accelerate detection and resolution of data reliability issues.
- Use AI-assisted development tools (e.g., Claude Code, GitHub Copilot, or similar) to accelerate development of monitoring frameworks, runbooks, and incident response automation.
- Establish patterns and best practices for integrating AI-driven observability into data systems while maintaining explainability and human oversight of critical alerts and decisions.
- Partner with Data Services to ensure downstream data delivery mechanisms (APIs, flat files, service-based access, event-driven integrations) meet defined reliability and performance expectations.
- Collaborate with DevOps and platform teams to improve infrastructure reliability supporting Databricks, cloud storage, and data delivery services.
- Work with quality assurance and testing teams to establish data quality testing standards and validate pipeline outputs.
- Advocate for a culture of data ownership, operational accountability, and continuous improvement across data teams through documentation, knowledge sharing, and mentorship.
- Support capacity planning and scaling efforts by analyzing pipeline performance, usage patterns, and failure modes to identify infrastructure and architectural improvements.
- Maintain comprehensive documentation of reliability standards, SLAs, incident runbooks, and observability architecture for both technical and non-technical stakeholders.
- Demonstrated experience improving reliability, observability, or operational quality of data systems with measurable SLI/SLO/SLA improvements.
- Hands-on experience supporting both data ingestion pipelines and downstream data consumption or delivery patterns.
- 1+ years of hands-on experience with machine learning-based monitoring, anomaly detection, or AI-assisted observability tools.
- Demonstrated experience with data quality testing, validation frameworks, and quality metrics definition.
- Experience with cloud-based data platforms (AWS, Databricks, or similar).
- Proficiency in Python and SQL, with experience building or supporting production-grade data pipelines.
- Experience implementing data quality frameworks, monitoring tools, and alerting systems.
- Demonstrated expertise with workflow orchestration tools (e.g., Databricks Workflows, Airflow) and version-controlled deployment practices.
- Familiarity with SRE and reliability engineering concepts including SLIs, SLOs, error budgets, and blameless postmortem culture.
- Strong troubleshooting and root cause analysis skills across complex, distributed systems.
- Experience designing and operating observability systems for data pipelines (metrics, logs, traces, alerts).
- Ability to communicate clearly with both technical and non-technical stakeholders during incidents, postmortems, and requirements discussions.
- Understanding of healthcare data, EMR integrations, or regulated data environments is strongly preferred.
- Experience defining and measuring data quality metrics; ability to establish and track reliability KPIs.
- Experience leveraging LLMs or AI-assisted tools (e.g., Claude Code, ChatGPT, GitHub Copilot) to accelerate development of monitoring code, incident response workflows, and documentation.
- Familiarity with healthcare data standards: FHIR, HL7, CCD, claims data formats, and value-based care metrics.
- Experience operating observability and incident management platforms (e.g., DataDog, New Relic, Sumo Logic, PagerDuty).
- On-call experience and demonstrated comfort with incident response, runbook creation, and blameless postmortem analysis.
- Experience with policy-as-code and data governance frameworks.
- Background in a startup or high-growth environment with exposure to scaling data systems.
- Familiarity with Tuva or similar clinical data normalization and quality frameworks.
How to Stand Out
- Highlight any experience you have building monitoring dashboards or alerts for data pipelines; concrete examples stand out.
- Include a brief case study in your resume showing how you reduced data latency or prevented a data‑quality incident.
- Be prepared to discuss specific SLIs/SLOs you have defined and how you measured compliance.
- Demonstrate familiarity with healthcare data compliance (HIPAA, HITECH) during interviews.
- Mention any AI‑assisted tools you’ve used for incident triage or code generation; employers value practical AI experience.
- When negotiating, emphasize the value of remote‑work flexibility and any home‑office setup costs you may have incurred.
- Watch for vague promises about “unlimited” PTO without clear policy; ask how time off is tracked and approved.
This is a remote position listed on WFA Digital, the platform for professionals who work from anywhere. Browse more remote jobs across all categories.