Senior Infrastructure Engineer
Software Development
Job Description
About
The Role
We're looking for a Senior Infrastructure Engineer to build and run Somnia's key backend services: the L1 and node fleet, RPC and indexing layers, product backends, and developer-facing services teams depend on. An SRE-minded role: you make reliability measurable, rollouts safe, and infrastructure repeatable, so every team can move fast without breaking things. You treat infrastructure as a product: automate relentlessly, measure everything, and leverage AI to accelerate development, operations, and incident response.Key Responsibilities
Define and maintain SLOs, SLIs, and error budgets, plus the observability—metrics, logs, traces and alerts—that catches regressions before users do. Build repeatable, self-service infrastructure through infrastructure-as-code, CI/CD and golden paths so teams can provision, deploy and recover without reinventing the wheel. Own rollouts end-to-end—progressive delivery, canaries, safe migrations and clean rollbacks. Operate the systems behind Somnia's nodes, validators, RPC and indexing, tuning for performance and cost across regions. Lead incident response and on-call, run blameless postmortems, and continuously harden the platform. Partner with product and protocol teams to design and operate production-ready services. You'll rotate between embedding with engineering teams and building the shared platform, tooling and operational standards that underpin the wider organisation.Requirements
Must Have - Strong experience operating production infrastructure at scale (cloud and/or bare metal), with deep Linux fundamentals.- Experience with infrastructure-as-code such as Terraform or Pulumi, alongside configuration management.
- Experience running containers and orchestration platforms (Docker, Kubernetes) in production.
- Strong programming skills, ideally in Go and/or TypeScript, for building automation and internal tooling.
- Experience with observability stacks (Prometheus, Grafana, OpenTelemetry or equivalents).
- Experience operating and monitoring distributed systems, including capacity planning and performance tuning.
- Comfortable operating in high-stakes production environments and responding to incidents.
- Genuine interest in crypto and on-chain systems.
- Experience with high-performance networking, low-latency systems or load balancing at scale.
- Multi-region and geo-distributed deployments with failover strategies.
- Security and key management (HSMs, secrets management, hardening).
- EVM tooling and the wider Web3 infrastructure ecosystem.
- Platform reliability is measurable, with well-defined SLOs and continuously improving service health.
- Infrastructure is automated, repeatable and increasingly self-service.
- Incidents become less frequent, easier to diagnose and faster to resolve.
- Product teams spend more time shipping features and less time managing infrastructure.
How to Stand Out
- Showcase a production‑grade IaC portfolio: Publish a public GitHub repo (or a private link with access) that contains end‑to‑end Terraform (or Pulumi) modules for provisioning a Kubernetes cluster, L1 node fleet, and RPC services, complete with CI/CD pipelines (GitHub Actions/GitLab CI) that run `terraform plan/apply` automatically. Include README docs, versioned releases, and a demo of self‑service provisioning via a CLI or web UI—Somnia’s interviewers will look for repeatable, product‑like infrastructure you’ve shipped.
- Demonstrate measurable reliability expertise: Prepare a concise case study (1–2 slides) describing how you defined SLOs/SLIs, set error budgets, and built observability stacks using Prometheus, Grafana, OpenTelemetry, and Alertmanager for a high‑traffic service. Highlight specific metrics (e.g., 99.9% request latency ≤ 100 ms) and the impact on rollout safety (e.g., reduced rollback frequency by 30%). Somnia’s SRE‑mindset interview will probe these numbers.
- Highlight AI‑augmented ops experience: Build a small prototype that uses an LLM (e.g., OpenAI GPT‑4 or a locally hosted model) to generate incident runbooks, suggest remediation steps from log patterns, or auto‑populate Terraform variable defaults. Share the code or a video walkthrough; Somnia explicitly wants engineers who “leverage AI to accelerate development, operations, and incident response.”
- Prepare for scenario‑based reliability questions: Practice answering “design a canary rollout for a new RPC endpoint with a 5‑minute error budget” and “how would you rebalance the node fleet after a sudden traffic spike?” Write out your step‑by‑step plan, include tooling (Argo Rollouts, Helm, Service Mesh), and be ready to whiteboard it over a shared doc during the interview.
- Negotiate with data and remote‑work value: Research senior infrastructure salaries (e.g., $170‑200k base + 0.15‑0.25% equity for fully remote U.S. roles). In your negotiation email, cite your proven IaC product releases, AI‑ops prototype, and the cost‑saving impact (e.g., $200k/year reduction in manual toil). Emphasize the flexibility remote work offers and ask for a total‑comp package that matches market benchmarks plus a signing bonus for relocation‑independent talent.
This is a remote position listed on WFA Digital, the platform for professionals who work from anywhere. Browse more remote jobs across all categories.