Senior Solutions Engineer
WFA Digital Insight
TensorWave’s Senior Solutions Engineer sits at the nexus of cutting‑edge AI infrastructure and customer success. Unlike generic support roles, this position acts as the final escalation point for teams training massive models—any downtime can cost billions, so the job demands relentless, kernel‑level debugging and an intimate grasp of GPU‑focused stacks. Candidates will be pulling apart Kubernetes controllers, GPU drivers, and high‑performance networking layers while translating those findings into actionable product upgrades. The role also blends technical depth with executive‑level communication, meaning you’ll regularly brief VPs on complex incidents. If you thrive on turning obscure system failures into resilient engineering solutions, TensorWave offers a rare blend of hands‑on problem solving and strategic influence.
Job Description
About TensorWave Our mission is simple: deliver seamless, secure, reliable, and resilient AI compute at scale. We've built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas, not infrastructure. About the Role We're looking for a Senior Solutions Engineer to serve as the elite escalation point between our Global Operations Center (GOC) and our Core Engineering teams. You are the technical backstop for our most sophisticated customers teams training large models who cannot afford a single hour of downtime. You'll own the problems that go beyond runbooks, sitting at the intersection of customer success and engineering: resolving the hardest technical blockers and translating those findings into a more resilient product. If you're an engineer who loves the detective work of kernel-level debugging and high-performance networking, and who also thrives in the high stakes environment of customer-facing resolution, this is your role. What You’ll Do - Resolve Complex Escalations: Act as the final authority on issues exceeding GOC scope, utilizing code-level debugging and architectural investigation.
- Direct Customer Engagement: Partner with customer technical leads to diagnose production issues, ensuring transparency and rapid resolution through active collaboration.
- Iterative Problem Solving: Develop diagnostic scripts and workarounds to maintain customer operations while long-term patches are in development.
- Drive Root Cause Analysis: Own end-to-end P1 resolution, partnering with TAMs to deliver clear, actionable post-incident analysis.
- Bridge to Engineering: Convert recurring customer pain points into evidence-based feature requests, influencing product roadmap to resolve systemic failures.
- Build Scalable Knowledge: Document non-obvious platform behaviors and refine GOC runbooks, ensuring institutional knowledge grows with every incident.
- Kubernetes Expert: Deep experience in cluster administration and scheduler internals; comfortable reading/modifying controller code.
- AI/GPU Infrastructure Specialist: Proficient in orchestrating GPU workloads and diagnosing training job failures using ROCm or CUDA.
- Network Pathologist: Skilled in RDMA/RoCEv2, SRIOV, and BGP; capable of interpreting switch telemetry to identify silent packet drops.
- Linux Power User: Expert in kernel networking, hugepages, and cgroups; able to debug at the OS layer when applications are silent.
- Builder Mindset: Proficient in Python and Ansible; capable of writing custom diagnostic tools to automate remediation.
- Executive Communicator: Strong technical rigor when presenting findings to VPs of Engineering, maintaining trust while delivering difficult updates.
- Experience in high-uptime environments where 24/7/365 availability is required. What We Offer - Stock Options - 100% paid Medical, Dental, and Vision insurance for Employees - Company Health Savings Account Contributions - 100% paid Short Term and Long Term Disability Insurance for Employees - Life and Voluntary Supplemental Insurance Options - Other Insurance Options, such as Pet & Legal Insurance - Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support - Flexible Spending Account - 401(k) - Employee Assistance Program - Flexible PTO - Paid Holidays - Parental Leave - Other In-Office Perks Equal Employment Opportunity TensorWave is an Equal Opportunity Employer.
How to Stand Out
- Brush up on GPU debugging tools (CUDA Nsight, ROCm ROCprof) and be ready to demonstrate a short troubleshooting walkthrough during the interview.
- Prepare concrete examples where you turned a complex production incident into a lasting product improvement.
- Highlight any open‑source contributions or internal tooling you built with Python/Ansible; bring a GitHub repo or code snippets.
- Practice explaining deep technical concepts (e.g., RDMA packet loss) in plain language for senior executives.
- Research TensorWave’s platform architecture (Kubernetes + GPU orchestration) so you can ask insightful questions about their stack.
- When negotiating, emphasize the equity component and remote‑work flexibility as key value drivers.
- Watch for vague promises about “unlimited PTO” without clear policy details; ask how time‑off is tracked and approved.
This is a remote position listed on WFA Digital, the platform for professionals who work from anywhere. Browse more remote jobs across all categories.