Industrial Compute
WFA Digital Insight
OpenAI’s Compute organization is the engine that turns cutting‑edge AI research into production‑ready systems. The Industrial Compute role sits at the intersection of software, hardware and physical operations, giving engineers a chance to shape the infrastructure that powers models like GPT‑5.6. Unlike typical data‑center jobs, this position demands a blend of systems thinking and hands‑on problem solving across distributed GPU fleets, power and cooling, and supply‑chain logistics. Candidates will partner with research, hardware and operations teams, so the work is both technically deep and highly collaborative. If you enjoy tackling ambiguous, large‑scale challenges and want your work to directly enable frontier AI, this is a rare chance to contribute at an unprecedented scale.
Job Description
About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities - Help build, scale, and operate OpenAI’s global compute infrastructure.
- Solve complex problems across software, hardware, manufacturing supply chain, and data center systems.
- Improve the reliability, performance, efficiency, and scalability of critical infrastructure.
- Partner with cross-functional teams to bring new compute capacity online quickly and reliably.
- Identify bottlenecks across technical, operational, and physical systems, and develop practical solutions.
- Build tools, processes, systems, or infrastructure that improve execution at scale.
- Contribute to the long-term architecture and operational maturity of OpenAI’s compute footprint.
- Enjoy working on ambiguous, high-impact problems where the path forward is not always defined.
- Are comfortable collaborating across disciplines, including software, hardware, operations, and physical infrastructure.
- Have strong technical judgment and a bias toward execution.
- Care deeply about reliability, speed, safety, and operational excellence.
- Are excited by the challenge of building infrastructure at unprecedented scale.
- Want your work to directly support the development and deployment of frontier AI.
- Have worked on hardware systems, manufacturing, supply chain, data center development, or large capital infrastructure projects.
- Have domain expertise in civil, controls, mechanical, hardware, electrical, thermal, power, networking, or facilities engineering.
- Have helped bring new technical platforms, data centers, factories, or large-scale systems from concept to production.
- Have experience operating in fast-moving environments where technical depth and execution speed both matter.
How to Stand Out
- Highlight any experience you have with large‑scale GPU clusters or data‑center operations, even if it was part of a side project.
- Prepare concrete examples of ambiguous problems you solved, focusing on your decision‑making process and outcomes.
- Include a short case study in your resume that showcases how you improved reliability or efficiency using Excel for analysis.
- Familiarize yourself with OpenAI’s recent research releases; being able to reference them shows genuine interest.
- During interviews, be ready to discuss trade‑offs between hardware cost, power consumption, and performance.
- Ask about the team’s current bottlenecks to demonstrate forward‑thinking and alignment with the role’s challenges.
- Negotiate for equity and remote‑work allowances early, as these are standard components of OpenAI’s compensation framework.
This is a remote position listed on WFA Digital, the platform for professionals who work from anywhere. Browse more remote jobs across all categories.