Machine Learning Engineer - Inference Maintainer & Developer Experience
Job Description
Our mission is to make the world programmable. Sight is one of the key ways we understand the world, and soon this will be true for the software we use, too. We’re building the tools, community, and resources needed to make the world programmable with artificial intelligence. Roboflow simplifies building and using computer vision models. Today, over 1M+ developers, including those from half the Fortune 100, use Roboflow’s machine learning open source and hosted tools. That includes counting cells https://blog.roboflow.com/cancer-research-computer-vision/ to accelerate cancer research, improving construction site safety https://blog.roboflow.com/preventing-accidents-on-construction-sites-with-computer-vision/, digitizing floor plans https://blog.roboflow.com/floor-plan-analysis-computer-vision/, preserving coral reef populations https://blog.roboflow.com/reefos-supercharging-coral-reef-restoration-with-ai/, guiding drone flight https://blog.roboflow.com/georeferencing-drone-videos/, and much more https://roboflow.com/templates. Our team is small relative to our impact, and we believe our user success is our success (not the inverse). A team member summarized: “Roboflow is a company full of giant brains and tiny egos.” We find software has a multiplier effect on all roles (not only product and engineering), so Roboflow employs developers across the company in design, sales, customer support, marketing, and beyond. We’re supported by great customers and investors, having raised over 63 million from Google Ventures, Y Combinator, Craft Ventures, Sam Altman, Lachy Groom, amongst other leading software investors. At the center of all of this is inference https://github.com/roboflow/inference — one of our most important open source projects and the engine that runs computer vision models everywhere, from cloud GPUs to edge devices in the field. It powers our commercial platform and is relied on by tens of thousands of developers. This role exists to be its steward. WHY THIS ROLE EXISTS Inference https://github.com/roboflow/inference is growing fast — and so is the volume of contributions, increasingly authored with the help of AI agents. That's a great problem to have, but it's outpacing our ability to keep quality high and cut releases on a predictable cadence. Today we ship roughly weekly, and it's a fight. We want to flip that equation. The goal is to build and continuously evolve an agentic-driven contribution and release pipeline — automated and semi-automated review, triage, CI/CD, and end-to-end testing — so that we can safely absorb a high volume of agent-generated PRs while staying firmly in control of quality. The ideal end state: nightly end-to-end tests across every target (both standalone and on-platform), backed by a growing, world-grounded suite that validates the real health of every build. With that foundation, daily releases become routine, and we can say "yes" to far more contributions without ever lowering the bar — pushing back, by design, according to strictly defined review standards. Alongside that, this person becomes the human face of inference https://github.com/roboflow/inference: teaching internal teams and customers how to get more out of it, partnering with marketing to tell its story, and owning the (genuinely fun) work of bringing new models into the engine. WHAT WE'RE LOOKING FOR Primarily, you like to make great things with passionate colleagues. You are someone who likes to own outcomes, not only inputs. You're motivated by having responsibility and accountability. You're eager to 'do the work,' big and small. You're motivated by the question, "How can I improve this?" and have a track record of doing so, even in ways adjacent to your role. Much of our current team is made up of former founders who thrive in the level of autonomy at Roboflow. Maybe you had a side hustle in high school or college. You care about open source and the developers who depend on it. One of the best ways to stand out among other applicants is to write about something you've built with Roboflow, or to contribute to one of our open source projects — inference https://github.com/roboflow/inference especially. What You'll Do - Build and maintain inference https://github.com/roboflow/inference, our flagship open source and commercial CV inference engine, keeping it healthy and high-quality as contribution volume scales.
- Build an agentic-driven contribution pipeline — automated and semi-automated review, triage, and CI/CD — so we can safely accept a high volume of agent-generated PRs and move from weekly releases toward daily ones.
- Design and grow a world-grounded, ever-expanding test suite that validates real build health across every target (standalone and on-platform), with the goal of nightly end-to-end runs across all of them.
- Define and enforce the "rules of the road" — the review standards and skills that agents and contributors must follow.
- Streamline how new models get added to inference https://github.com/roboflow/inference (the most fun part of the job) — making it dramatically faster and easier to bring the latest computer vision and ML models to our users.
- Teach and enable internal teams and customers.
- Be the bridge between core engineering and clients — translating new capabilities into docs, demos, stories, and launches which would help people use inference https://github.com/roboflow/inference more effectively.
- Contribute to and grow the broader open source community around the project.
- 5+ years of hands-on experience building and operating production-grade ML systems, ideally involving large-scale deployment of modern AI models.
- A real CV/ML foundation — you understand what inference does: how computer vision models work internally, how they're deployed across diverse environments, and how to adapt them for real-world, high-impact use.
- Stellar agentic skills.
- Strong CS and systems background, with the ability to independently tackle complex programming, architecture, and reliability challenges and exercise sound judgment on when to move fast and when rigor is essential.
- Hands-on experience with CI/CD, release engineering, and test infrastructure — you've built or substantially improved automated testing and delivery pipelines before.
- Practical expertise with core ML technologies, including several of the following: PyTorch, TensorFlow, ONNX, TensorRT, vLLM (or other LLM/model deployment tools).
- Strong proficiency in image and video processing, including several of the following: OpenCV, DeepStream, Pillow, PyAV, hardware-accelerated video decoding.
- Excellent communication and soft skills. You can teach, write clearly, and collaborate across engineering, support, field, and marketing — and you actually enjoy it. You're comfortable being a public-facing voice for a project.
- Open source maintenance experience is a strong plus — you know what it takes to steward a busy repo and a community of contributors.
- Level-up your performance with AI agents.
- The best way to stand out is to write about something you’ve built with Roboflow or contribute to one of our open source projects https://roboflow.com/open-source.
- We may send you a technical screen if applicable. Introduction Phase: - [15m] Technical Assessment Team Interview Phase: - Live coding [45m] - Home assignment - [30m] Meet with Inference Core team member - [60m] Meet with hiring manager - Use this time to review specifics about the job description - Begin working through your 30/60/90 projects - Ask questions!
How to Stand Out
- Build a concise demo that uses Roboflow’s public API to deploy a vision model and serve real‑time inference (e.g., a Flask or FastAPI app). Include a short README that explains the deployment steps, performance metrics (latency, throughput), and how you’d instrument monitoring—this directly shows you can maintain and improve their inference stack.
- Contribute a pull request to an open‑source computer‑vision repo (preferably Roboflow’s own SDK or a related tool) that adds a new inference‑optimisation feature (e.g., ONNX export, TensorRT acceleration, or batch inference). Highlight the PR link in your resume and be ready to discuss the design decisions and code‑review feedback in the interview.
- Prepare a one‑page “Inference Performance Dashboard” built in Excel (or Google Sheets) that visualizes latency, memory usage, and cost across different hardware back‑ends (CPU, GPU, edge TPU). Use formulas, conditional formatting, and pivot tables to show you can turn raw metrics into actionable insights for product teams—exactly the skill set they list under “Excel”.
- Emphasize remote‑first collaboration: list specific tools you use (e.g., GitHub Actions for CI, Notion for docs, Slack/Discord for async communication, and Miro for design reviews). In the interview, walk through a recent cross‑time‑zone sprint where you delivered a production inference upgrade, focusing on how you kept stakeholders aligned without face‑to‑face meetings.
- Research Roboflow’s recent case studies (cell counting, construction safety, floor‑plan analysis) and prepare a short “product‑impact” story showing how you would improve inference latency or model‑size for one of those use‑cases. Quantify the benefit (e.g., “reduce per‑image latency from 150 ms to <80 ms → 30 % cost saving on GPU usage”).
- When negotiating salary, cite the 2024 market median for senior ML inference engineers in NY/SF ($170k–$190k base) and remote benchmarks ($150k–$170k). Mention that you bring added value through open‑source contributions and proven cost‑saving inference optimisations, which justifies positioning your ask toward the top of that range plus a performance‑based bonus.
This is a remote position listed on WFA Digital, the platform for professionals who work from anywhere. Browse more remote jobs across all categories.