HomeSearchAi Engineer Jobs › Staff ML Platform Engineer (MLOps)

Staff ML Platform Engineer (MLOps)

FutureFit AI

Remote (US); Remote (Canada) · staff
More Ai Engineer jobs: Ai Engineer jobsAi Engineer salary

Come join our Data team! High velocity, high trust, and high impact with a will to win. If that resonates deeply with you, this could be your next career move. We're seeking someone who leads with humility, pursues audacious goals, and is motivated by meaningful impact on people and the world. At FutureFit AI, our core mission is to help more people get to better jobs faster and cheaper, with a specific focus on those facing barriers to opportunity.

Our work helps resolve the growing issue of economic inequality, ensuring that no one is left behind in the future of work. Our AI-powered platform brings efficiency and insight to workforce development, replacing outdated systems and unlocking human potential at scale. Ready to make an impact? Apply today. Important note: Data shows that men typically apply when meeting 3/10 requirements, while women often wait until it's 10/10.

We encourage you to apply if you see a strong (not necessarily perfect) fit. The Opportunity We're seeking a Staff ML Platform Engineer (MLOps) to build the platform our ML and LLM-powered products run on. Our ML footprint has grown fast, but the layer underneath it has not kept pace. You'll own that layer end to end: how models get built, deployed, evaluated, and served; how compute and environments get provisioned and managed; how our LLM calls get routed and optimized for cost; and how we know quickly when a recommender goes down or goes off the rails.

This is a build role and an operate role: when a model regresses or a recommendation looks wrong, you can trace it back to the inputs that produced it and help fix it. Your Role Our ML footprint has grown quickly: batch models, real-time recommendation models, LLM-powered features, and daily pipelines processing every available job across the US and Canada. What we haven't built is the platform underneath it: consistent compute and environments, a disciplined path from experiment to production, cost-aware routing across LLMs, and the monitoring that tells us fast when something breaks.

You'll assess our current pipelines and ML workflows with clear eyes, decide what to build and in what order, then build it. This is greenfield platform work with direct influence on production models and how our ML team operates, and it comes with real operational ownership: you will be close enough to the running systems to debug them, not one step removed. What You'll Own Assessment and plan: Evaluate our current pipelines, data architecture, and ML workflows, and produce a prioritized, opinionated plan for what needs to change.

Platform foundations and optimization: Own compute provisioning and environment management, keep training and serving environments reproducible, keep frameworks and packages current across services and model images, and tune latency, throughput, and spend, all without destabilizing production. LLM infrastructure and smart routing: Build the layer our LLM features run on, including smart routing that sends each request to the cheapest model that can handle it well, plus the prompt and response evaluation needed to prove quality holds when we route down.

Experimentation and safe rollout: Give us a real discipline for A/B testing models before they are fully ramped: shadow deploys, canaries, holdouts, and success criteria agreed in advance, so a model earns its way into production instead of being switched on. Observability and traceability: Know within minutes when a recommender goes down or starts drifting, and be able to explain why: model and data monitoring, alerting, regression detection, lineage, and enough traceability to reproduce a questionable recommendation on demand or trace a prediction back to the inputs that produced it (we currently use Braintrust; comparable tooling counts too).

Hands-on operations: Stay close enough to the running systems to operate them. You will work with the team to keep models online, but when something breaks in production you can dive in and help fix it, in

Search all live jobs — free, no account →