Own the end-to-end lifecycle of production ML serving systems for a top-performing AI Shopping Agent.
Posted by employer 5 months ago
First seen on Joblaze 1 week ago
Last verified on the company career page 1 day ago
AI in the day-to-day
Our ML models power the core of our platform, focusing on reliable and efficient production ML serving.
Requirements
Not disclosed in this posting: compensation, work arrangement, visa sponsorship.
Joblaze summary
In this role, the Senior Machine Learning Engineer will manage the entire lifecycle of production ML serving systems, ensuring they operate efficiently and reliably under real-world conditions. Key skills include strong Python programming, experience with cloud platforms, and a deep understanding of inference performance metrics. This position is ideal for seasoned professionals with a background in software or ML engineering, particularly those who thrive in fast-paced startup environments. The engineer will play a crucial role in shaping the architecture of Wizard's inference platform, directly impacting the performance of its AI shopping agent.
Joblaze insights
Quick facts
From the original posting
At Wizard AI, we’re building the top-performing AI Shopping Agent that delivers the best products from across the web with unmatched accuracy, quality, and trust. Our ML models power the core of our platform, and we’re looking for a Senior Machine Learning Engineer to own how they run in production reliably, efficiently, and at scale.
As a Senior ML Engineer on our Inference Platform, you’ll own the end-to-end lifecycle of production ML serving systems from model packaging and deployment to monitoring, optimization, and scaling. This is not a traditional MLOps role focused solely on pipelines and tooling. You’ll be responsible for the inference infrastructure powering a live conversational shopping agent, operating multiple specialized serving engines under real-world production load.
You’ll own critical decisions around serving architecture, performance, reliability, and scalability, working closely with ML Engineers, Data teams, Product, and DevOps to ensure models move seamlessly from experimentation into high-performance production systems.
Production serving infrastructure operates with clear SLAs, strong observability, and minimal downtime. Latency, availability, throughput, and GPU utilization are actively measured and optimized as platform demands grow.
You own the complete serving lifecycle — from deployment and release management through monitoring, optimization, and scaling — enabling ML engineers to ship quickly while maintaining reliability and reproducibility.
You shape the future of Wizard's inference platform, driving key architectural decisions that improve performance, reduce infrastructure costs, and support the next generation of AI-powered shopping experiences.