Build and own the infrastructure and pipelines for machine learning models in production at a fast-growing startup.
Posted by employer 1 month ago
First seen on Joblaze 1 week ago
Last verified on the company career page 5 days ago
Skills & Technologies
Requirements
Not disclosed in this posting: compensation, work arrangement, visa sponsorship.
Joblaze summary
In this role, the Senior MLOps Engineer is responsible for developing and maintaining the infrastructure and pipelines essential for training and deploying machine learning models in a production environment. Key skills include proficiency in Python, cloud technologies, Docker, Kubernetes, and CI/CD practices, alongside experience in managing production ML systems and GPU workloads. This position is ideal for candidates with 5 to 7 years of relevant experience, particularly those who thrive in fast-paced startup settings. The role emphasizes collaboration with cross-functional teams to enhance the scalability and reliability of the ML platform.
Joblaze insights
Quick facts
From the original posting
Build and own the infrastructure and pipelines used to train, evaluate, package, deploy, and operate machine learning models in production.
Develop reliable MLOps capabilities across experiment tracking, model and data versioning, reproducibility, orchestration, automated testing, monitoring, and controlled model rollouts.
Partner with Data Science, Engineering, and Infrastructure teams to productionize models and continuously improve the scalability, reliability, observability, and cost efficiency of our ML platform.
5–7+ years of experience in Machine Learning Engineering, MLOps, ML Infrastructure, Platform Engineering, or a related production engineering role.
Strong hands-on experience with Python, cloud infrastructure, Docker, Kubernetes, CI/CD, infrastructure-as-code, workflow orchestration, and production observability.
Proven experience building and operating production ML systems, including training pipelines, experiment tracking, model registries, versioning, monitoring, data-quality checks, staged deployments, and rollback mechanisms.
Experience managing GPU-based training workloads and/or distributed training infrastructure, and cloud cost optimization.
Experience working in a fast-growing startup, with the ability to operate in a dynamic, fast-paced environment.