Join Omnifold's Infrastructure Team to build robust systems for AI model training and deployment in a fast-paced environment.
Posted by employer 6 months ago
First seen on Joblaze 1 week ago
Last verified on the company career page 1 day ago
AI in the day-to-day
We train custom AI models for forecasting, requiring unique infrastructure for model training and inference.
Requirements
Not disclosed in this posting: compensation, years of experience, visa sponsorship.
Joblaze summary
In this role, the individual will focus on developing and maintaining the infrastructure that supports the deployment and monitoring of custom AI models for forecasting. Key skills include expertise in cloud computing, particularly with GPU workloads, and proficiency in Python, alongside a solid understanding of security practices and CI/CD processes. This position is ideal for someone with a strong computer science background who thrives in a fast-paced environment and is eager to engage with unique machine learning workflows. Omnifold's Infrastructure Team plays a crucial role in ensuring the reliability and efficiency of their innovative AI-driven solutions.
Joblaze insights
Quick facts
From the original posting
Omnifold trains custom AI models that help planners forecast the future. We are hiring for our Infrastructure Team, who own the systems that make everything else possible.
What makes this job interesting:
We train a unique model for each customer, which means model training and inference work differently here than at any other company. You’ll never get more reps building model training infrastructure!
Our team has very fast iteration speed but needs robust monitoring to pick up signal on user patterns. This is especially important as our application interface for AI-driven forecasting is unique on the market.
What you’ll own
Deployment: Reliable processes for getting models and services into production
Security: Data isolation between customers, product security, infrastructure hardening (SOC2 compliance and beyond)
Cloud resource management: GPU allocation, instance sizing, cost optimization
Monitoring and logging: Visibility into what's running, what's failing, and why
Data and ML ops: ETL pipelines from varied customer data sources, model versioning and lifecycle management
Automated testing: Building the test infrastructure that lets us ship with confidence
What we’re looking for
Experience with cloud computing (especially GPU workloads), CI/CD infrastructure-as-code. We run on AWS
Familiarity with or interest in ML workflows
Security fundamentals: encryption, access controls, compliance basics
Python proficiency
Must have a strong Computer Science background
Location: San Francisco (in-person, 5 days per week)
Omnifold’s Mission
Every bad forecast has a physical consequence. Unnecessary goods are manufactured, shipped, and stored. Emergency air freight is needed for misallocated products. Poor production planning means workers show up with nothing to do, or work frantic overtime. Inefficiency is everywhere.
Our mission is to eliminate waste and accelerate growth for every company with physical products.
Explore more