Lead the infrastructure team at Omnifold, focusing on AI model training and deployment in a fast-paced startup environment.
Posted by employer 6 months ago
First seen on Joblaze 1 week ago
Last verified on the company career page 1 day ago
Role intensity
40% coding
AI in the day-to-day
We train custom AI models for forecasting, requiring robust infrastructure for model training and inference.
Requirements
Not disclosed in this posting: compensation, visa sponsorship.
Joblaze summary
The Infrastructure Tech Lead at Omnifold is responsible for developing and maintaining the systems that support custom AI model training and deployment, ensuring reliable processes and robust monitoring. Key skills include expertise in cloud computing, particularly with GPU workloads, and a solid understanding of security practices and CI/CD pipelines, with a strong foundation in Python. This role is ideal for a seasoned professional with around ten years of experience in infrastructure or DevOps, particularly in startup environments. Omnifold's unique approach to AI-driven forecasting presents an opportunity for impactful work in a fast-paced setting.
Joblaze insights
Quick facts
From the original posting
Omnifold trains custom AI models that help planners forecast the future. We are hiring our first infrastructure tech lead, who will own the systems that make everything else possible.
What makes this job interesting:
We train a unique model for each customer, which means model training and inference work differently here than at any other company. You’ll never get more reps building model training infrastructure!
Our team has very fast iteration speed but needs robust monitoring to pick up signal on user patterns. This is especially important as our application interface for AI-driven forecasting is unique on the market.
What you’ll own
Deployment: Reliable processes for getting models and services into production
Security: Data isolation between customers, product security, infrastructure hardening (SOC2 compliance and beyond)
Cloud resource management: GPU allocation, instance sizing, cost optimization
Monitoring and logging: Visibility into what's running, what's failing, and why
Data and ML ops: ETL pipelines from varied customer data sources, model versioning and lifecycle management
Automated testing: Building the test infrastructure that lets us ship with confidence
What we’re looking for
Experience with cloud computing (especially GPU workloads), CI/CD infrastructure-as-code. We run on AWS
Familiarity with or interest in ML workflows
Security fundamentals: encryption, access controls, compliance basics
Python proficiency
Ideally ~10 years of experience, including startup experience, with at least 3 years in a tech lead role. 5+ years in infrastructure, DevOps, or platform engineering roles
Must have a strong Computer Science background
Location: San Francisco (in-person, 5 days per week)
Omnifold’s Mission
Every bad forecast has a physical consequence. Unnecessary goods are manufactured, shipped, and stored. Emergency air freight is needed for misallocated products. Poor production planning means workers show up with nothing to do, or work frantic overtime. Inefficiency is everywhere.
Our mission is to eliminate waste and accelerate growth for every company with physical products.