Is the Generative AI Inference Engineer role remote?

Yes — Stability AI lists this as a fully remote position.

How much experience is required?

At least 7 years of relevant experience for this Generative AI Inference Engineer role.

What's the tech stack?

Joblaze extracted these technologies from the posting: AWS, NVIDIA Nsight, Triton, TensorRT, PyTorch, OpenCV.

What seniority level is this role?

Stability AI targets senior candidates for this position.

Is this full-time or contract?

Full-time for this Generative AI Inference Engineer role at Stability AI.

Generative AI Inference Engineer at Stability AI

Joblaze summary

The Generative AI Inference Engineer at Stability AI focuses on developing and optimizing multi-modal machine learning inference systems, driving innovation in generative AI applications. Candidates should possess extensive experience in productionizing machine learning systems, particularly with Python, PyTorch, and high-performance inference frameworks. This role is ideal for seasoned professionals with a strong background in deep learning and cloud deployment, who thrive in collaborative environments alongside leading researchers. The position offers a chance to work with cutting-edge technology and contribute to impactful projects in a rapidly evolving field.

Joblaze insights

Listed about a month ago on Joblaze — tracked directly from Stability AI's career page.
Stability AI has 10 other open roles including Junior Software Engineer, Senior Site Reliability Engineer , Solutions Engineer.
2358 active Python roles on Joblaze right now.
1250 active AWS roles on Joblaze right now.
1056 active Kubernetes roles on Joblaze right now.

Quick facts

Is the Generative AI Inference Engineer role remote?: Yes — Stability AI lists this as a fully remote position.
How much experience is required?: At least 7 years of relevant experience for this Generative AI Inference Engineer role.
What's the tech stack?: Joblaze extracted these technologies from the posting: AWS, NVIDIA Nsight, Triton, TensorRT, PyTorch, OpenCV.
What seniority level is this role?: Stability AI targets senior candidates for this position.
Is this full-time or contract?: Full-time for this Generative AI Inference Engineer role at Stability AI.

From the original posting

Generative AI Inference Engineer

About the role:

We are seeking passionate Machine Learning Engineers to join our Inference team, focusing on the creative applications of generative AI models. The ideal candidate will have substantial experience developing and running inference for multi-modal models. A deep understanding of diffusion model architectures and familiarity with workflow tools like ComfyUI are a big plus. You will be expected to leverage and push the boundaries of state-of-the-art inference optimization techniques for multi-modal generative models. This role offers the opportunity to work alongside top researchers and engineers, utilizing cutting-edge high-performance computing resources to make a significant impact in the rapidly evolving field of generative AI.

Responsibilities:

Lead efforts to drive the design, development of customer-facing multi modal ML inference systems.
Work with the Platform and Inference teams on building inference systems for the next generation of models, where you will work on areas such as optimization, model tuning and deployment.
Partner with leading cloud providers to deliver hosted Stability AI inference solutions.
Be a strategic thought partner for leaders across the organization on driving business impact through machine learning
Be part of the team to bring new Stability models and pipelines into existence
Prototype and productionize inference platform improvements and new features

Qualifications:

7+ years working on productionizing machine learning systems, including inference pipeline development
Expert level knowledge on writing and running python services at scale
5+ years working on python scientific stack, pyTorch and at least one high-performance inference framework (e.g. Triton and TensorRT)
Deep understanding of Diffusion Architecture
Experience profiling and optimizing deep neural networks on Nvidia GPUs, using profiling tools such as NVIDIA Nsight
Experience with python-based image manipulation/encoding/decoding frameworks, such as OpenCV
Experience deploying to cloud orchestration systems such as Kubernetes and cloud providers such as AWS, GCP, and Azure
Experience with Docker
Ability to rapidly prototype solutions and iterate on them with tight product deadlines
Strong communication, collaboration, and documentation skills
Experience with the open-source ML ecosystem (HuggingFace, W&B, etc.)

Equal Employment Opportunity:

We are an equal opportunity employer and do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, veteran status, disability or other legally protected statuses.

Generative AI Inference Engineer

Similar positions