← Back to results

Senior Software Engineer, Machine Learning Infrastructure & Automation

Build automation and infrastructure for ML models in a high-impact role at a growing generative media ecosystem.

Location
United States
Compensation
Not disclosed
Level
senior
Type
full time · Remote

Posted by employer 8 hours ago

First seen on Joblaze 5 hours ago

Last verified on the company career page 5 hours ago

Apply at Fal → Save job Scanned from fal.ai

Skills & Technologies

What you'll build

  • Own ML CI/CD infrastructure
  • Accelerate development cycles
  • Build automated model validation
  • Automate performance benchmarking
  • Improve deployment reliability

Must have

  • 5+ years of software engineering background
  • Proficiency in Python
  • Experience designing and operating CI/CD systems
  • Deep understanding of automated testing
  • Experience working with containerized workloads

Nice to have

  • Experience with ML infrastructure
  • Familiarity with NVIDIA GPU architectures
  • Experience building automated inference benchmarks
  • Experience with agentic coding tools
  • Experience optimizing CI/CD pipelines at scale

AI in the day-to-day

Develop AI-powered automation and agentic coding systems that help engineers iterate faster.

Requirements

Experience
5+ years

Not disclosed in this posting: compensation, visa sponsorship.

Benefits

Health Insurance

Joblaze summary

In this role, the Senior Software Engineer focuses on enhancing fal's machine learning infrastructure by developing automation tools and CI/CD systems that streamline the deployment and testing of generative AI models. Key skills include proficiency in Python, experience with CI/CD technologies like GitHub Actions, and a solid understanding of automated testing and cloud infrastructure. This position is well-suited for experienced engineers who are adept at identifying inefficiencies and implementing solutions that boost developer productivity. The role involves collaboration with applied ML teams, emphasizing a culture of innovation and continuous improvement.

Joblaze insights

  • Listed today — first seen on Joblaze October 9, 2026. Last confirmed on Fal's careers page October 9, 2026.
  • Python appears in 52.9% of 546 comparable senior ai/ml roles in United States; GitHub Actions appears in 0.4% of 546 comparable senior ai/ml roles in United States.

Quick facts

Is the Senior Software Engineer, Machine Learning Infrastructure & Automation role remote?
Yes — Fal lists this as a fully remote position.
How much experience is required?
At least 5 years of relevant experience for this Senior Software Engineer, Machine Learning Infrastructure & Automation role.
What's the tech stack?
Joblaze extracted these technologies from the posting: Docker, GitHub Actions, PyTorch, Python.
What seniority level is this role?
Fal targets senior candidates for this position.
Is this full-time or contract?
Full-time for this Senior Software Engineer, Machine Learning Infrastructure & Automation role at Fal.

From the original posting

About this role:

Help fal's ML team move faster by building the automation, infrastructure, and developer tooling that makes developing, testing, and deploying generative AI models seamless.

You'll own and improve the CI/CD systems supporting our rapidly growing collection of ML models and inference pipelines. Your focus will be on eliminating manual work, accelerating development cycles, and building reliable systems that allow ML engineers to ship new models and optimizations with confidence.

This is a high-impact engineering role where you'll work closely with our Applied ML and ML Performance teams. You'll build everything from automated model validation and performance benchmarking to AI-powered development workflows that help engineers iterate faster.

The ideal candidate thinks beyond traditional CI/CD and sees automation as a force multiplier for the entire engineering organization.

What you’ll do:

  • Own ML CI/CD infrastructure Design, build, and maintain automated testing, validation, and deployment pipelines for our ML models and inference services.

  • Accelerate development cycles.Dramatically reduce CI execution times through intelligent parallelization, caching, test selection, and efficient use of compute resources.

  • Build automated model validation.Develop systems that test model outputs, detect quality regressions, and validate changes across different models, GPU architectures, and configurations.

  • Automate performance benchmarking. Build continuous performance testing that detects regressions in inference latency, throughput, GPU utilization, and cost.

  • Build automated pricing and deployment checks. Ensure model pricing, billing configurations, API schemas, and deployments are validated automatically before reaching production.

  • Extend our agentic engineering workflows. Develop AI-powered automation and agentic coding systems that automatically diagnose CI failures, identify regressions, propose fixes, and streamline engineering workflows.

  • Improve deployment reliability. Build automated safeguards, deployment verification, rollback mechanisms, and monitoring to ensure new model releases are reliable.

  • Eliminate engineering toil. Identify repetitive tasks across the ML team and build tools and systems that automate them, allowing engineers to focus on developing new models and improving performance.

Qualifications/Nice-to-haves:

  • 5+ years of software engineering background with proficiency in Python and experience building production infrastructure and developer tooling.

  • Experience designing and operating CI/CD systems using GitHub Actions or comparable technologies.

  • Deep understanding of automated testing, build systems, dependency management, caching, and parallel execution.

  • Experience working with containerized workloads, Docker, and cloud infrastructure.

  • Ability to design reliable distributed systems and debug complex infrastructure failures.

  • Strong understanding of observability, including logs, metrics, tracing, and automated alerting.

  • A passion for developer productivity and a demonstrated ability to eliminate manual processes through automation.

  • Comfortable working independently, identifying high-impact problems, and building end-to-end solutions.

  • Experience with ML infrastructure, PyTorch, GPU workloads, or model-serving systems.

  • Familiarity with NVIDIA GPU architectures and multi-GPU environments.

  • Experience building automated inference benchmarks or ML quality evaluation frameworks.

  • Experience with agentic coding tools such as Codex or Claude Code, or building custom AI engineering agents.

  • Experience optimizing CI/CD pipelines at scale, including distributed test execution and ephemeral compute environments.

  • Experience developing internal developer platforms or infrastructure-as-code tooling.

Tech Stack:

  • Python, PyTorch, Docker, GitHub Actions

  • NVIDIA GPUs and distributed GPU infrastructure

  • Model serving and inference pipelines

  • Observability and performance monitoring tools

  • AI coding agents and automated engineering workflows

  • Access to fal's massive GPU cluster for testing and development

What we offer at fal:

  • Interesting and challenging work

  • A lot of learning and growth opportunities

  • Health, dental, and vision insurance (US)

  • Regular team events and offsites

Standard company text repeated across Fal's postings is omitted here.

Similar positions

Fal
Fal
Software Engineer, Applied Machine Learning
Fal · San Francisco, United States
Fal