← Back to results

Software Engineer, Distributed Systems

Join fal as a Software Engineer to build large-scale distributed systems for AI products in a growth-focused environment.

Location
San Francisco
Compensation
$180k–$250k/yr
Level
mid
Type
full time · On-site

Posted by employer 1 year ago

First seen on Joblaze 1 week ago

Last verified on the company career page 1 day ago

Apply at Fal → Save job Scanned from fal.ai

Skills & Technologies

Rust Python Flexible on stack

AI in the day-to-day

Leverage AI to automate the mundane parts of building complex but reliable systems.

Requirements

Experience
3+ years

Not disclosed in this posting: visa sponsorship.

Benefits

Health Insurance Relocation Assistance

Joblaze summary

In this role, the software engineer focuses on developing and enhancing a robust distributed systems platform that supports high-performance AI workloads. Proficiency in Python or Rust is essential, along with a solid grasp of distributed systems principles such as fault tolerance and scheduling. This position is ideal for experienced engineers who have a proven track record in building scalable systems under production loads and are eager to drive technical decisions. The team at fal is dedicated to pushing the boundaries of generative media, making this an exciting opportunity in a rapidly evolving field.

Joblaze insights

Quick facts

Is the Software Engineer, Distributed Systems role remote?
No — this is an on-site role in San Francisco.
What's the salary range?
Fal lists $180,000–$250,000 for this role.
How much experience is required?
At least 3 years of relevant experience for this Software Engineer, Distributed Systems role.
Where is the role based?
Fal is hiring for this position in San Francisco.
What's the tech stack?
Joblaze extracted these technologies from the posting: Python, Rust.
What seniority level is this role?
Fal targets mid-level candidates for this position.
Is this full-time or contract?
Full-time for this Software Engineer, Distributed Systems role at Fal.

From the original posting

fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.

As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.

About this role: 

You are an experienced software engineer who thrives on building large-scale computing platforms. You have deep expertise in large scale distributed systems that deal with high complexity, a lot of traffic and data. You know how to achieve reliability and scale with minimum operational load.

Key responsibilities

  • Build our core Python/Rust platform: request routing, AI workload orchestration, scheduling, GPU autoscaling, large scale file storage, queueing, etc

  • Produce forward designs for platform evolution as we scale to 100x current traffic and need to provide low latency across the world

  • Leverage AI to an extreme level to automate the mundane parts of building complex but reliable systems

  • Profile and tune low level CPU and memory performance

Requirements

  • 3+ years experience building distributed compute and orchestration platforms in Python or Rust

  • Strong understanding of distributed systems fundamentals: consensus, scheduling, fault tolerance, capacity planning

  • Deep understanding of computational complexity and memory allocation

  • Track record of designing systems that scale under real production load

  • Experience building and using observability to drive performance and reliability decisions

  • Excellent communication and ability to drive technical decisions across teams

  • Self-starter who executes quickly, takes ownership, and constantly seeks improvement

Nice to have

  • Experience with AI/ML inference or training infrastructure

  • Experience with high-performance systems programming (async runtimes, zero-copy, memory-safe concurrency)

  • Background in building multi-tenant compute platforms

  • Understanding of networking fundamentals and performance characteristics

  • Familiarity with GPU workload characteristics and scheduling constraints

Compensation

  • $180,000-250,000 plus equity + benefits (This range is across all 3 levels Mid, Senior and Staff)

Location

  • San Francisco, CA

What we offer at fal

  • Interesting and challenging work

  • A lot of learning and growth opportunities

  • We are currently hiring in downtown San Francisco.

  • We offer relocation assistance to San Francisco.

  • Health, dental, and vision insurance (US)

  • Regular team events and offsites

U.S. EQUAL EMPLOYMENT OPPORTUNITY INFORMATION:

fal provides equal employment opportunities to applicants and employees without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status, disability, or any other classification protected by applicable law.

Similar positions

Fal
Software Engineer, Platform
Fal · San Francisco
Fal
Software Engineer, Platform
Fal · Remote - Global