Design and develop reliable software for managing large-scale GPU clusters in a dynamic engineering environment.
Posted by employer 14 hours ago
First seen on Joblaze 2 hours ago
Last verified on the company career page 2 hours ago
Skills & Technologies
What you'll build
Must have
Nice to have
Practical constraints
Requirements
Not disclosed in this posting: compensation, work arrangement, visa sponsorship.
Joblaze summary
In this role, the Embedded Software Engineer will focus on developing and deploying reliable software that manages extensive GPU clusters, ensuring optimal performance and fault tolerance. Key skills include proficiency in C, C++, or assembly, along with experience in Linux-based distributed systems. This position is ideal for candidates with a strong engineering background and problem-solving abilities, particularly those with experience in high-performance computing environments. The team operates in a collaborative, cross-disciplinary setting, emphasizing innovation in data center technology.
Joblaze insights
Quick facts
From the original posting
SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.
EMBEDDED SOFTWARE ENGINEER, DATA CENTER ENGINEERING (STARMIND)
The Data Center Engineering team designs training and inference compute clusters from first principles, pushing thermal, electrical, and optical systems to the limit of physics. We operate as a highly cross disciplinary team comprised of mechanical, electrical, optical, and software engineers, all working to maximize the efficiency, reliability, and deployment speed of terrestrial computing to serve rapidly scaling demand.
As an Embedded Software Engineer on our team, you will design, validate, and ship code that reliably and autonomously manages hundreds of thousands of coherently connected GPUs. Your software will span the entire hardware stack from individual firmware ICs up through cluster aggregation networking, ensuring our compute fabric operates at peak efficiency and gracefully handles various fault conditions. To empirically define the limit of physics in our computing hardware, you will also work closely with mechanical, electrical, and optical engineers to develop automated hardware test environments for thermal, radiation, and power testing.
RESPONSIBILITIES:
BASIC QUALIFICATIONS:
PREFERRED SKILLS AND EXPERIENCE:
ADDITIONAL REQUIREMENTS:
ITAR REQUIREMENTS:
Standard company text repeated across SpaceX's postings is omitted here.