← Back to results

GPU System Reliability Engineer Lead

Lead the RAS validation strategy for GPU server platforms in a pioneering space-based solar energy company.

Location
San Carlos, CA or Seattle, WA
Compensation
$150k–$225k/yr
Level
lead
Type
full time · Hybrid

Posted by employer 1 week ago

First seen on Joblaze 1 week ago

Last verified on the company career page 1 day ago

Skills & Technologies

Requirements

Experience
5+ years
Education
Bachelor's degree

Not disclosed in this posting: visa sponsorship.

Benefits

401k Match Daily Lunch Equity/Stock Options Paid Time Off Health Insurance Relocation Assistance Parental Leave

Joblaze summary

The GPU System Reliability Engineer Lead at Cowboy Space Corporation focuses on developing and executing strategies for reliability, availability, and serviceability (RAS) of GPU server systems in Low Earth Orbit. This role requires a strong background in hardware validation and deep knowledge of CPU and GPU architectures, along with hands-on experience in fault injection and error recovery mechanisms. Ideal candidates will have over five years of relevant experience and a solid understanding of system management interfaces. The position is part of a pioneering team dedicated to transforming energy solutions from space.

Joblaze insights

Quick facts

Is the GPU System Reliability Engineer Lead role remote?
It's hybrid — Cowboy Space Corporation expects some on-site time in San Carlos, CA or Seattle, WA.
What's the salary range?
Cowboy Space Corporation lists $150,000–$225,000 for this role.
How much experience is required?
At least 5 years of relevant experience for this GPU System Reliability Engineer Lead role.
Where is the role based?
Cowboy Space Corporation is hiring for this position in San Carlos, CA or Seattle, WA.
What's the tech stack?
Joblaze extracted these technologies from the posting: CPU, DDR, GPU, HBM, NVLink, PCIe.
What seniority level is this role?
Cowboy Space Corporation targets lead candidates for this position.
Is this full-time or contract?
Full-time for this GPU System Reliability Engineer Lead role at Cowboy Space Corporation.

From the original posting

About Cowboy Space Corp:

Mission
Cowboy Space Corporation is solving the global energy crisis by building the infrastructure for abundant, resilient space-based solar energy. We are tackling one of humanity’s most complex engineering challenges with a world-class team dedicated to delivering a revolutionary power platform. Cowboy Space Corporation is transforming how civilization powers, computes and connects - from orbit to Earth.

Background
Cowboy Space Corporation is building the infrastructure to power and connect the orbital economy. Our modular, scalable satellites collect sunlight in Low Earth Orbit to enable multiple integrated applications: transmitting energy via infrared lasers (space-to-earth and space-to-space), powering on-orbit high-performance computing clusters (GPU/TPU), and providing secure, high bandwidth optical data transport.

Current energy and data systems rely on complex logistics and outdated infrastructure. Cowboy Space Corporation overcomes these challenges by enabling direct, on-demand, secure, and scalable energy distribution and data processing from space. This will revolutionize how we operate in orbit and on Earth, supporting the rapidly expanding space industrial base, ISAM (In-Space Servicing, Assembly, and Manufacturing) operations, remote regions, and military bases.

Baiju Bhatt founded Cowboy Space Corporation in 2024. Inspired by his father’s work with NASA Langley Research Center, Baiju earned his B.S. in Physics and M.S. in Mathematics at Stanford before co-founding Robinhood, now a public company that has helped over 20 million Americans access the financial system. Cowboy Space Corporation has raised ~$365 million from Index Ventures, Interlagos, Construct Ventures, Breakthrough Energy Ventures, Andreessen Horowitz, NEA, and others.

This is an ambitious mission that demands extraordinary talent. Cowboy Space Corporation ’s team has worked at places like SpaceX, Blue Origin, Stoke Space, Astranis and NASA, and is based in San Carlos, CA. If you're ready to solve complex technical challenges and help build the most important energy company in the world, we want to hear from you.

The Role

Deploying high-performance GPU compute in Low Earth Orbit introduces a fundamentally different fault landscape than ground-based datacenter operation. This role sits at the frontier of that problem. When a fault occurs 500km above Earth, the system must detect it, classify it, contain it, and recover from it autonomously. You will own the end-to-end RAS validation strategy for GPU server systems, working directly with GPU and HBM silicon partners to analyze failures, characterize fault propagation paths, and ensure detection and recovery mechanisms function correctly. The right candidate combines deep knowledge of processor and memory architecture with hands-on system-level validation experience and the ability to drive partner engagements to resolution. This role is located in San Carlos or Seattle.

Key Responsibilities:

  • Lead RAS validation strategy and execution for GPU server platforms, including fault injection, detection coverage, and recovery verification.

  • Partner directly with GPU system designers to analyze hardware failures, review silicon errata, and align on fault handling requirements for DDR, HBM, CPU, and GPU subsystems.

  • Characterize fault propagation paths from hardware detection through firmware and OS layers, and validate that error signals are correctly classified, logged, and acted upon.

  • Validate BMC and out-of-band management visibility into hardware health events via IPMI, Redfish, and MCTP/PLDM protocols.

  • Debug complex failure modes spanning GPU and CPU architecture, memory subsystems, PCIe/NVLink fabric, and system management firmware.

  • Drive root-cause analysis for RAS failures discovered during validation and work with partners to provide input on platform design decisions that affect fault detection and serviceability.

  • Define RAS coverage metrics and maintain traceability from hardware fault models to test coverage.

  • Collaborate with firmware, software, and platform teams to validate OS-level error handling, ACPI error interfaces (EINJ, BERT,HEST), and runtime error recovery flows.

 

Basic Qualifications:

  • Bachelors degree in Electrical Engineering or a related discipline.

  • 5+ years of experience in hardware validation, platform reliability engineering, or silicon validation on server-class compute systems.

  • Deep understanding of CPU and GPU architecture, including memory subsystems (DDR, HBM), cache hierarchies, and interconnect fabrics (PCIe, NVLink, XGMI).

  • Strong knowledge of RAS concepts: error detection and correction (ECC), fault containment, error propagation, machine check architecture (MCA/MCI), and recovery mechanisms.

  • Hands-on experience with fault injection methodologies at hardware, firmware, and software levels.

  • Familiarity with system management interfaces including BMC, IPMI, Redfish, and MCTP/PLDM.

  • Experience working directly with silicon vendors or ODM partners on hardware failure analysis and RAS gap closure.

  • Strong scripting skills in Python or equivalent for test automation and log analysis.

 
 

Compensation and Benefits:

The salary range for this position is $150,000 – $225,000 annually. The actual base salary offered will depend on factors such as job-related skills, experience, qualifications, and internal equity.

  • Equity in Cowboy Space Corp.

  • Employees and their eligible dependents may enroll in medical, dental, and vision insurance

  • 401(k) retirement savings plan

  • Paid time off

  • 10 paid holidays per calendar year

  • Paid parental leave

  • Relocation assistance if applicable

  • Daily lunch in the office and a fully stocked kitchen with beverages and snacks

ITAR Requirements

  • Export Control Requirement: To conform to U.S. Government space technology export regulations, including the International Traffic in Arms Regulations (ITAR), applicants must be a U.S. citizen, lawful permanent resident of the U.S., protected individual as defined by 8 U.S.C. 1324b(a)(3), or eligible to obtain the required authorizations from the U.S. Department of State. Learn more about ITAR here.

Disclaimer

This job description is a summary of the primary duties and responsibilities of the job and position. It is not intended to be a comprehensive or all-inclusive listing of duties and responsibilities. Contents are subject to change at Cowboy Space Corp.’s discretion.

Cowboy Space Corp. is an equal employment opportunity employer. We consider individuals for employment or promotion according to their skills, abilities and experience. Cowboy Space Corp. is committed to complying with all applicable laws prohibiting discrimination based on race, color, religious creed, age, national origin, ancestry, physical, mental or developmental disability, sex (which includes pregnancy, childbirth, breastfeeding and medical conditions relating to pregnancy, childbirth or breastfeeding), veteran status, military status, marital or registered domestic partnership status, medical condition (including cancer or genetic characteristics), genetic information, gender, gender identity, gender expression, sexual orientation, as well as any other category protected by federal, state or local laws.

Similar positions

Cowboy Space Corporation
Lead Mechanical Engineer, Compute Systems
Cowboy Space Corporation · San Carlos, California
Cowboy Space Corporation
Technical Program Manager
Cowboy Space Corporation · Kent, Washington
Cowboy Space Corporation
Software Engineer
Cowboy Space Corporation · San Carlos, California
Cowboy Space Corporation
DSP Software Engineer
Cowboy Space Corporation · San Carlos, California
Cowboy Space Corporation
FPGA Engineer
Cowboy Space Corporation · San Carlos, California