← Back to results

Head of Platform Product Reliability

Lead reliability engineering for AI infrastructure products at a cutting-edge hardware startup in San Jose.

Location
San Jose, United States
Compensation
Not disclosed
Level
lead
Type
full time · On-site

Posted by employer 4 months ago

First seen on Joblaze 10 hours ago

Last verified on the company career page 10 hours ago

Apply at Etched → Save job Scanned from etched.com

What you'll build

  • Define reliability strategy for AI servers
  • Establish reliability requirements and validation methodologies
  • Lead root-cause investigations for reliability failures
  • Develop system reliability models
  • Build fleet reliability infrastructure

Must have

  • BS, MS, or PhD in Electrical Engineering, Mechanical Engineering, Reliability Engineering, or related field
  • 10+ years of reliability engineering experience
  • Experience leading reliability programs for AI accelerator or GPU-class compute systems
  • Deep understanding of system-level failure mechanisms
  • Hands-on experience with FMEA and reliability statistics
  • Strong technical judgment

Nice to have

  • Experience with liquid-cooled systems
  • Direct experience supporting hyperscale or cloud datacenter deployments
  • Demonstrated experience building a reliability organization
  • Familiarity with fleet telemetry systems
  • Experience working closely with ODM or JDM partners
  • Background in high-speed digital systems

Requirements

Experience
10+ years
Education
Bachelor's degree

Not disclosed in this posting: compensation, visa sponsorship.

Benefits

Wellness Benefits Daily Meals Housing Subsidy Health Insurance Relocation Assistance

Joblaze summary

The Head of Platform Product Reliability at Etched is responsible for overseeing the reliability engineering of the company's server and datacenter products, ensuring they meet high standards from design through deployment. This role requires extensive experience in reliability engineering, particularly with complex AI infrastructure systems, and proficiency in methodologies like FMEA and Weibull analysis. Ideal candidates will have a strong technical background and a proven track record in leading cross-functional teams in fast-paced environments. Etched emphasizes a collaborative culture where engineering and research intersect, fostering innovation in hardware for frontier intelligence.

Joblaze insights

  • Listed today — first seen on Joblaze September 21, 2026. Last confirmed on Etched's careers page September 21, 2026.
  • AI/ML appears in 5.2% of 557 comparable lead other roles in United States; Reliability Engineering appears in 0.4% of 557 comparable lead other roles in United States.

Quick facts

Is the Head of Platform Product Reliability role remote?
No — this is an on-site role in San Jose, United States.
How much experience is required?
At least 10 years of relevant experience for this Head of Platform Product Reliability role.
Where is the role based?
Etched is hiring for this position in San Jose, United States.
What's the tech stack?
Joblaze extracted these technologies from the posting: AI servers, AI/ML, Hardware, Reliability Engineering, datacenter infrastructure.
What seniority level is this role?
Etched targets lead candidates for this position.
Is this full-time or contract?
Full-time for this Head of Platform Product Reliability role at Etched.

From the original posting

Job Summary

We are seeking a highly technical and execution-focused Head of Platform Product Reliability to lead reliability engineering across Etched's server, rack, and datacenter platform products.

This role owns system-level product reliability from architecture through fleet deployment. You will define reliability strategy, qualification methodologies, accelerated stress testing programs, failure analysis processes, and long-term reliability standards for complex AI infrastructure systems. This team focuses specifically on product reliability engineering for platform hardware and deployed systems — ensuring every Etched product ships with the reliability profile that enterprise and hyperscale customers demand.

You will work cross-functionally with Platform Engineering, Mechanical Engineering, Thermal, Firmware, Manufacturing, Supply Chain, Datacenter Operations, and Program teams to ensure Etched products achieve exceptional reliability at scale.

Key Responsibilities

  • Define and own the end-to-end reliability strategy for AI servers, accelerator platforms, rack systems, and datacenter infrastructure, from design requirements through field deployment

  • Establish reliability requirements, qualification standards, and validation methodologies that scale across product generations

  • Build and institutionalize reliability engineering processes spanning the full product lifecycle:

    • EVT / DVT / PVT qualification gates and exit criteria

    • Accelerated life testing (ALT) and accelerated stress testing (AST)

    • Environmental testing: temperature, humidity, altitude, contamination

    • HALT / HASS programs for design margin and production screening

    • Vibration, shock, and transportation stress testing

    • Power cycling, thermal cycling, and long-duration soak testing

  • Lead root-cause investigations for reliability failures surfaced during development, manufacturing, and field deployment, driving corrective actions across hardware, firmware, thermal, and mechanical domains

  • Develop comprehensive system reliability models including MTBF projections, FIT rate analysis, Weibull lifetime modeling, component derating methodologies, and reliability growth tracking

  • Ensure reliability is considered early, partnering with Platform Engineering architects and design leads so reliability requirements shape decisions before they become expensive to change

  • Work closely with ODMs, JDMs, contract manufacturers, and component suppliers to validate and enforce long-term platform reliability commitments

  • Build fleet reliability infrastructure: telemetry analysis pipelines, field feedback loops, and monitoring frameworks that give Etched visibility into deployed system health at scale

  • Drive reliability signoff criteria and lead product release readiness reviews across engineering and program teams

  • Build and lead a high-performing product reliability engineering organization — hiring, developing, and retaining technical talent as the company scales

You may be a good fit if you have (Must-have qualifications)

  • BS, MS, or PhD in Electrical Engineering, Mechanical Engineering, Reliability Engineering, or a related technical field

  • 10+ years of reliability engineering experience in hardware-centric organizations, with meaningful time spent on complex systems rather than component-level work

  • Experience leading reliability programs for one or more of:

    • AI accelerator or GPU-class compute systems

    • Hyperscale or cloud server infrastructure

    • Networking platforms, storage systems, or rack-scale infrastructure

  • Deep understanding of system-level failure mechanisms — including thermal, power delivery, mechanical, and connector/interconnect failure modes — and how design decisions affect long-term field reliability

  • Hands-on experience with FMEA, Weibull analysis, HALT/HASS, qualification planning, failure analysis methodologies, and reliability statistics and modeling

  • A track record of driving cross-functional root-cause investigations in fast-moving hardware organizations where schedule pressure is real and accountability is high

  • Strong technical judgment — capable of making defensible tradeoffs between reliability targets, cost, schedule, and performance without losing sight of customer expectations

  • Excellent communication skills and the credibility to influence design decisions with engineering leads, program managers, and executive stakeholders

Strong candidates may also have experience with (Nice-to-have qualifications)

  • Experience with liquid-cooled systems, high-density power delivery, or thermal management for high-power AI infrastructure

  • Direct experience supporting hyperscale or cloud datacenter deployments at scale, including customer-facing reliability commitments and SLA management

  • Demonstrated experience building a reliability organization from early-stage — establishing processes, tooling, and team norms in environments without established infrastructure

  • Familiarity with fleet telemetry systems, large-scale field reliability analytics, and data-driven approaches to proactive reliability management

  • Experience working closely with ODM or JDM partners in Taiwan or broader Asia, including NPI support and on-site qualification engagement

  • Background in high-speed digital systems, GPU compute platforms, or accelerator-based architectures — with an understanding of how these affect system-level reliability behavior

Benefits

  • Medical, dental, and vision packages with generous premium coverage

    • $500 per month credit for waiving medical benefits

  • Housing subsidy of $2,500 per month for those living within walking distance of the office

  • Daily lunch and dinner in our office

  • Unlimited compute budget subject to ROI justification

 

Standard company text repeated across Etched's postings is omitted here.

Similar positions

Etched
Product Engineer, Silicon Validation
Etched · San Jose, United States
Etched
Infrastructure Software Engineer
Etched · San Jose, United States
Etched
Security Engineer
Etched · San Jose, United States
Etched
Applied AI Engineer, Manufacturing Execution
Etched · San Jose, United States
Etched
Head of Supercomputing
Etched · San Jose, United States