Head of Platform Product Reliability

JOB INFO

Apply

Apply for this job directly on SHORTList.

Referral

Share your custom referral link for this job with qualified candidates. Earn the referral you lead to a hire.

COMPANYEtched

We are seeking a highly technical and execution-focused Head of Platform Product Reliability to lead reliability engineering across Etched's server, rack, and datacenter platform products.

This role owns system-level product reliability from architecture through fleet deployment. You will define reliability strategy, qualification methodologies, accelerated stress testing programs, failure analysis processes, and long-term reliability standards for complex AI infrastructure systems. This team focuses specifically on product reliability engineering for platform hardware and deployed systems — ensuring every Etched product ships with the reliability profile that enterprise and hyperscale customers demand.

You will work cross-functionally with Platform Engineering, Mechanical Engineering, Thermal, Firmware, Manufacturing, Supply Chain, Datacenter Operations, and Program teams to ensure Etched products achieve exceptional reliability at scale.

Key Responsibilities:

  • Define and own the end-to-end reliability strategy for AI servers, accelerator platforms, rack systems, and datacenter infrastructure, from design requirements through field deployment.
  • Establish reliability requirements, qualification standards, and validation methodologies that scale across product generations.
  • Build and institutionalize reliability engineering processes spanning the full product lifecycle:
    • EVT / DVT / PVT qualification gates and exit criteria.
    • Accelerated life testing (ALT) and accelerated stress testing (AST).
    • Environmental testing: temperature, humidity, altitude, contamination.
    • HALT / HASS programs for design margin and production screening.
    • Vibration, shock, and transportation stress testing.
    • Power cycling, thermal cycling, and long-duration soak testing.
  • Lead root-cause investigations for reliability failures surfaced during development, manufacturing, and field deployment, driving corrective actions across hardware, firmware, thermal, and mechanical domains.
  • Develop comprehensive system reliability models including MTBF projections, FIT rate analysis, Weibull lifetime modeling, component derating methodologies, and reliability growth tracking.
  • Ensure reliability is considered early, partnering with Platform Engineering architects and design leads so reliability requirements shape decisions before they become expensive to change.
  • Work closely with ODMs, JDMs, contract manufacturers, and component suppliers to validate and enforce long-term platform reliability commitments.
  • Build fleet reliability infrastructure: telemetry analysis pipelines, field feedback loops, and monitoring frameworks that give Etched visibility into deployed system health at scale.
  • Drive reliability signoff criteria and lead product release readiness reviews across engineering and program teams.
  • Build and lead a high-performing product reliability engineering organization — hiring, developing, and retaining technical talent as the company scales.

You may be a good fit if you have (Must-have qualifications):

  • BS, MS, or PhD in Electrical Engineering, Mechanical Engineering, Reliability Engineering, or a related technical field.
  • 10+ years of reliability engineering experience in hardware-centric organizations, with meaningful time spent on complex systems rather than component-level work.
  • Experience leading reliability programs for one or more of:
    • AI accelerator or GPU-class compute systems.
    • Hyperscale or cloud server infrastructure.
    • Networking platforms, storage systems, or rack-scale infrastructure.
  • Deep understanding of system-level failure mechanisms — including thermal, power delivery, mechanical, and connector/interconnect failure modes — and how design decisions affect long-term field reliability.
  • Hands-on experience with FMEA, Weibull analysis, HALT/HASS, qualification planning, failure analysis methodologies, and reliability statistics and modeling.
  • A track record of driving cross-functional root-cause investigations in fast-moving hardware organizations where schedule pressure is.