Reliability Engineer

Opens intel.wd1.myworkdayjobs.com in a new tab

About This Role

  • Join us to help build the next generation of AI hardware solutions.
  • You will be part of a highly skilled, agile team developing cutting-edge hardware for the AI domain, where we push the boundaries of what silicon can do for emerging AI workloads.
  • With a startup-like culture, we move quickly and give engineers the opportunity to drive significant technical and business impact.
  • We are continuously developing modern and effective working methods, including hands-on adoption of AI tools throughout the chip development flow.
  • Mission: Define and own the pod-level reliability specifications that ensure the availability, resilience, and serviceability of a large-scale data center across hardware, thermal, and operational dimensions.

What You'll Do

  • Define and maintain pod-level reliability/availability specs and targets (MTBF, AFR, RAS) for compute, memory, storage, network, power, and cooling subsystems.
  • Translate system/SLA requirements into pod and subsystem level reliability specs; flow requirements down to silicon, platform, and facilities teams.
  • Lead FMEA, root-cause analysis, and pod fleet failure-data analytics to drive corrective actions and spec updates.
  • Architect RAS features (ECC, memory mirroring, predictive failure, telemetry) and graceful degradation/redundancy against pod-level specs.
  • Partner with facilities on pod power/cooling redundancy (N+1, 2N), thermal margins, and disaster-recovery readiness.
  • Establish HALT/HASS, burn-in, qualification processes; track field returns and KPIs against pod spec.

Requirements

  • BS/MS/PhD in EE/ME Reliability or related; and/or at least 4-6 yrs experience.
  • Experience authoring and owning reliability specs and requirement flow-down.
  • Strong RAS, FMEA, statistical reliability (Weibull, FIT) skills.
  • Experience with large-scale fleet telemetry and thermal/power redundancy.

Nice to Have

  • AI cluster operations, data analytics (Python/SQL).

Benefits

  • We offer a total compensation package that ranks among the best in the industry.
  • It consists of competitive pay, stock bonuses, and benefit programs which include health, retirement, and vacation.
  • Find out more about the benefits of working at Intel .
  • Annual Salary Range for jobs which could be performed in the US: $122,440.00-232,190.00 USD The range displayed on this job posting reflects the minimum and maximum target compensation for the position across all US locations.
  • Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training.
  • Your recruiter can share more about the specific compensation range for your preferred location during the hiring process.

Tools & Skills

Languages

Sourced directly from Intel’s career page

Your application goes straight to Intel.

Intel logo

Intel

2 Locations

Specialisation
Open roles at Intel
677 positions
Job ID
/job/US-Massachusetts-Beaver-Brook/Reliability-Engineer_JR0285885

Get matched to roles like this

Upload your resume once. We’ll notify you when matching roles open up.

Join talent pool — free

Similar Other roles