Distinguished Engineer, Production Engineering, Cluster Management

Opens nvidia.wd5.myworkdayjobs.com in a new tab

What You'll Do

  • , and increase consistency, traceability, and release safety within DGX Cloud environments Partner closely with platform teams, hardware and provider engineering, service owners, and other Production Engineering leaders to identify repeated friction and convert it into durable improvements in software, processes, and operational interfaces Raise the engineering bar for operability, resilience, scalability, and performance across cluster operations through build leadership, architecture review, and technical standards What we need to see: BS, MS, or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent experience 18+ years of experience building and operating large-scale distributed systems, infrastructure platforms, or production environments Confirmed company-level technical leadership at principal, distinguished, or equivalent scope in production engineering, SRE, infrastructure software, or cloud platforms Consistent track record of defining operating models, architectural direction, and engineering standards across multiple technical domains and organizations Consistent record leading large, cross-team technical efforts from concept through production, including aligning collaborators, navigating for clarity and delivering measurable outcomes Deep experience with one or more of these areas: Kubernetes-based production systems, infrastructure automation, or distributed systems operations.
  • Strong software engineering skills in languages such as Python, Go, or similar low-level programming languages Deep understanding of distributed systems, Linux, networking, containers, and production reliability concerns Experience crafting operational workflows, APIs, service interfaces, or automation frameworks that become the standard way teams run production systems Strong architectural judgment and a validated history of simplifying complex operational problems through reusable software, clear technical strategy, and durable engineering direction Ways to stand out from the crowd: Defined the structural foundation for a large, heterogeneous infrastructure environment spanning multiple platforms or providers Established widely used operating standards, architectures, APIs, or workflows that improved reliability, operability, or performance at company scale Built automation and engineering interfaces that connect platform teams, infrastructure teams, and service owners into a consistent production system Experience improving production readiness, restoration, runtime safety, or release quality for large-scale infrastructure Equally comfortable setting technical strategy, reviewing architecture at scale, writing code, and driving adoption across organizational boundaries This role is purposely assigned to the cross-domain production operating model for DGX Cloud capacity.
  • It does not involve a shared platform-software function for typical automation services.
  • Success is achieved by making sure the larger DGX Cloud cluster estate functions optimally in a live environment.
  • The position calls for strong technical leadership, consistent workflows, clear limits, and productive cross-team engineering coordination. #LI-Hybrid Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.
  • The base salary range is 320,000 USD - 488,750 USD.
  • You will also be eligible for equity and benefits .
  • Applications for this job will be accepted at least until September 7, 2026.
  • This posting is for an existing vacancy.
  • NVIDIA uses AI tools in its recruiting processes.
  • As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Tools & Skills

Languages

Sourced directly from NVIDIA’s career page

Your application goes straight to NVIDIA.

NVIDIA logo

NVIDIA

2 Locations

Specialisation
Open roles at NVIDIA
2000 positions
Job ID
/job/US-CA-Santa-Clara/Distinguished-Engineer--Production-Engineering--Cluster-Management_JR2024993

Get matched to roles like this

Upload your resume once. We’ll notify you when matching roles open up.

Join talent pool — free

Similar Other roles