Senior HPC Cluster Engineer - AI, ML

Opens nvidia.wd5.myworkdayjobs.com in a new tab

Overview

  • NVIDIA is a pioneer in accelerated computing, known for inventing the GPU and driving breakthroughs in gaming, computer graphics, high-performance computing, and artificial intelligence.
  • Our technology powers everything from generative AI to autonomous systems, and we continue to shape the future of computing through innovation and collaboration.
  • Within this mission, our team, Managed AI Superclusters (MARS) builds and scales the infrastructure, platforms, and tools that enable researchers and engineers to develop the next generation of AI/ML systems.
  • By joining us, you’ll help design solutions that power some of the world’s most advanced computing workloads.
  • NVIDIA is looking for a Senior AI/ML HPC Cluster Engineer to join our MARS team.
  • You will provide leadership and strategic guidance on the management of large-scale HPC systems including the deployment of compute, networking, and storage.
  • You will be working with a team of passionate and skilled engineers across NVIDIA that are continuously working to provide better tools to build and manage this infrastructure.
  • Ideal candidate is strong in building and maintaining distributed clusters, driving improvements, and has the ability to understand researcher computing needs.
  • What you'll be doing: Provide leadership in systems administration and service delivery on our AI/HPC fleet by coordinating system upgrades, responding to incidents, and delivering reliability improvements.
  • Collaborate closely with global teams to deliver a world class user experience in AI and HPC research.
  • Own day-to-day operations of production AI/HPC clusters, ensuring system health, user satisfaction, and efficient resource utilization.
  • Develop and improve our ecosystem around GPU-accelerated computing including developing scalable automation solutions.
  • Build and maintain heterogeneous AI/ML clusters on-premises and in the cloud.
  • Create and cultivate customer and cross-team relationships to meet user evolving user needs.
  • Support our researchers to run their workloads including performance analysis and optimizations Analyze and optimize cluster efficiency, job fragmentation, and GPU waste to meet internal SLA targets.
  • Conduct root cause analysis and suggest corrective action Proactively find and fix issues before they occur.
  • Lead SEV triage and postmortems for reliability incidents affecting users or infrastructure.
  • Participate in on-call rotation and incident response for critical production GPU clusters.
  • What we need to see: Bachelor’s degree in Computer Science, Electrical Engineering or related field or equivalent experience Minimum 5 years of experience designing and operating large scale compute infrastructure Experience with AI/HPC advanced job schedulers, such as Slurm, K8s, PBS, RTDA, BCM, or LSF Proficient in administering Centos/RHEL and/or Ubuntu Linux distributions Solid understanding of cluster configuration management tools (BCM, Terraform, Ansible, Puppet, Salt, etc.), container technologies (Docker, Singularity, Podman, Shifter, Charliecloud), Python programming, and bash scripting.
  • Applied experience with AI/HPC workflows that use MPI Experience analyzing and tuning performance for a variety of AI/HPC workloads.
  • Passion for continual learning and staying ahead of emerging technologies and effective approaches in the HPC and AI/ML infrastructure fields.
  • Ways to stand out from the crowd: Background with NVIDIA GPUs, CUDA Programming, NCCL and MLPerf benchmarking Experience with AI/ML concepts, algorithms, models, and frameworks (PyTorch, Tensorflow) Experience with InfiniBand with IPoIB and RDMA Understanding of fast, distributed storage systems such as Lustre and GPFS for AI/HPC workloads.

Tools & Skills

Languages

Sourced directly from NVIDIA’s career page

Your application goes straight to NVIDIA.

NVIDIA logo

NVIDIA

3 Locations

Specialisation
Open roles at NVIDIA
2000 positions
Job ID
/job/India-Bengaluru/Senior-HPC-Cluster-Engineer---AI--ML_JR2021817-1

Get matched to roles like this

Upload your resume once. We’ll notify you when matching roles open up.

Join talent pool — free

Similar Other roles