Opens nvidia.wd5.myworkdayjobs.com in a new tab
Overview
- NVIDIA is a pioneer in accelerated computing, known for inventing the GPU and driving breakthroughs in gaming, computer graphics, high-performance computing, and artificial intelligence.
- Our technology powers everything from generative AI to autonomous systems, and we continue to shape the future of computing through innovation and collaboration.
- Within this mission, our team, Managed AI Superclusters (MARS) builds and scales the infrastructure, platforms, and tools that enable researchers and engineers to develop the next generation of AI/ML systems.
- By joining us, you’ll help design solutions that power some of the world’s most advanced computing workloads.
- NVIDIA is looking for a Senior AI/ML HPC Cluster Engineer to join our MARS team.
- You will provide leadership and strategic guidance on the management of large-scale HPC systems including the deployment of compute, networking, and storage.
- You will be working with a team of passionate and skilled engineers across NVIDIA that are continuously working to provide better tools to build and manage this infrastructure.
- Ideal candidate is strong in building and maintaining distributed clusters, driving improvements, and has the ability to understand researcher computing needs.
- What you'll be doing: Provide leadership in systems administration and service delivery on our AI/HPC fleet by coordinating system upgrades, responding to incidents, and delivering reliability improvements.
- Collaborate closely with global teams to deliver a world class user experience in AI and HPC research.
- Own day-to-day operations of production AI/HPC clusters, ensuring system health, user satisfaction, and efficient resource utilization.
- Develop and improve our ecosystem around GPU-accelerated computing including developing scalable automation solutions.
- Build and maintain heterogeneous AI/ML clusters on-premises and in the cloud.
- Create and cultivate customer and cross-team relationships to meet user evolving user needs.
- Support our researchers to run their workloads including performance analysis and optimizations Analyze and optimize cluster efficiency, job fragmentation, and GPU waste to meet internal SLA targets.
- Conduct root cause analysis and suggest corrective action Proactively find and fix issues before they occur.
- Lead SEV triage and postmortems for reliability incidents affecting users or infrastructure.
- Participate in on-call rotation and incident response for critical production GPU clusters.
- What we need to see: Bachelor’s degree in Computer Science, Electrical Engineering or related field or equivalent experience Minimum 5 years of experience designing and operating large scale compute infrastructure Experience with AI/HPC advanced job schedulers, such as Slurm, K8s, PBS, RTDA, BCM, or LSF Proficient in administering Centos/RHEL and/or Ubuntu Linux distributions Solid understanding of cluster configuration management tools (BCM, Terraform, Ansible, Puppet, Salt, etc.), container technologies (Docker, Singularity, Podman, Shifter, Charliecloud), Python programming, and bash scripting.
- Applied experience with AI/HPC workflows that use MPI Experience analyzing and tuning performance for a variety of AI/HPC workloads.
- Passion for continual learning and staying ahead of emerging technologies and effective approaches in the HPC and AI/ML infrastructure fields.
- Ways to stand out from the crowd: Background with NVIDIA GPUs, CUDA Programming, NCCL and MLPerf benchmarking Experience with AI/ML concepts, algorithms, models, and frameworks (PyTorch, Tensorflow) Experience with InfiniBand with IPoIB and RDMA Understanding of fast, distributed storage systems such as Lustre and GPFS for AI/HPC workloads.
Sourced directly from NVIDIA’s career page
Your application goes straight to NVIDIA.
Opens nvidia.wd5.myworkdayjobs.com in a new tab
Specialisation
Open roles at NVIDIA
2000 positions
Job ID
/job/India-Bengaluru/Senior-HPC-Cluster-Engineer---AI--ML_JR2021817-1
Get matched to roles like this
Upload your resume once. We’ll notify you when matching roles open up.
Join talent pool — freeSimilar Other roles
Samsung Semiconductor
Sr. Manager, Foundry Sales Business Development
San Jose, California, United States|Other
Micron Technology
Sr. Engineer, Photo Track PE
Taichung - Fab 16, Taiwan|Other
Micron Technology
広島大学プロジェクト研究 手続き用
Hiroshima - Fab 15, Japan|Other
Micron Technology
SOFTWARE DEVELOPMENT ENGINEER
Taichung - AATT, Taiwan|Other