Opens nvidia.wd5.myworkdayjobs.com in a new tab
Overview
- NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for over 30 years.
- It’s a unique legacy of innovation fueled by great technology—and dynamic people.
- Today, we’re tapping into the unlimited potential of AI to define the next era of computing—an era in which our GPUs act as the brains of computers, robots, and self-driving cars that can understand the world.
- Doing what’s never been done before takes vision, innovation, and the world’s best talent.
- NVIDIANs immerse themselves in a diverse, supportive environment that encourages everyone to do their best work.
- Join the team and see how you can make a lasting impact on the world.
- Slurm is a critical workload manager for many of the world’s most demanding AI and high-performance computing environments.
- We are looking for a Technical Support Engineer dedicated to supporting Slurm for NVIDIA’s customers.
- You will join a dedicated team of Slurm subject-matter guides, owning sophisticated support cases and helping customers run reliable, efficient, and highly scalable clusters.
- This role requires extensive production experience with Slurm and the ability to diagnose issues across the scheduler and the surrounding Linux, networking, storage, authentication, database, and GPU infrastructure.
- What you’ll be doing: Own Slurm support cases from initial investigation through resolution for customers running production AI and HPC clusters.
- Diagnose complex problems involving slurmctld, slurmd , slurmdbd, job scheduling, node management, resource allocation, accounting, authentication, and high availability.
- Solve Slurm configuration and policy features, including partitions, reservations, priorities, fair-share, quality of service, backfill, preemption, GRES/TRES, cgroups, and job constraints.
- Investigate performance, reliability, and scalability issues using logs, diagnostic data, configuration analysis, reproductions, and source-level debugging when required.
- Isolate problems across Slurm and its surrounding dependencies, including Linux, MUNGE, databases, networking, parallel storage, containers, GPUs, and cluster-management systems.
- Advise customers on Slurm configuration, upgrades, operational practices, managing system resources, and safe recovery from production incidents.
- Collaborate with engineering teams by producing clear technical descriptions, reproducible test cases, and well-supported defect reports.
- Develop guides, knowledge-base articles, diagnostic tools, and internal training that strengthen Slurm expertise across the support organization.
- What we need to see: BS degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 5+ years of hands-on experience administering and supporting Slurm in production HPC or AI environments including business-critical outage incidents.
- Expert-level understanding of Slurm architecture, daemons, configuration, scheduling behavior, accounting, resource management, and failure modes.
- Capacity to identify sophisticated Slurm incidents independently and guide them to a technically sound resolution.
- In-depth Linux system-administration and solve experience, including systemd, cgroups, authentication, networking, and database-backed services.
- Experience operating Slurm across multi-user clusters with complex scheduling policies and heterogeneous compute resources.
- Strong analytical and research skills, showing proficiency in distinguishing Slurm defects from configuration, integration, infrastructure, and workload problems.
- Excellent written and verbal communication skills, including the ability to turn detailed technical findings into clear explanations and actionable recommendations.
- Ways to stand out from the crowd: Experience supporting large-scale Slurm environments containing thousands of nodes or GPUs.
- Experience diagnosing scheduler performance, job-throughput, controller-load, and database-scaling issues.
- Familiarity with Slurm source code, plugins, SPANK, Lua job-submit plugins, or upstream issue investigation.
- Experience with containers and HPC integration technologies such as Pyxis, Enroot, Apptainer, or Singularity.
- Previous experience integrating Slurm with NVIDIA Base Command Manager, Bright Cluster Manager, or another cluster-management platform.
Tools & Skills
Languages
Sourced directly from NVIDIA’s career page
Your application goes straight to NVIDIA.
Opens nvidia.wd5.myworkdayjobs.com in a new tab
Specialisation
Open roles at NVIDIA
2000 positions
Job ID
/job/India-Remote/Technical-Support-Engineer---Slurm_JR2025521
Get matched to roles like this
Upload your resume once. We’ll notify you when matching roles open up.
Join talent pool — freeSimilar Other roles
Samsung Semiconductor
Technical Director, Large-Scale AI Model Inferencing
San Jose, California, United States|Other
Samsung Semiconductor
Technical Account Manager, DRAM Business Enablement
San Jose, California, United States|Other
Samsung Semiconductor
Staff Engineer, Storage Product Planning
San Jose, California, United States|Other
Samsung Semiconductor
Staff Engineer, Storage Business Enablement
San Jose, California, United States|Other