Opens nvidia.wd5.myworkdayjobs.com in a new tab
Overview
- NVIDIA is looking for an exceptional engineer to grow and thrive alongside our HPC Infrastructure team that designs, evaluates, and optimizes the compute inrastructure powering NVIDIA's next-generation silicon design and AI workloads.
- You will be the primary architecture reviewer and performance champion for new data center clusters being built across multiple sites globally.
- What you'll be doing: Own data center architecture reviews for new HPC clusters — evaluating compute, storage,networking, and cooling decisions and challenging assumptions before clusters are built.
- Serve as the BDC representative in cluster build meetings across multiple simultaneous programs (GPU compute clusters, AI infrastructure, and EDA environments), driving architecture alignment independently.
- Analyze and validate cluster design choices across storage-to-compute distance, cross-mount latency, rack layout, and multi-site topology — and surface risk and tradeoff recommendations to leadership.
- Lead performance benchmarking and profiling of HPC cluster infrastructure, running sanity benchmarks, regression suites, and cluster health checks to identify bottlenecks early.
- Drive infrastructure optimization at multiple layers: scheduler-level tuning (LSF/Slurm),hardware-level tuning (GPU, CPU, networking, storage), and OS/kernel-level tuning (NUMA binding, socket binding, huge page configuration, kernel image selection).
- Collaborate with platform and operations teams on cluster health, capacity planning, and operational readiness for tapeout milestones.
- Partner with vendors and internal teams to evaluate new hardware, storage systems, and networking fabrics; produce architecture recommendations backed by data.
- Continuously improve infrastructure observability, benchmarking frameworks, and architecture documentation to raise the bar across the BDC HPC platform.
- What we need to see: B.E./B.Tech or M.Tech/M.S.
- with 5+ years of hands-on experience in HPC infrastructure, data center architecture, systems engineering, or a senior SRE/platform engineering role at scale.
- Deep understanding of data center architecture fundamentals: compute (CPU/GPU servers), storage (parallel file systems, NVMe, tiered storage), and high-speed networking (InfiniBand,Ethernet, NVLink).
- Proven ability to evaluate and challenge infrastructure design decisions — including rack layout, power/cooling constraints, storage-to-compute distance, and network fabric topology.
- Experience with OS and kernel-level performance tuning: NUMA binding, socket affinity, huge page configuration, kernel image selection, and system parameter optimization.
- Hands-on experience with large-scale Linux HPC cluster administration using workload managers such as LSF and/or Slurm, including job scheduling optimization and resource utilization analysis.
- Strong Linux/Unix system administration skills and proficiency in scripting (Python, Bash, or Perl) for automation and analysis.
- Experience running HPC performance benchmarks, cluster health checks, and profiling tools to identify infrastructure bottlenecks.
- Excellent problem-solving and communication skills — able to synthesize complex architectural tradeoffs and present clear recommendations to engineering leadership.
Sourced directly from NVIDIA’s career page
Your application goes straight to NVIDIA.
Opens nvidia.wd5.myworkdayjobs.com in a new tab
Specialisation
Open roles at NVIDIA
2000 positions
Job ID
/job/India-Bengaluru/Senior-HPC-Platform-Architect_JR2022882
Get matched to roles like this
Upload your resume once. We’ll notify you when matching roles open up.
Join talent pool — free