Senior HPC Platform Architect

Opens nvidia.wd5.myworkdayjobs.com in a new tab

Overview

  • NVIDIA is looking for an exceptional engineer to grow and thrive alongside our HPC Infrastructure team that designs, evaluates, and optimizes the compute inrastructure powering NVIDIA's next-generation silicon design and AI workloads.
  • You will be the primary architecture reviewer and performance champion for new data center clusters being built across multiple sites globally.
  • What you'll be doing: Own data center architecture reviews for new HPC clusters — evaluating compute, storage,networking, and cooling decisions and challenging assumptions before clusters are built.
  • Serve as the BDC representative in cluster build meetings across multiple simultaneous programs (GPU compute clusters, AI infrastructure, and EDA environments), driving architecture alignment independently.
  • Analyze and validate cluster design choices across storage-to-compute distance, cross-mount latency, rack layout, and multi-site topology — and surface risk and tradeoff recommendations to leadership.
  • Lead performance benchmarking and profiling of HPC cluster infrastructure, running sanity benchmarks, regression suites, and cluster health checks to identify bottlenecks early.
  • Drive infrastructure optimization at multiple layers: scheduler-level tuning (LSF/Slurm),hardware-level tuning (GPU, CPU, networking, storage), and OS/kernel-level tuning (NUMA binding, socket binding, huge page configuration, kernel image selection).
  • Collaborate with platform and operations teams on cluster health, capacity planning, and operational readiness for tapeout milestones.
  • Partner with vendors and internal teams to evaluate new hardware, storage systems, and networking fabrics; produce architecture recommendations backed by data.
  • Continuously improve infrastructure observability, benchmarking frameworks, and architecture documentation to raise the bar across the BDC HPC platform.
  • What we need to see: B.E./B.Tech or M.Tech/M.S.
  • with 5+ years of hands-on experience in HPC infrastructure, data center architecture, systems engineering, or a senior SRE/platform engineering role at scale.
  • Deep understanding of data center architecture fundamentals: compute (CPU/GPU servers), storage (parallel file systems, NVMe, tiered storage), and high-speed networking (InfiniBand,Ethernet, NVLink).
  • Proven ability to evaluate and challenge infrastructure design decisions — including rack layout, power/cooling constraints, storage-to-compute distance, and network fabric topology.
  • Experience with OS and kernel-level performance tuning: NUMA binding, socket affinity, huge page configuration, kernel image selection, and system parameter optimization.
  • Hands-on experience with large-scale Linux HPC cluster administration using workload managers such as LSF and/or Slurm, including job scheduling optimization and resource utilization analysis.
  • Strong Linux/Unix system administration skills and proficiency in scripting (Python, Bash, or Perl) for automation and analysis.
  • Experience running HPC performance benchmarks, cluster health checks, and profiling tools to identify infrastructure bottlenecks.
  • Excellent problem-solving and communication skills — able to synthesize complex architectural tradeoffs and present clear recommendations to engineering leadership.

Tools & Skills

Sourced directly from NVIDIA’s career page

Your application goes straight to NVIDIA.

NVIDIA logo

NVIDIA

India, Bengaluru

Specialisation
Open roles at NVIDIA
2000 positions
Job ID
/job/India-Bengaluru/Senior-HPC-Platform-Architect_JR2022882

Get matched to roles like this

Upload your resume once. We’ll notify you when matching roles open up.

Join talent pool — free

Similar Other roles