Opens nvidia.wd5.myworkdayjobs.com in a new tab
What You'll Do
- will include building AI/HPC infrastructure for new and existing customers.
- Support operational and reliability aspects of large-scale AI clusters, focusing on performance at scale, real-time monitoring, logging, and alerting.
- Engage in and improve the whole lifecycle of services—from inception and design through deployment, operation, and refinement.
- Develop tooling to automate and manage of large-scale infrastructure environments, to automate operational monitoring and alerting, and to enable self-service consumption of resources.
- Deploy monitoring solutions for the servers, network and storage.
- Perform troubleshooting bottom up from bare metal, operating system, software stack and application level.
- Being a technical resource, develop, re-define and document standard methodologies to share with customer and internal teams Support activities and engage in POCs/POVs for future improvements.
- Experience with system administration on Linux systems required (CentOS, RHEL, and Ubuntu preferred) What We Need to See: BS/MS/PhD or equivalent experience in Computer Science, Data Science, Electrical/Computer Engineering, Physics, Mathematics, other Engineering fields with at least 5+ years’ work or research experience in networking fundamentals, TCP/IP stack, and data centre compute architecture.
- Advance knowledge of HPC, AI & EVPN, BGP, OSPF, VXLAN protocols.
- Deep understanding of DC architecture fundamentals such as compute, storage (PFS) & InfiniBand, Ethernet, NVLink.
- Experience running HPC performance benchmarks, cluster health checks, and profiling tools to identify infrastructure bottlenecks.
- Python programming, bash scripting experience and Advance Linux knowledge.
- Extensive experience delivering automated network provisioning and comfortable with automation and configuration management tools including Jenkins, Ansible, Puppet.
- Possess solid working knowledge of Ethernet/InfiniBand/RDMA core principles.
- Excellent customer-facing and communication skills (verbal and written in both languages), enabling effective engagement with customers, partners, and cross-functional teams across the India region.
- Listening skills in English are critical.
- Willingness to travel Ways To Stand Out from The Crowd: Knowledge of CPU and/or GPU architecture including of Kubernetes, container related microservice technologies.
- Background with RDMA (InfiniBand or RoCE) fabrics.
- Linux or Networking Certifications (e.g., CCNP, CCIE) or NVIDIA-related certifications.
- Deep Knowledge on observability stack and build experience.
Sourced directly from NVIDIA’s career page
Your application goes straight to NVIDIA.
Opens nvidia.wd5.myworkdayjobs.com in a new tab
Specialisation
Open roles at NVIDIA
2000 positions
Job ID
/job/India-Gurugram/Senior-Solutions-Architect--Networking-and-Compute-Infrastructure_JR2025322
Get matched to roles like this
Upload your resume once. We’ll notify you when matching roles open up.
Join talent pool — freeSimilar Other roles
Samsung Semiconductor
Workplace Project Manager, Project Management Office, Workplace Solutions Group
San Jose, California, United States|Other
Samsung Semiconductor
Staff Engineer, SSD Storage and Systems Architecture
San Jose, California, United States|Other
Samsung Semiconductor
Staff Engineer, SRAM Circuit Design
San Jose, California, United States|Other
Samsung Semiconductor
Staff Engineer, AI Platform Enablement and Emerging Applications
San Jose, California, United States|Other