Opens nvidia.wd5.myworkdayjobs.com in a new tab
Overview
- NVIDIA has been reinventing computer graphics, PC gaming, and accelerated computing for 30 years.
- It is a unique legacy of innovation that’s fueled by great technology and amazing people.
- Today, we’re tapping into the unlimited potential of AI to define the next era of computing.
- An era in which our GPU acts as the brains of computers, generative AI, robots, and self-driving cars that can understand the world.
- Doing what’s never been done before takes vision, innovation, and the world’s best talent.
- As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work.
- We are seeking a highly skilled Senior Staff SRE to join our dynamic team.
- Our company is at the forefront of technological innovation, and we are dedicated to driving efficiency and optimizing the performance of our infrastructure both on-prem and cloud.
- Join us in this exciting endeavor! What You Will Be Doing: Lead initiatives to transform IT Compute Core Team, architecture to build new service offerings across On-Prem and Cloud You will design, scale, and deploy core infrastructure services including DNS, NTP/PTP, DHCP, and LDAP.
- This includes building for performance and reliability at global scale, covering automation, monitoring, high availability, capacity planning, and lifecycle management.
- Define and implement metrics to measure the efficiency of services and drive efficiency with software and hardware optimizations (SR-IOV/ DPU) Experience with Technologies like eBPF and XDP for Observability & DDoS mitigation Collect and review system data for capacity and planning purposes, analyze capacity data and develop plans for appropriate level enterprise-wide systems, and coordinate with management personnel in implementing changes.
- Develop and maintain tools for collecting, analyzing, and visualizing data for reporting, alerting, monitoring.
- Collaborate with NVIDIA leadership, senior engineers, program managers, and product managers to develop compelling IT products and services that meet customer needs.
- What We Need To See: Bachelor’s degree in Engineering, Computer Science, Mathematics, or related field, or equivalent experience 15+ years of proven experience in compute platform engineering with a focus on automation.
- Experience in designing and deploying Containerization architectures and Distributed Systems Infrastructure Proven experience evaluating existing application architectures and identify opportunities for containerization to improve scalability, reliability, and efficiency.
- Strong analytical skills with the ability to define and track key performance metrics.
- Experience in developing tools for data analysis and performance profiling, Development with Terraform, Config Management tools.
- Proficiency in programming languages such as Go and/or Python.
- Linux OS Proficiency with Kernel Internals Experience with running large environments consisting of BareMetal Build Infrastructure Understanding of Network Protocols and Architectures (VLAN/VxLAN/SDN/BGP/Anycast) Ways To Stand Out From The Crowd: Deep understanding of other infrastructure components like, DNS, LDAP, Security Tools etc.
- Hands-on experience with containers and its implementation Deploying and Managing Services like DNS , LDAP at scale Solid understanding of microservices architecture, infrastructure as code (IaC) and configuration management tools.
- NVIDIA is widely considered to be one of the technology world’s most desirable employers.
- We have some of the most forward-thinking and passionate people on the planet working for us.
- If you're creative and autonomous, we want to hear from you! #LI-Hybrid.
Sourced directly from NVIDIA’s career page
Your application goes straight to NVIDIA.
Opens nvidia.wd5.myworkdayjobs.com in a new tab
Specialisation
Open roles at NVIDIA
2000 positions
Job ID
/job/India-Bengaluru/Senior-Staff-Site-Reliability-Engineer_JR2024324
Get matched to roles like this
Upload your resume once. We’ll notify you when matching roles open up.
Join talent pool — freeSimilar Other roles
Samsung Semiconductor
Technical Director, Large-Scale AI Model Inferencing
San Jose, California, United States|Other
Samsung Semiconductor
Technical Account Manager, DRAM Business Enablement
San Jose, California, United States|Other
Samsung Semiconductor
Staff Engineer, Storage Product Planning
San Jose, California, United States|Other
Samsung Semiconductor
Staff Engineer, Storage Business Enablement
San Jose, California, United States|Other