Senior Software Engineer, Automation Infrastructure

Opens nvidia.wd5.myworkdayjobs.com in a new tab

Overview

  • NVIDIA's Performance Lab (PerfLab) builds the systems and automation used to evaluate the performance and quality of accelerated computing and AI workloads.
  • We turn complex benchmark experiments into reliable, scalable, and reproducible workflows that help engineering teams make better decisions faster.
  • We are looking for an experienced and highly self-motivated System Software Engineer to help build the next generation of our infrastructure.
  • You will independently own meaningful platform components and lead projects from problem discovery and technical design through production deployment and adoption.
  • The ideal candidate enjoys finding important engineering problems, understanding their root causes, and using technology to create simple, reusable solutions.
  • You will collaborate with NVIDIA teams around the world and influence how we evaluate evolving areas such as large language models, agentic AI, accelerated computing, and other emerging AI workloads.
  • What You'll Be Doing Define the technical direction and architecture for major areas of PerfLab's benchmark infrastructure, translating evolving business and engineering needs into clear roadmaps and scalable platform capabilities.
  • Lead the design and implementation of reusable software, services, and workflows that automate benchmark definition, execution, result collection, validation, and reporting across local, cluster, and cloud-native environments.
  • Remain hands-on with performance testing and analysis, developing a deep understanding of existing workflows and using that knowledge to guide platform investments and technical decisions.
  • Establish engineering approaches that improve the reliability, scalability, observability, maintainability, and reproducibility of large benchmark campaigns.
  • Lead the diagnosis of complex, cross-layer issues spanning applications, Linux systems, containers, distributed jobs, compute resources, networking, and storage.
  • Build strong partnerships with performance engineers, QA teams, product teams, and other customers; create alignment across organizations and drive high-impact ideas and projects from concept through adoption.
  • Provide technical leadership through architecture and code reviews, clear decision-making, high engineering standards, and mentoring of other engineers.
  • Identify emerging technologies, including AI-assisted automation, and determine where they can deliver meaningful improvements in benchmark creation, failure triage, data analysis, or engineering productivity.
  • Contribute to the strategy and development of internal and open-source infrastructure projects, and help grow their adoption across teams.
  • What We Need to See Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent practical experience.
  • 6+ years of relevant software engineering experience in system software, infrastructure, developer platforms, distributed systems, or production automation.
  • Strong Python programming and software engineering expertise, with a track record of building and operating production-quality tools, services, or automation frameworks.
  • Solid understanding of Linux and system-level concepts such as processes, concurrency, networking, storage, resource management, and failure handling.
  • Extensive experience with containers and workload orchestration, scheduling, or distributed computing platforms, including the design of reliable systems with clear interfaces, testing, observability, and recovery behavior.
  • Knowledge of machine learning, AI, or accelerated-computing workloads and experience reasoning about their performance, quality, and operational tradeoffs.
  • Proven technical leadership on sophisticated, multi-functional projects, including defining architecture, managing technical risk, resolving ambiguity, and driving solutions through delivery and adoption.
  • Strong analytical and problem-solving abilities, with the judgment to prioritize effectively, manage multiple initiatives, and adapt in a dynamic, constantly evolving environment.
  • Excellent communication, organizational, and influencing skills, with the ability to align globally distributed collaborators and drive decisions.
  • A record of mentoring engineers, elevating engineering quality, and helping teams make better technical decisions.
  • Ways to Stand Out From the Crowd Experience leading major initiatives in GPU or AI infrastructure, distributed training or inference, model evaluation, or performance benchmarking.
  • Experience architecting workflow engines, schedulers, experiment platforms, test frameworks, or developer infrastructure used by multiple teams.
  • Experience operating distributed or cloud-native systems at scale, including performance profiling, capacity analysis, resource scheduling, or multi-node workloads.
  • Practical experience establishing AI-agent, tool-calling, or coding-agent strategies that improved engineering workflows at team or organizational scale.
  • Demonstrated success turning loosely defined, cross-organizational problems into durable platforms or programs with measurable engineering impact.
  • We have some of the most forward-thinking and hardworking people in the world working for us.
  • If you are creative, autonomous, and passionate about providing technical leadership while remaining hands-on in building systems that make complex AI performance work repeatable and scalable, we want to hear from you.

Tools & Skills

Languages

Sourced directly from NVIDIA’s career page

Your application goes straight to NVIDIA.

NVIDIA logo

NVIDIA

China, Shanghai

Specialisation
Open roles at NVIDIA
2000 positions
Job ID
/job/China-Shanghai/Senior-Software-Engineer--Automation-Infrastructure_JR2022807

Get matched to roles like this

Upload your resume once. We’ll notify you when matching roles open up.

Join talent pool — free

Similar Other roles