Senior Software Engineer, Automation Infrastructure

Opens nvidia.wd5.myworkdayjobs.com in a new tab

Overview

  • NVIDIA's Performance Lab (PerfLab) builds the systems and automation used to evaluate the performance and quality of accelerated computing and AI workloads.
  • We turn complex benchmark experiments into reliable, scalable, and reproducible workflows that help engineering teams make better decisions faster.
  • We are looking for an experienced and highly self-motivated System Software Engineer to help build the next generation of PerfLab's benchmark infrastructure.
  • You will independently own meaningful platform components and take projects from problem discovery and technical design through production deployment and adoption.
  • The ideal candidate enjoys finding important engineering problems, understanding their root causes, and using technology to create simple, reusable solutions.
  • You will collaborate with NVIDIA teams around the world and work on evolving areas such as large language models, agentic AI, accelerated computing, and other emerging AI workloads.
  • What You'll Be Doing: Design, build, and maintain reusable software, services, and workflows that automate benchmark definition, execution, result collection, validation, and reporting across local, cluster, and cloud-native environments.
  • Carry out performance testing and analysis as needed.
  • Develop a deep understanding of existing performance workflows and infrastructure, and translate real-world needs into scalable, user-friendly solutions.
  • Improve the reliability, scalability, observability, and reproducibility of large benchmark campaigns.
  • Diagnose complex issues across applications, Linux systems, containers, distributed jobs, compute resources, networking, and storage.
  • Build strong partnerships with performance engineers, QA teams, product teams, and other customers; agree on goals, organize execution, and drive new ideas and projects from concept through adoption.
  • Contribute to technical designs, code reviews, documentation, and internal or open-source infrastructure projects.
  • Apply AI-assisted automation where it can meaningfully improve benchmark creation, failure triage, data analysis, or engineering productivity.
  • What We Need to See: Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience in practice.
  • 5+ years of relevant software engineering experience in system software, infrastructure, developer platforms, distributed systems, or production automation.
  • Strong Python programming and software engineering skills, with experience building production-quality tools, services, or automation frameworks.
  • Solid understanding of Linux and system-level concepts such as processes, concurrency, networking, storage, resource management, and failure handling.
  • Hands-on experience with containers and at least one workload orchestration, scheduling, or distributed computing platform, as well as designing reliable systems or pipelines with clear interfaces, testing, observability, and recovery behavior.
  • Working knowledge of machine learning, AI, or accelerated-computing workloads and an interest in how their performance and quality are evaluated.
  • Strong analytical and problem-solving abilities, with the capacity to manage multiple priorities effectively and adapt in a dynamic, fast-changing environment.
  • High self-motivation and demonstrated ability to identify valuable problems, turn ambiguous needs into clear technical plans, and drive projects through delivery and adoption.
  • Excellent communication and organizational skills, with the ability to align cross-functional stakeholders, collaborate with globally distributed teams, and move new ideas toward concrete outcomes.
  • Ways to Stand Out From the Crowd: Experience with GPU or AI infrastructure, distributed training or inference, model evaluation, or performance benchmarking.
  • Experience building workflow engines, schedulers, experiment platforms, test frameworks, or developer infrastructure.
  • Experience operating distributed or cloud-native systems in production, including performance profiling, capacity analysis, resource scheduling, or multi-node workloads.
  • Practical experience using AI agents, tool-calling, or coding agents to automate engineering workflows.
  • Contributions to open-source infrastructure or developer tooling, or a track record of initiating automation that reduced manual work, improved reliability, or helped other teams move faster.
  • We have some of the most forward-thinking and hardworking people in the world working for us.
  • If you are creative, autonomous, and passionate about building systems that make complex AI performance work repeatable and scalable, we want to hear from you.

Tools & Skills

Languages

Sourced directly from NVIDIA’s career page

Your application goes straight to NVIDIA.

NVIDIA logo

NVIDIA

China, Shanghai

Specialisation
Open roles at NVIDIA
1999 positions
Job ID
/job/China-Shanghai/Senior-Software-Engineer--Automation-Infrastructure_JR2022054-1

Get matched to roles like this

Upload your resume once. We’ll notify you when matching roles open up.

Join talent pool — free

Similar Other roles