Senior Site Reliability Engineer in Test, SDET

Opens nvidia.wd5.myworkdayjobs.com in a new tab

Overview

  • At NVIDIA, we are at the forefront of technological innovation, pushing the boundaries of AI and accelerated computing.
  • Our team in Shanghai, China is looking for a Senior Site Reliability Engineer focused on Test Environment Management to join us.
  • This is an opportunity to build and operate highly reliable test infrastructure, CI/CD systems, and environments that power validation of NVIDIA enterprise offerings.
  • If you are passionate about reliability engineering, test infrastructure, and AI-scale systems, this role is for you.
  • What you’ll be doing Design, build, operate, and continuously improve reliable, scalable test environments and automation infrastructure that support validation of NVIDIA enterprise offerings.
  • Own end-to-end CI/CD pipelines using GitLab CI, GitHub Actions, and ArgoCD (GitOps) — including pipeline design, reliability, performance, and progressive delivery of test workloads.
  • Manage Software Bills of Materials (SBOMs): generation, continuous monitoring, vulnerability correlation, policy enforcement, and integration into CI/CD and release gate.
  • Provision, scale, observe, and lifecycle-manage ephemeral and long-lived test environments (Kubernetes-based and hybrid) with strong emphasis on isolation, reproducibility, and rapid recovery.
  • Define and drive reliability practices for test systems: SLIs/SLOs/error budgets for test environments and pipelines, toil reduction, chaos/resilience testing of infrastructure, and automated remediation.
  • Collaborate closely with development and platform teams to triage environment and pipeline failures, perform root-cause analysis, verify fixes, and continuously harden test infrastructure.
  • Apply AI/ML/Agentic techniques and internal tools to accelerate environment provisioning, flaky-test detection, capacity planning, anomaly detection, and overall Quality Assurance velocity.
  • What we need to see MS or PhD in Computer Science, related field, or equivalent experience with 8+ years of professional experience in Site Reliability Engineering, Test Environment Management, CI/CD platform engineering, or software testing infrastructure.
  • Strong proficiency with Linux, shell scripting, and Python (or equivalent automation languages).
  • Hands-on experience designing and operating CI/CD systems with GitLab CI and/or GitHub Actions, practical experience with ArgoCD (or equivalent GitOps tooling) for CD of applications and infrastructure.
  • Solid background in containerization and orchestration (Docker, Kubernetes) and virtualization technologies.
  • Deep understanding of SRE principles: SLIs/SLOs, error budgets, incident response, postmortems, toil elimination, and reliability engineering for complex distributed systems.
  • Experience building and operating test environments (ephemeral, multi-tenant, or production-like) with focus on reliability, isolation, and rapid turnaround.
  • Strong knowledge of QA principles and how test infrastructure enables high-quality software delivery.
  • Comfort working with AI/LLM-related workloads and toolings.
  • Excellent problem-solving skills, clear written and verbal communication, and the ability to collaborate across engineering teams.
  • Self-motivated, proactive, and passionate about learning new technology at scale.
  • Ways to stand out from the crowd Experience operating large-scale Kubernetes platforms and GitOps workflows in production or high-stakes test environments.
  • Background in software supply-chain security, SBOM tooling ecosystems, vulnerability management, and policy enforcement (OPA/Gatekeeper, Kyverno, etc.).
  • Hands-on work with NVIDIA GPU hardware, multi-GPU environments, or accelerated computing infrastructure.
  • Experience with parallel programming, high-performance computing, or large-scale AI model training/inference test harnesses.
  • Track record of applying AI/observability techniques to detect flaky tests, optimize environment utilization, or automate root-cause analysis.
  • Prior experience defining and driving reliability programs (error budgets, chaos engineering, capacity forecasting) for CI/CD or test platforms.
  • NVIDIA is widely considered one of the technology world’s most desirable employers.
  • We have some of the most brilliant people on the planet working for us.
  • If you’re creative, autonomous, and excited about making test infrastructure as reliable as the products it validates, we want to hear from you!.

Tools & Skills

Sourced directly from NVIDIA’s career page

Your application goes straight to NVIDIA.

NVIDIA logo

NVIDIA

China, Shanghai

Specialisation
Open roles at NVIDIA
2000 positions
Job ID
/job/China-Shanghai/Senior-Site-Reliability-Engineer-in-Test--SDET_JR2023203-1

Get matched to roles like this

Upload your resume once. We’ll notify you when matching roles open up.

Join talent pool — free

Similar Other roles