Senior Software Engineer

Opens nvidia.wd5.myworkdayjobs.com in a new tab

Overview

  • We are seeking a Senior Software Engineer with strong infrastructure expertise to design, build, and operate the next generation of our enterprise Observability, Automation, and AI-driven Reliability Platform.
  • This role will build highly scalable distributed systems and platform services spanning Storage, Compute, Network, VMware, OpenShift, and bare-metal infrastructure.
  • The engineer will help transform infrastructure operations from reactive monitoring and manual remediation to proactive, predictive, and AI-driven autonomous operations.
  • What You Will Be Doing: Design, build, and operate distributed software platforms for enterprise observability, telemetry, automation, and infrastructure reliability at large scale.
  • Develop reusable platform services, APIs, automation frameworks, and control planes that enable self-service, reduce operational toil, and automate infrastructure operations across multiple engineering teams.
  • Build scalable telemetry and event-processing systems spanning metrics, logs, traces, events, topology, and alerts, with the performance and efficiency to process billions of infrastructure signals.
  • Build intelligent and AI-native reliability capabilities, including agentic workflows for anomaly detection, forecasting, root-cause analysis, automated debugging, and closed-loop remediation.
  • Drive technical architecture and engineering direction across Storage, Compute, Network, and Platform domains, solving complex and ambiguous problems that span multiple teams.
  • Engineer for production at scale, with strong focus on software quality, scalability, security, performance, observability, maintainability, and operational readiness.
  • Provide technical leadership and mentorship, influence engineering standards and architecture decisions, and deliver measurable improvements in reliability, MTTR, operational toil, engineering productivity, and infrastructure efficiency.
  • What We Need To See: Bachelor's or Master's degree in Computer Science, Engineering, or equivalent practical experience, with 10+ years of software engineering, SRE, infrastructure, or distributed-systems experience and demonstrated technical leadership.
  • Strong software engineering expertise in Go, Python, or equivalent languages, with experience designing and building production-grade distributed systems, platform services, APIs, and automation.
  • Proven experience owning complex software/platform initiatives across multiple teams or infrastructure domains, from architecture and implementation through adoption and measurable impact.
  • Deep understanding of distributed systems, event-driven architectures, microservices, APIs, and high-throughput data processing, including technologies such as Kafka, NATS, gRPC, or equivalent.
  • Strong experience with modern observability and telemetry platforms, including OpenTelemetry, Prometheus, VictoriaMetrics, Vector, Loki, Grafana, ClickHouse, or equivalent technologies.
  • Strong SRE and infrastructure knowledge across Kubernetes/OpenShift, VMware, bare-metal, storage, networking, and/or cloud environments, with experience using Terraform, Ansible, or equivalent automation technologies.
  • Demonstrated ability to solve ambiguous problems, influence technical direction without direct authority, mentor engineers, establish engineering standards, and deliver measurable operational and business outcomes.
  • Ways To Stand Out From The Crowd: Experience building software and reliability platforms for large-scale on-premises infrastructure, particularly Storage, Compute, Networking, VMware, and Kubernetes/OpenShift.
  • Deep understanding of storage and infrastructure telemetry, including IOPS, latency, NVMe health, SAN/NAS topology, block/object storage, and infrastructure failure domains.
  • Experience building self-healing systems, automated remediation, predictive operations, or autonomous SRE capabilities.
  • Production experience applying Generative AI, AIOps, LLMs, or Agentic AI to incident triage, RCA, operational intelligence, debugging, or remediation; experience with LangChain, LlamaIndex, AutoGen, or equivalent is a plus.
  • Experience building high-performance platform services using FastAPI, gRPC, or equivalent technologies.

Tools & Skills

Languages

Sourced directly from NVIDIA’s career page

Your application goes straight to NVIDIA.

NVIDIA logo

NVIDIA

India, Bengaluru

Specialisation
Open roles at NVIDIA
1998 positions
Job ID
/job/India-Bengaluru/Senior-Software-Engineer_JR2024294

Get matched to roles like this

Upload your resume once. We’ll notify you when matching roles open up.

Join talent pool — free

Similar Other roles