Opens nvidia.wd5.myworkdayjobs.com in a new tab
Overview
- NVIDIA is leading company of AI computing.
- At NVIDIA, our employees are passionate about AI, HPC , VISUAL, GAMING.
- SA team is more focusing to bring NVIDIA new technology into difference industries.
- This role focuses on NVIDIA Inference Microservices (NIM), inference / RL rolloutperformance, and AI workflow enablement for LLM, VLM, and other generative AI workloads.
- It is a highly hands-on position at the intersection of model optimization, inference infrastructure, and customer solution delivery.
- What you’ll be doing: Drive the implementation, deployment, and optimization of NVIDIA Inference Microservices (NIM) solutions for enterprise and industry AI workloads.
- Package and serve open-source, NVIDIA, and customer-proprietary models through NIM with standardized, containerized APIs for on-premises, cloud, and hybrid environments.
- Optimize high-volume inference and rollout workloads for LLMs and VLMs.
- Evaluate and tune the NIM models.
- Deliver technical projects, demos and client support tasks as directed by the Solution Architecture Leadership.
- Provide technical support and guidance to customers, facilitating the adoption and implementation of NVIDIA technologies and products.
- Collaborate with cross-functional teams to enhance and expand our AI solutions portfolio.
- What we need to see: Master’s degree or higher in Computer Science, Machine Learning, Electrical Engineering, Mathematics, or a related technical field, or equivalent experience.
- 2+ years of hands-on experience in machine learning engineering, applied research, LLM/VLM inference, or RL rollout.
- Production-quality Python and PyTorch skills, including distributed GPU training, solution, profiling, debugging, memory optimization.
- Working knowledge of transformer architectures, performance optimization, rollout sampling strategies, structured generation, and model-quality evaluation.
- Strong written and verbal communication skills, with the ability to collaborate effectively across research, engineering, infrastructure, product, and customer-facing teams.
- Ways to stand out from the crowd: Publications, open-source contributions, or significant technical projects, LLM/VLM, agent systems.
- Experience applying programmatic verification, simulators, compilers, execution sandboxes, APIs, or external tools as reward sources for model training.
- Familiar with oss RL framework such as SLIME, Nemo-RL.
- Familiarity with enterprise AI deployment, customer adaptation, or adapting foundation models to specialized vertical domains.
Sourced directly from NVIDIA’s career page
Your application goes straight to NVIDIA.
Opens nvidia.wd5.myworkdayjobs.com in a new tab
Specialisation
Open roles at NVIDIA
2000 positions
Job ID
/job/China-Shanghai/NIM-Solution-Architect_JR2025729
Get matched to roles like this
Upload your resume once. We’ll notify you when matching roles open up.
Join talent pool — freeSimilar Other roles
Samsung Semiconductor
Specialist, Foundry Sales Operations
San Diego, California, United States (12265 El Camino Real); San Jose, California, United States|Other
Samsung Semiconductor
Senior Manager, Memory Marketing MI (Market Intelligence)
San Jose, California, United States|Other
Micron Technology
Intern, Photo Manufacturing Data Analytics and AI
Fab 10N/X, Singapore|Other
Micron Technology
Intern - Probe Automation
Fab 10N/X, Singapore|Other