AI Infrastructure Engineer

Opens intel.wd1.myworkdayjobs.com in a new tab

About This Role

  • We are looking for a performance-obsessed AI Infrastructure Engineer to push LLM inference to its absolute limits on Intel's next-generation GPU architectures.
  • In this role, you will dive deep into the inference stack and redefine peak performance.
  • You will work end-to-end across the stack: profiling bottlenecks, writing custom GPU kernels, and upstreaming your optimizations directly into industry-standard serving frameworks like vLLM and SGLang.
  • Your optimizations will be instrumental in unlocking the full potential of Intel hardware for state-of-the-art generative AI workloads.
  • What You Will Do • Drive Inference Performance: Own the end-to-end optimization pipeline for running state-of-the-art LLMs on Intel GPUs. • Deep Stack Optimization: Profile, diagnose, and resolve cross-stack performance bottlenecks. • Kernel Development and Integration: Design, write, and optimize custom high-performance kernels for critical attention mechanisms, MoE, quantization, and operator fusions. • Open Source Leadership: Upstream your architectural improvements and hardware backends directly into open-source repositories like vLLM, SGLang, and PyTorch, acting as a bridge between the hardware teams and the open-source community. • Shape the Hardware Roadmap: Apply roofline analysis and systematic profiling to decompose bottlenecks.
  • You will partner with our architecture and compiler teams to shape future GPU roadmaps based on real-world GenAI workload data. • Show passion about AI infrastructure and performance optimization.

Requirements

  • Bachelors Degree in Computer Science, Software Engineering, Artificial Intelligence/Machine Learning, or related field and 4+ years experience, Masters Degree and 3+ years, OR PhD. • 3+ years of relevant software engineering experience in GPU computing, AI systems, or high-performance computing (HPC). • Proficiency in modern C++ and Python.
  • You are comfortable reading and modifying complex systems-level code.
  • Preferred Qualifications • Understanding of CPU/GPU architecture. • Understanding of modern LLM architectures and inference paradigms: attention mechanisms, KV caching, continuous batching, speculative decoding, and prefill-decode disaggregation. • Prior open-source contributions to inference engines (vLLM, SGLang, PyTorch, llama.cpp). • Hands-on experience writing and optimizing custom GPU kernels using Triton, SYCL, CUDA/CUTLASS, or other DSLs. • Experience with scale-out inference orchestration across multi-node topologies. • You leverage AI coding agents daily to accelerate your own workflow and benchmark generation.
  • Your expertise will play a vital role in advancing Intel's AI technology.
  • We invite you to bring your skills, experience, and passion for AI to make an impact-apply today.

Benefits

  • We offer a total compensation package that ranks among the best in the industry.
  • It consists of competitive pay, stock bonuses, and benefit programs which include health, retirement, and vacation.
  • Find out more about the benefits of working at Intel .
  • Annual Salary Range for jobs which could be performed in the US: $170,500.00-315,490.00 USD The range displayed on this job posting reflects the minimum and maximum target compensation for the position across all US locations.
  • Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training.
  • Your recruiter can share more about the specific compensation range for your preferred location during the hiring process.

Tools & Skills

Sourced directly from Intel’s career page

Your application goes straight to Intel.

Intel logo

Intel

4 Locations

Specialisation
Open roles at Intel
656 positions
Job ID
/job/US-California-Santa-Clara/AI-Infrastructure-Engineer_JR0286233

Get matched to roles like this

Upload your resume once. We’ll notify you when matching roles open up.

Join talent pool — free

Similar Other roles