Software Engineer — Distributed LLM Inference Systems

Opens intel.wd1.myworkdayjobs.com in a new tab

About This Role

  • The Role and Impact: As a Software Engineer on Intel’s Artificial Intelligence Frameworks team, you will contribute to designing, developing, and optimizing distributed inference systems for large language models.
  • Your day-to-day work will involve implementing distributed inference algorithms, optimizing model execution and communication, transforming neural network models, and developing software components that improve inference performance across diverse hardware architectures.
  • You may work on areas such as disaggregated serving, request scheduling, KV cache management, parallel execution, and efficient communication between inference components.
  • By collaborating with researchers and engineers, you will play a key role in advancing Intel's AI capabilities and ensuring industry-leading solutions.
  • Business Group: Intel's Artificial Intelligence Frameworks team is dedicated to empowering transformative AI solutions by developing and optimizing software frameworks for machine learning and deep learning.
  • This group works on enhancing the performance of AI applications across diverse computing hardware backends while contributing to open-source communities.
  • As part of Intel, this team supports the mission to drive technological innovation and deliver impactful AI advancements globally.
  • Key Responsibilities Design, develop, and optimize distributed LLM inference systems and related AI framework components.
  • Implement distributed algorithms, including model/data parallel frameworks and asynchronous communication for deep learning.
  • Develop and optimize components such as request schedulers, model workers, communication layers, and KV cache management mechanisms.
  • Profile distributed inference workloads to identify computation, communication, memory, and scheduling bottlenecks.
  • Collaborate with component teams to improve end-to-end latency, throughput, scalability, and resource utilization.
  • Contribute high-quality code, tests, and documentation to internal and open-source projects while following industry engineering standards.

Requirements

  • – Master’s degree in computer science, Artificial Intelligence, Software Engineering, or a related field, with 0-1 years of hands-on experience demonstrated through internships, academic projects, coursework, or training.
  • Proficiency in Python and modern C++ programming.
  • Foundational knowledge of deep learning and AI frameworks, such as PyTorch.
  • Experience debugging and optimizing software for performance.
  • Basic understanding of machine learning algorithms and techniques.
  • Strong problem-solving skills and the ability to learn unfamiliar systems quickly.
  • Preferred Qualifications Experience or project exposure related to distributed LLM inference and serving.
  • Experience in contributing to open-source projects or collaborating within open-source ecosystems.
  • Understanding of LLM inference concepts such as prefill and decode, KV cache management, continuous batching, parallelism strategies, and disaggregated serving.
  • Familiarity with inference engines or serving frameworks such as vLLM, SGLang, TensorRT-LLM, or similar technologies.
  • Knowledge of large language models and inference optimization techniques.
  • Knowledge of AI Agent architecture and execution workflows, including tool calling, planning, memory, context management, and multi-agent coordination.
  • Effective communication skills, including fluency in written and spoken English.
  • Take the opportunity to be part of Intel's journey in redefining AI frameworks and enabling groundbreaking innovations.
  • Take the opportunity to be part of Intel's journey in redefining AI frameworks and enabling groundbreaking innovations.
  • Your contributions will shape the future of AI software and its real-world impact.
  • Apply now and be a part of advancing transformative AI capabilities.

Tools & Skills

Languages

Sourced directly from Intel’s career page

Your application goes straight to Intel.

Intel logo

Intel

PRC, Shanghai

Specialisation
Open roles at Intel
627 positions
Job ID
/job/PRC-Shanghai/Software-Engineer---Distributed-LLM-Inference-Systems_JR0286399

Get matched to roles like this

Upload your resume once. We’ll notify you when matching roles open up.

Join talent pool — free

Similar Other roles