About this role
Senior AI Inference Engineer
At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.
Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.
THE ROLE:
AMD is looking for an experienced software engineer to join our growing team. As a key contributor you will be part of a leading team to drive and enhance AMD’s abilities to deliver the highest quality, industry-leading technologies to market.
THE PERSON:
If you are passionate about AI/ML frameworks, model optimization, and efficient AI deployment on accelerator hardware, this is your opportunity. We are looking for an engineer with strong software development skills, hands-on experience in LLM inference and training optimization, and the ability to use AI-assisted tools effectively while maintaining strong technical judgment and code quality.
You will work with AI researchers, framework engineers, compiler/runtime teams, hardware experts, and customer-facing teams to build robust optimization software for real-world AI workloads across AMD platforms.
KEY RESPONSIBILITIES
- Design, implement, and maintain model optimization features for AI workloads on AMD hardware platforms.
- Develop quantization, low-precision, and compression capabilities for CNN, Transformer, LLM, and multimodal models.
- Build production-quality Python tools, libraries, APIs, and framework components.
- Support training and fine-tuning workflows for optimized models.
- Analyze and debug accuracy, latency, memory usage, and deployment tradeoffs.
- Collaborate across framework, compiler, runtime, hardware, and application teams.
PREFERRED EXPERIENCE:
- Experience in one or more areas: AI/ML frameworks, vLLM/SGLang, model optimization, quantization, training workflows, runtime integration, or accelerator-oriented deployment.
- Hands-on experience with ML frameworks such as PyTorch, ONNX/ONNX Runtime, or similar.
- Strong C++/Python development and debugging skills.
- Experience building production-quality tools, libraries, APIs, or framework components.
- Solid understanding of CNN, LLM, or multimodal model architectures.
- Ability to reason about accuracy, performance, latency, memory footprint, and deployment constraints.
- Demonstrated ability to use AI-assisted tools effectively for software development, debugging, technical exploration, and productivity improvement.
- Strong software engineering fundamentals and ability to work with geographically distributed teams.
ACADEMIC CREDENTIALS:
- BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Machine Learning, or related technical fields.