About this role
NVIDIA is seeking a Deep Learning Compiler Engineer - CUDA in Shanghai, China to join the Architecture group. You will contribute to advancing GPU architectures and parallel programming models by designing and implementing the domain-specific language (DSL) and the core compiler for a tile-aware GPU programming model, and by evaluating next-generation architectures to optimize performance in AI/ML workloads.
Responsibilities
- Design and implement the DSL and core compiler for a tile-aware GPU programming model on emerging GPU architectures
- Continuously innovate and iterate on the compiler architecture to optimize performance
- Investigate next-generation GPU architectures and develop DSL and compiler solutions accordingly
- Perform performance analysis on emerging AI/LLM workloads and integrate results with AI/ML frameworks
Requirements
- Masters, PhD, or equivalent experience in a relevant field (CE, CS&E, CS, AI)
- 2+ years of relevant work experience
- Excellent C/C++ programming and software engineering skills; ACM background is a plus
- Strong fundamentals in computer architecture
- Ability to abstract problems and apply methodical problem-solving approaches
- Strong compiler background; experience with MLIR/TVM/Triton/LLVM is desirable
- Good knowledge of GPU architecture and fast kernel programming is a plus
- Knowledge of LLM algorithms or a specific HPC domain is a plus
- Knowledge of multi-GPU distributed communication is a plus
- Excellent oral communication in English is a plus