About this role
About the Team & Mission
LLM Engineer (Optimization) focuses on maximizing inference performance of large language models to develop AI systems optimized for real service environments. Research and development of Inference Engine, Runtime, Compiler, and Model Optimization techniques across hardware from servers (GPU clusters) to edge and on-device. Leverage latest LLM Serving techniques and GPU/Accelerator optimization to deliver high-performance, low-latency, cost-efficient AI services.
Responsibilities
-
LLM Inference Optimization
-
Optimize inference performance (latency, throughput, memory efficiency) of large language models
-
Research and apply optimization techniques suited for various model architectures and inference environments
-
Improve performance in real service contexts such as long context and multi-turn conversations
-
Inference Engine and Runtime Development
-
Develop and optimize GPU and accelerator-based LLM inference engines
-
Utilize or improve latest inference frameworks such as vLLM, TensorRT-LLM, SGLang, llama.cpp, ONNX Runtime, MLX
-
Apply cutting-edge serving technologies such as speculative decoding and prefill-decode disaggregation
-
Model Compression and Compiler Optimization
-
Research and apply model slimming techniques like quantization (MXFP8, NVFP4, AWQ, GPTQ, etc.), pruning, distillation
-
Use compilers and kernel optimization (CUDA, Triton, TensorRT, TVM, MLIR) to improve inference performance
-
Analyze and optimize GPU memory and kernel efficiency
-
Edge AI and On-device Optimization
-
Develop optimization techniques for running LLMs efficiently on mobile, embedded, and edge devices
-
Formulate optimization strategies for CPU, GPU, NPU and other architectures
-
Ensure high performance with low power consumption under constrained resources
-
Performance Analysis and Benchmark
-
Analyze performance across various hardware and inference backends and conduct benchmarks
-
Use profiling tools such as NVIDIA Nsight Systems, Nsight Compute, py-spy to identify GPU and runtime bottlenecks and continuously improve inference performance
-
Design serving architectures that balance model quality, inference performance (latency/throughput), and cost according to given prompts and service requirements
Qualifications
-
3+ years of experience in LLM, Machine Learning Infrastructure, or Inference Optimization
-
Experience with LLM Inference Engine or AI Runtime development
-
Understanding of GPU architecture, CUDA programming, or parallel computing
-
Knowledge of model optimization techniques such as quantization, model compression, and compiler optimization
-
Experience with DL frameworks like PyTorch, ONNX, TensorRT
-
Proficiency in Python or C/C++ with software engineering skills
-
Strong problem-solving, performance analysis, and collaboration abilities
-
Preferred Qualifications
-
Experience with LLM Serving Frameworks such as vLLM, TensorRT-LLM, SGLang, llama.cpp, MLX, ONNX Runtime
-
Experience with CUDA, Triton kernels, or custom operator development
-
Experience with speculative decoding, prefill-decode disaggregation, KV cache compression, expert parallelism (MoE) and other modern LLM serving tech
-
Experience optimizing for NVIDIA GPUs (H100, A100, etc.), AMD GPUs or diverse AI accelerators
-
Edge AI and on-device LLM optimization experience (Orin, Thor, Qualcomm, Apple Silicon)
-
Kubernetes-based AI serving or large-scale GPU cluster operations
-
Understanding of LLM fine-tuning and distributed training
-
Contributions to AI Systems, ML Systems, or LLM infrastructure open source projects
-
Research experience or publications in AI/ML Systems or related conferences
-
Interview Process
-
Document screening
-
Coding and assignment tests
-
First interview (video, about 1 hour)
-
Second interview (in-person or video, about 3 hours)
-
Compensation discussion and onboarding
-
Additional Information
-
The process may change based on schedule and progress, and results will be communicated to the registered email.
-
Please omit information prohibited by hiring laws from the application.
-
For inquiries, use the official contact channel.
-
Veterans, job security beneficiaries, and disabled job seekers are facilitated according to related laws.
-
42dot does not accept unsolicited résumés from recruitment agencies and will not pay fees for unsolicited résumés.
-
If false information is found in the application, employment may be canceled.
-
After the interview process, a background check may be conducted with the candidate’s consent.
-
A 3-month probation period may apply.