About this role
Job title: AI Compiler & Inference Engineer
Overview: Oversee deployment of customer AI inference workloads to a K3 CPU/NPU environment, ensuring precise numerical equivalence, reliable deployment, and safe operator handling.
Responsibilities: Deploy AI inference workloads to target environments; ensure numerical equivalence and safe operator handling; build and maintain CI-testable model pipelines; manage model conversion versus retraining considerations; preserve original weights where applicable.
Requirements: 5+ years ML systems/inference engineering experience with production model deployment; strong expertise in ONNX, PyTorch export, model compilers/runtimes, and quantisation workflows; hands-on experience with numerical validation, calibration, and edge/NPU/heterogeneous inference environments; proficiency in Python/C++ and Linux deployment; experience building CI-testable model pipelines; deep understanding of model conversion vs. retraining and original weight preservation.
Contract Duration: 15 October – 14 December 2026 (Optional 2–4 weeks extension).
Notes: Global candidate search with priority given to model compiler/runtime deployment experts.