About this role
About the Role
We are seeking an AI Compiler & Inference Engineer to support the deployment of customer AI inference workloads to a K3 CPU/NPU environment, with a focus on measured equivalence, reliable deployment, and safe handling of unsupported operators and models.
Key Responsibilities
- Profile the K3 NPU compiler/runtime environment, supported operators, model formats, quantisation modes, and deployment constraints.
- Implement model intake and conversion for agreed model families, including supported ONNX/PyTorch-exportable models.
- Design authorised quantisation and calibration workflows while preserving model provenance.
- Build operator-support matrices and deterministic fallback/rejection behaviour.
- Create reference outputs and statistical/numerical equivalence tests.
- Integrate model conversion into the One-Click planner and artefact manifest.
- Optimise inference packaging and deployment without hidden vendor-tool dependencies.
- Document customer limitations, performance considerations, and recovery/reconversion procedures.
- Transfer model-support runbooks and test assets to the permanent engineering team.
Requirements
-
5+ years of ML systems/inference engineering experience with production model deployment.
-
Strong experience with ONNX, PyTorch export, model compilers/runtimes, and quantisation.
-
Experience with numerical validation and calibration workflows.
-
Familiarity with edge, accelerator NPU, or heterogeneous inference environments.
-
Strong Python/C++ and Linux deployment skills.
-
Experience developing CI-testable model pipelines.
-
Understanding of the distinction between model conversion and model retraining, including cases where original learned weights cannot simply be recreated.
-
Role Details
-
Contract: 15 October – 14 December 2026
-
Extension: Optional 2–4 weeks
-
Location: Overseas / Remote
-
Search: Global, with priority for engineers experienced in model compiler/runtime deployment.