About this role
AI Compiler & Inference Engineer
Location: Remote Employment Type: Contract Work Arrangement: Remote Sector: Information Technology & Software Experience Level: Senior (5-8 years) Application Deadline: October 14, 2026
Job Description We are seeking a highly skilled AI Compiler & Inference Engineer to join our team on a fixed-term contract basis. This role is critical for supporting the deployment of customer AI inference workloads onto a K3 CPU/NPU environment. Your focus will be on ensuring measured equivalence, reliable deployment, and the safe handling of unsupported operators and models. Key responsibilities include profiling the K3 NPU compiler/runtime, supported operators, model formats, and deployment constraints. You will implement model intake and conversion for agreed model families, including supported ONNX/PyTorch-exportable models. The role also involves designing authorized quantisation and calibration workflows while preserving model provenance, and building operator-support matrices with deterministic fallback/rejection behavior. You will create reference outputs and statistical/numerical equivalence tests, integrate model conversion into the One-Click planner, and optimize inference packaging and deployment.
Required Skills
- AI Compiler/Inference Engineering, K3 CPU/NPU, ONNX, PyTorch Export, Model Compilers, Quantisation, Numerical Validation, Python, C++, Linux, CI/CD
Key Responsibilities
- Profile the K3 NPU compiler/runtime environment, supported operators, model formats, quantisation modes, and deployment constraints. Implement model intake and conversion for agreed model families, including supported ONNX/PyTorch-exportable models. Design authorised quantisation and calibration workflows while preserving model provenance. Build operator-support matrices and deterministic fallback/rejection behaviour. Create reference outputs and statistical/numerical equivalence tests. Integrate model conversion into the One-Click planner and artefact manifest. Optimise inference packaging and deployment without hidden vendor-tool dependencies. Document customer limitations, performance considerations, and recovery/reconversion procedures. Transfer model-support runbooks and test assets to the permanent engineering team.
Qualifications
- 5+ years of ML systems/inference engineering experience with production model deployment. Strong experience with ONNX, PyTorch export, model compilers/runtimes, and quantisation. Experience with numerical validation and calibration workflows. Familiarity with edge, accelerator NPU, or heterogeneous inference environments. Strong Python/C++ and Linux deployment skills. Experience developing CI-testable model pipelines. Understanding of the distinction between model conversion and model retraining, including cases where original learned weights cannot simply be recreated.