About this role
About the Role We are seeking a Principal Architect to define and own the architecture of our Unified Edge AI Engine. This role designs the software backbone that orchestrates the entire lifecycle of video analytics on embedded hardware and serves as the bridge between Systems Engineering and AI Research. The candidate will build a modular, graph-based execution engine that balances data throughput with limited compute resources to run advanced AI models in real-time across diverse hardware platforms. What You'll Do
- Architect a modular, high-throughput Video & AI Pipeline Engine with a dynamic DAG framework for assembling AI models, logic nodes, and image processing blocks.
- Design heterogeneous resource orchestration: scheduling across CPU, GPU, DSP, and NPU; manage PCIe bandwidth, memory bandwidth, and thermal constraints to prevent system stalls.
- Define architectural support for AI model optimization: quantization (INT8/FP16) and network pruning; ensure the engine can natively handle optimized model formats, balancing inference speed, memory footprint, and accuracy.
- Architect zero-copy data transport mechanisms (DMA, Shared Memory, Ring Buffers) to pass high-resolution video frames and tensor data between pipeline stages without CPU intervention.
- Create a hardware-agnostic HAL that allows the engine to scale from low-power cameras to high-performance AI processors with minimal code changes.
- Act as a Systems Architect to guide decisions on resource allocation, memory management, and hardware selection when needed. What We're Looking For
- Bachelor’s degree or higher in Embedded AI, Computer Science, Computer Engineering, or related field; 10+ years of industry experience in Embedded Software Architecture or High-Performance Computing; proven track record of designing complex software engines for Video Processing, Computer Vision, or Autonomous Systems.
- Mastery of Modern C++ (14/17/20) and architectural patterns for high-concurrency, real-time systems; deep understanding of Producer-Consumer models, lock-free queues, and multi-threaded synchronization.
- Strong knowledge of AI model optimization techniques: Quantization (PTQ/QAT), Weight Pruning, and Knowledge Distillation; experience integrating these optimized models into C++ runtimes (e.g., TensorRT, SNPE/QNN, OpenVINO).
- Deep understanding of SoC architectures (Qualcomm, Ambarella, NVIDIA); ability to analyze how Cache Coherency, Bus Arbitration, and Memory Controllers impact AI performance.
- Proficiency with system profilers (Perf, eBPF, Nsight Systems) to visualize the "Critical Path" of the pipeline and optimize instruction-level performance.
- Experience designing or heavily customizing media pipelines (similar to GStreamer, MediaPipe, or DeepStream).
- Excellent communication and leadership as a System Architect with a track record of making critical resource allocation decisions.
- Travel: None; Relocation: None.