About this role
Thesis Worker 30 hp - Adaptive Hierarchical Inference with Foundation models
30 hp - Adaptive Hierarchical Inference with Foundation models for Commercial-Vehicle Edge-Cloud Systems
Introduction
Thesis work is an excellent way to get closer to Scania and build relationships for the future. Many of today's employees began their Scania career with their degree project.
Background
TRATON GROUP is one of the world’s leading commercial vehicle manufacturers. Its brands include Scania, MAN, International and Volkswagen Truck & Bus. The group offers light commercial vehicles, trucks and buses, supported by financing, charging and digital logistics services. Through its global operations, production sites, and sales and service networks, TRATON has access to diverse vehicle platforms, real-world data and fleet environments. This provides a strong basis for developing and validating solutions for more sustainable and efficient transportation.
Within TRATON’s Cloud and Embedded Platform department, we develop solutions for connected vehicles and IoT platforms. This work supports TRATON’s growing focus on communication, digital services and smart transport. Advanced data analysis is an important part of these developments. Modern commercial vehicles generate diverse data from vehicle telemetry, time-series sensors, cameras, radar, LiDAR and operational logs. High-performance units and embedded platforms make it possible to process much of this data close to the vehicle, reducing latency, network dependence and unnecessary data transfer. Foundation models could support predictive maintenance, anomaly detection, diagnostic assistance, multimodal perception and reasoning. However, the most capable models require cloud-scale computing resources, while onboard systems face strict limits on computation, memory, energy and communication. This research therefore focuses on hierarchical inference. A compact model performs real-time analysis onboard, while a larger cloud-based model is used only when the local prediction is uncertain and the latency and resource conditions allow it.
Objective
The objective is to design and evaluate an uncertainty-aware, adaptive inference cascade across sensing or embedded devices, a vehicle high-performance unit (HPU), and cloud resources. The focus of the work will be on efficient inference, confidence estimation, and orchestration rather than on development of foundation models. The problem definition of the thesis is: How can calibrated uncertainty and vehicle operating conditions be used to decide, per input, whether a compact onboard model should answer locally or defer to a larger model, while meeting task-quality and latency requirements and limiting communication and resource use?
Together with the supervisors, the student will select one representative commercial-vehicle use case, such as predictive maintenance or anomaly detection, visual diagnostic assistance, time-series reasoning, or multimodal situational awareness. To keep the 20-week scope feasible, the thesis should prioritize one use case, one main data modality (or a clearly defined multimodal pair), and one primary algorithmic contribution: calibrated uncertainty and adaptive cloud escalation. The primary scope is the vehicle HPU-to-cloud cascade; sensor-level processing may be included only when it directly supports the chosen use case. The exact data, models, and target hardware will be defined based on availability and confidentiality requirements.
Illustrative three-tier deployment concept
(The focus will most likely be on Tier 2 and 3; Tier 1 is mentioned for context)
Execution tier
Candidate approach
Primary purpose
Tier 1: Sensor / Embedded layer
Lightweight preprocessing or tiny agent; signal-quality and uncertainty indicators
Reduce data, detect degraded sensing, and prepare features for onboard inference
Tier 2: Vehicle HPU layer
Small specialized or multimodal foundation model
Low-latency local inference with a or uncertainty score
Tier 3: Cloud layer
Large specialized or multimodal foundation model
Enhanced reasoning for difficult cases when latency, connectivity, and policy permit
Job description
- Conduct a literature study on compact foundation models, hierarchical inference, uncertainty estimation and calibration, learning to defer, and adaptive resource aware edge-cloud offloading.
- Select and deploy a compact open-source model on a representative edge or HPU platform. Characterize model quality, latency, throughput, memory footprint, model size, and, where feasible, energy consumption.
- Implement and evaluate a confidence or uncertainty-estimation method for the local model.
- Develop an adaptive threshold-learning or cloud-escalation strategy that jointly considers uncertainty, task quality, and other relevant factors such as latency and operational cost.
- Compare local-only, remote-only, static-threshold, adaptive-routing, and retrospective-oracle baselines. Report repeated trials and uncertainty for task quality, calibration, offload rate, latency, communication, memory, and energy where feasible.
- Deliver reproducible code, experiment configurations, documented datasets and models, analysis scripts, and a clear account of limitations and deployment assumptions.
An optional extension is a theoretical performance model, regret analysis, or a clearly scoped multimodal experiment.
Expected outcome
The expected result is a reproducible prototype and evaluation pipeline that combines an edge-deployable foundation model, calibrated uncertainty estimates, and an adaptive escalation policy. Subject to result quality and confidentiality constraints, the work may also contribute to a scientific publication.
References
[1] V. N. Moothedath, J. P. Champati, and J. Gross, "Getting the Best Out of Both Worlds: Algorithms for Hierarchical Inference at the Edge," IEEE Transactions on Machine Learning in Communications and Networking, 2024.
[2] C.-H. Chang, A. P. Behera, S. Zhan