Motivated by the potential of machine-learning- based (ML) algorithms for radio access network (RAN) control and management, we consider the problem of energy-aware O-RAN service orchestration subject to ML inference time constraints. While ML applications enable complex operations in RAN control, guaranteeing service level agreements to close RAN operations in real time is a key requirement to facilitating their wider adoption. In this paper, we focus on orchestrating ML/AI workloads as near-real-time applications in O-RAN Cloud (O-Cloud). We propose PERX, an energy-efficient and performance-aware O-RAN orchestrator that predicts the performance of diverse sets of colocated ML/AL applications by learning a pairwise characterization of application inference times via hierarchical Bayesian learning. We formulate a latency-constrained integer optimization problem for application orchestration and propose an iterative procedure to solve the problem. In line with industry standards, we adopt Kubernetes as the orchestration framework to develop a latency-aware O-Cloud orchestrator. Experimental results reveal up to 50 % increase in profit with guaranteed service level agreements, compared to state of the art benchmarks.
Part of ISBN 9798331543709
QC 20251106