What Oracle added to OKE

Oracle Cloud Infrastructure Kubernetes Engine (OKE) now supports three optional cluster software add-ons designed to improve the experience of running GPU-accelerated workloads: the NVIDIA GPU Operator, the NVIDIA Network Operator, and Node Feature Discovery. The add-ons are available on enhanced clusters and give administrators a way to deploy and configure supported GPU, networking, and node-discovery components through OKE, with the goal of simplifying initial setup and ongoing maintenance for AI, machine learning, and high-performance computing workloads on Kubernetes.
What the new add-ons manage
The three new add-ons are delivered through OKE's cluster add-ons framework and are optional, meaning they are disabled by default. On enhanced clusters, administrators can enable or disable an add-on, select a supported version, choose automatic updates or pin a fixed version, and apply supported key/value configuration arguments. Oracle manages the add-on lifecycle, reducing the amount of operational software that teams must deploy and maintain manually. When automatic updates are selected, OKE deploys add-on updates that are compatible with the cluster's supported Kubernetes version.
NVIDIA GPU Operator's role
The NVIDIA GPU Operator add-on provides a broader, operator-based approach to managing NVIDIA software components used by GPU workloads. It manages components such as the NVIDIA device plugin, the NVIDIA Container Toolkit, the NVIDIA Multi-Instance GPU (MIG) Manager, the NVIDIA Data Center GPU Manager (DCGM), and DCGM Exporter, with supported configuration arguments allowing administrators to tailor the deployment for their environment. DCGM and DCGM Exporter provide GPU telemetry and monitoring capabilities for Kubernetes environments, and for supported GPUs administrators can also enable the MIG Manager. NVIDIA MIG can divide a supported GPU into multiple separate instances with dedicated compute and memory resources.
NVIDIA Network Operator and node discovery
OCI positions the NVIDIA Network Operator as the add-on for teams that need networking components beyond basic GPU exposure, including networking drivers, secondary networking, and remote direct memory access components that must align with the cluster's Kubernetes version, worker-node image, GPU shape, and network configuration. Node Feature Discovery handles the node-labeling side of that alignment, advertising hardware capabilities to Kubernetes so workloads can be scheduled to compatible nodes. Together, the three add-ons cover the GPU, networking, and node-discovery layers that OCI describes as previously requiring manual installation and configuration.
What remains outside the add-ons
The OCI post notes that provisioning GPU worker nodes is only one part of preparing a Kubernetes environment for accelerated workloads, and that administrators still need to align NVIDIA GPU drivers, container runtime and toolkit components, Kubernetes device plugins, node labels, monitoring components, and, depending on architecture, networking drivers, secondary networking, and RDMA components. The new add-ons are aimed at reducing the amount of GPU-related software that teams install, configure, and maintain manually while retaining configuration flexibility for specialized environments, but the post does not list specific supported GPU shapes, Kubernetes versions, or regional availability for the three add-ons.
Share this article







