docs: add DRA GPUCluster/NVIDIADriver install guide for OpenShift - #518
nikitabugrovsky wants to merge 1 commit into
Conversation
Add openshift/dra-gpu-ocp.rst documenting how to install the NVIDIA GPU Operator with Dynamic Resource Allocation (DRA) on Red Hat OpenShift Container Platform, using the GPUCluster and NVIDIADriver custom resources introduced in GPU Operator v26.7 as an alternative to the Device Plugin-based ClusterPolicy workflow. The procedure, prerequisites, CRD examples, and command output were validated end-to-end on a live OpenShift 4.22 cluster with NFD and the GPU Operator v26.7.1 certified operator, including a sample DRA workload requesting a full GPU through a ResourceClaimTemplate. Cross-reference the new page from install-gpu-ocp.rst and add it to the openshift/index.rst toctree, right after the ClusterPolicy-based installation guide. Existing ClusterPolicy content is unchanged. Signed-off-by: Nikita Bugrovsky <nbugrovs@redhat.com>
840ad63 to
cf82e60
Compare
mikemckiernan
left a comment
There was a problem hiding this comment.
Thanks for the update! I made some nits and suggestions. plmk if I can clarify anything.
| .. note:: | ||
|
|
||
| This procedure requires GPU Operator v26.7 or later and Kubernetes v1.34.2 or later, which corresponds to Red Hat OpenShift Container Platform 4.21 or later. | ||
| This guide is validated on Red Hat OpenShift Container Platform 4.22. |
There was a problem hiding this comment.
sugg: Would these be better as prereqs in the next section?
imo, no need to repeat the GPU Op version, it's mentioned on L15 and the docs are versioned. To me, also no need to mention the k8s version (but not climbing that hill so I can die on it)--my suspicion is that the OCP version is the key requirement.
|
|
||
| .. note:: Refer to :ref:`install-gpu-ocp` for detailed namespace and ``OperatorGroup`` creation steps if you have not already created them. | ||
|
|
||
| #. Select the ``v26.7`` channel explicitly and look up the starting CSV. DRA support requires v26.7 or later, and the certified-operator default channel might still point to an earlier release: |
There was a problem hiding this comment.
nit: My reading is that the admin sets the channel and that there's no need to mention the certified-operator default channel. If I misunderstood, please disregard. Also, the 26.7 in the text is going to be stale two milliseconds after merging. My sugg is to add source_subtitutions like GPU Op does.
| #. Select the ``v26.7`` channel explicitly and look up the starting CSV. DRA support requires v26.7 or later, and the certified-operator default channel might still point to an earlier release: | |
| #. Select the ``v${version}`` channel explicitly and look up the starting CSV: |
|
|
||
| .. code-block:: console | ||
|
|
||
| $ CHANNEL=v26.7 |
There was a problem hiding this comment.
same nit: This would become toil and easy to forget.
| $ CHANNEL=v26.7 | |
| $ CHANNEL=v${version} |
|
|
||
| $ oc get csv -n nvidia-gpu-operator $STARTING_CSV -o jsonpath='{.metadata.annotations.alm-examples}' | jq -r 'map(select(.kind == "GPUCluster")) | .[0]' > gpucluster.json | ||
|
|
||
| The default example is similar to the following: |
There was a problem hiding this comment.
sugg: Up to you, not a biggie either way.
| The default example is similar to the following: | |
| The default example, stored in ``gpucluster.json``, is similar to the following: |
|
|
||
| $ oc get csv -n nvidia-gpu-operator $STARTING_CSV -o jsonpath='{.metadata.annotations.alm-examples}' | jq -r 'map(select(.kind == "NVIDIADriver")) | .[0]' > nvidiadriver.json | ||
|
|
||
| The default example is similar to the following (abbreviated): |
There was a problem hiding this comment.
sugg: same suggestion
| The default example is similar to the following (abbreviated): | |
| The default example, stored in ``nvidiadriver.json``, is similar to the following (abbreviated): |
| nvidia-dra-validator-thqnh 1/1 Running 0 8m45s | ||
| nvidia-gpu-driver-rhel9-59f4c8db85-pn5c5 2/2 Running 0 8m45s | ||
|
|
||
| The driver ``DaemonSet`` is named ``nvidia-gpu-driver-rhel9-<hash>`` and its node selector includes the RHCOS ``OSTREE_VERSION`` label, similar to the ``nvidia-driver-daemonset-<RHCOS-version>`` naming used by ``ClusterPolicy``. |
There was a problem hiding this comment.
sugg: treat the daemon set as a general reference
| The driver ``DaemonSet`` is named ``nvidia-gpu-driver-rhel9-<hash>`` and its node selector includes the RHCOS ``OSTREE_VERSION`` label, similar to the ``nvidia-driver-daemonset-<RHCOS-version>`` naming used by ``ClusterPolicy``. | |
| The driver daemon set is named ``nvidia-gpu-driver-rhel9-<hash>`` and its node selector includes the RHCOS ``OSTREE_VERSION`` label, similar to the ``nvidia-driver-daemonset-<RHCOS-version>`` naming used by ``ClusterPolicy``. |
|
|
||
| $ oc get resourceslice -o yaml | ||
|
|
||
| *Example Output* (abbreviated) |
There was a problem hiding this comment.
| *Example Output* (abbreviated) | |
| *Partial Output* |
|
|
||
| .. code-block:: console | ||
|
|
||
| $ oc new-project gpu-example |
There was a problem hiding this comment.
sugg: Not critical, but something like oc new-project gpu-dra-demo that is tailored to the procedure is a nice touch, imo.
| resourceClaimTemplateName: single-gpu | ||
| EOF | ||
|
|
||
| .. note:: OpenShift can report a ``PodSecurity`` admission warning similar to the following. The warning does not block pod creation under the default namespace Pod Security level and can be ignored for this example: |
There was a problem hiding this comment.
sugg: this page has a number of admonitions. My experience is that they can be visually distraction after one or two on a page. If someone runs the command, receives the barfy output, I think that person will go back to the docs for guidance. Trying to grab attention with the admonition isn't necessary, imo.
| .. note:: OpenShift can report a ``PodSecurity`` admission warning similar to the following. The warning does not block pod creation under the default namespace Pod Security level and can be ignored for this example: | |
| OpenShift can report a ``PodSecurity`` admission warning similar to the following. The warning does not block pod creation under the default namespace Pod Security level and can be ignored for this example: |
| .. tip:: | ||
|
|
||
| Starting with GPU Operator v26.7, you can alternatively install the DRA Driver for NVIDIA GPUs by using the ``GPUCluster`` and ``NVIDIADriver`` custom resources instead of ``ClusterPolicy``. | ||
| A cluster can use one or the other, but not both. | ||
| Refer to :doc:`dra-gpu-ocp` for the DRA installation procedure on OpenShift. | ||
|
|
There was a problem hiding this comment.
I won't stop you, but this approach does not scale. You are welcome to make updates in the "Post-Release Documentation Updates" of the GPU Op release notes file.
Documentation preview |
Add openshift/dra-gpu-ocp.rst documenting how to install the NVIDIA GPU Operator with Dynamic Resource Allocation (DRA) on Red Hat OpenShift Container Platform, using the GPUCluster and NVIDIADriver custom resources introduced in GPU Operator v26.7 as an alternative to the Device Plugin-based ClusterPolicy workflow.
The procedure, prerequisites, CRD examples, and command output were validated end-to-end on a live OpenShift 4.22 cluster with NFD and the GPU Operator v26.7.1 certified operator, including a sample DRA workload requesting a full GPU through a ResourceClaimTemplate.
Cross-reference the new page from install-gpu-ocp.rst and add it to the openshift/index.rst toctree, right after the ClusterPolicy-based installation guide. Existing ClusterPolicy content is unchanged.