Skip to content

docs: add DRA GPUCluster/NVIDIADriver install guide for OpenShift - #518

Open
nikitabugrovsky wants to merge 1 commit into
NVIDIA:mainfrom
nikitabugrovsky:nbugrovs/dra-gpu-ocp
Open

nikitabugrovsky wants to merge 1 commit into
NVIDIA:mainfrom
nikitabugrovsky:nbugrovs/dra-gpu-ocp

Conversation

@nikitabugrovsky

Copy link
Copy Markdown
Contributor

Add openshift/dra-gpu-ocp.rst documenting how to install the NVIDIA GPU Operator with Dynamic Resource Allocation (DRA) on Red Hat OpenShift Container Platform, using the GPUCluster and NVIDIADriver custom resources introduced in GPU Operator v26.7 as an alternative to the Device Plugin-based ClusterPolicy workflow.

The procedure, prerequisites, CRD examples, and command output were validated end-to-end on a live OpenShift 4.22 cluster with NFD and the GPU Operator v26.7.1 certified operator, including a sample DRA workload requesting a full GPU through a ResourceClaimTemplate.

Cross-reference the new page from install-gpu-ocp.rst and add it to the openshift/index.rst toctree, right after the ClusterPolicy-based installation guide. Existing ClusterPolicy content is unchanged.

Add openshift/dra-gpu-ocp.rst documenting how to install the NVIDIA
GPU Operator with Dynamic Resource Allocation (DRA) on Red Hat
OpenShift Container Platform, using the GPUCluster and NVIDIADriver
custom resources introduced in GPU Operator v26.7 as an alternative
to the Device Plugin-based ClusterPolicy workflow.

The procedure, prerequisites, CRD examples, and command output were
validated end-to-end on a live OpenShift 4.22 cluster with NFD and
the GPU Operator v26.7.1 certified operator, including a sample DRA
workload requesting a full GPU through a ResourceClaimTemplate.

Cross-reference the new page from install-gpu-ocp.rst and add it to
the openshift/index.rst toctree, right after the ClusterPolicy-based
installation guide. Existing ClusterPolicy content is unchanged.

Signed-off-by: Nikita Bugrovsky <nbugrovs@redhat.com>
@nikitabugrovsky nikitabugrovsky changed the title doc: add DRA GPUCluster/NVIDIADriver install guide for OpenShift docs: add DRA GPUCluster/NVIDIADriver install guide for OpenShift Sep 30, 2026

@mikemckiernan mikemckiernan left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the update! I made some nits and suggestions. plmk if I can clarify anything.

Comment thread openshift/dra-gpu-ocp.rst
Comment on lines +22 to +25
.. note::

This procedure requires GPU Operator v26.7 or later and Kubernetes v1.34.2 or later, which corresponds to Red Hat OpenShift Container Platform 4.21 or later.
This guide is validated on Red Hat OpenShift Container Platform 4.22.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

sugg: Would these be better as prereqs in the next section?

imo, no need to repeat the GPU Op version, it's mentioned on L15 and the docs are versioned. To me, also no need to mention the k8s version (but not climbing that hill so I can die on it)--my suspicion is that the OCP version is the key requirement.

Comment thread openshift/dra-gpu-ocp.rst

.. note:: Refer to :ref:`install-gpu-ocp` for detailed namespace and ``OperatorGroup`` creation steps if you have not already created them.

#. Select the ``v26.7`` channel explicitly and look up the starting CSV. DRA support requires v26.7 or later, and the certified-operator default channel might still point to an earlier release:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: My reading is that the admin sets the channel and that there's no need to mention the certified-operator default channel. If I misunderstood, please disregard. Also, the 26.7 in the text is going to be stale two milliseconds after merging. My sugg is to add source_subtitutions like GPU Op does.

Suggested change
#. Select the ``v26.7`` channel explicitly and look up the starting CSV. DRA support requires v26.7 or later, and the certified-operator default channel might still point to an earlier release:
#. Select the ``v${version}`` channel explicitly and look up the starting CSV:

Comment thread openshift/dra-gpu-ocp.rst

.. code-block:: console

$ CHANNEL=v26.7

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

same nit: This would become toil and easy to forget.

Suggested change
$ CHANNEL=v26.7
$ CHANNEL=v${version}

Comment thread openshift/dra-gpu-ocp.rst

$ oc get csv -n nvidia-gpu-operator $STARTING_CSV -o jsonpath='{.metadata.annotations.alm-examples}' | jq -r 'map(select(.kind == "GPUCluster")) | .[0]' > gpucluster.json

The default example is similar to the following:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

sugg: Up to you, not a biggie either way.

Suggested change
The default example is similar to the following:
The default example, stored in ``gpucluster.json``, is similar to the following:

Comment thread openshift/dra-gpu-ocp.rst

$ oc get csv -n nvidia-gpu-operator $STARTING_CSV -o jsonpath='{.metadata.annotations.alm-examples}' | jq -r 'map(select(.kind == "NVIDIADriver")) | .[0]' > nvidiadriver.json

The default example is similar to the following (abbreviated):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

sugg: same suggestion

Suggested change
The default example is similar to the following (abbreviated):
The default example, stored in ``nvidiadriver.json``, is similar to the following (abbreviated):

Comment thread openshift/dra-gpu-ocp.rst
nvidia-dra-validator-thqnh 1/1 Running 0 8m45s
nvidia-gpu-driver-rhel9-59f4c8db85-pn5c5 2/2 Running 0 8m45s

The driver ``DaemonSet`` is named ``nvidia-gpu-driver-rhel9-<hash>`` and its node selector includes the RHCOS ``OSTREE_VERSION`` label, similar to the ``nvidia-driver-daemonset-<RHCOS-version>`` naming used by ``ClusterPolicy``.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

sugg: treat the daemon set as a general reference

Suggested change
The driver ``DaemonSet`` is named ``nvidia-gpu-driver-rhel9-<hash>`` and its node selector includes the RHCOS ``OSTREE_VERSION`` label, similar to the ``nvidia-driver-daemonset-<RHCOS-version>`` naming used by ``ClusterPolicy``.
The driver daemon set is named ``nvidia-gpu-driver-rhel9-<hash>`` and its node selector includes the RHCOS ``OSTREE_VERSION`` label, similar to the ``nvidia-driver-daemonset-<RHCOS-version>`` naming used by ``ClusterPolicy``.

Comment thread openshift/dra-gpu-ocp.rst

$ oc get resourceslice -o yaml

*Example Output* (abbreviated)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
*Example Output* (abbreviated)
*Partial Output*

Comment thread openshift/dra-gpu-ocp.rst

.. code-block:: console

$ oc new-project gpu-example

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

sugg: Not critical, but something like oc new-project gpu-dra-demo that is tailored to the procedure is a nice touch, imo.

Comment thread openshift/dra-gpu-ocp.rst
resourceClaimTemplateName: single-gpu
EOF

.. note:: OpenShift can report a ``PodSecurity`` admission warning similar to the following. The warning does not block pod creation under the default namespace Pod Security level and can be ignored for this example:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

sugg: this page has a number of admonitions. My experience is that they can be visually distraction after one or two on a page. If someone runs the command, receives the barfy output, I think that person will go back to the docs for guidance. Trying to grab attention with the admonition isn't necessary, imo.

Suggested change
.. note:: OpenShift can report a ``PodSecurity`` admission warning similar to the following. The warning does not block pod creation under the default namespace Pod Security level and can be ignored for this example:
OpenShift can report a ``PodSecurity`` admission warning similar to the following. The warning does not block pod creation under the default namespace Pod Security level and can be ignored for this example:

Comment on lines +10 to +15
.. tip::

Starting with GPU Operator v26.7, you can alternatively install the DRA Driver for NVIDIA GPUs by using the ``GPUCluster`` and ``NVIDIADriver`` custom resources instead of ``ClusterPolicy``.
A cluster can use one or the other, but not both.
Refer to :doc:`dra-gpu-ocp` for the DRA installation procedure on OpenShift.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I won't stop you, but this approach does not scale. You are welcome to make updates in the "Post-Release Documentation Updates" of the GPU Op release notes file.

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown

Documentation preview

https://nvidia.github.io/cloud-native-docs/review/pr-518

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants