GPU batch jobs
Aim: Provide the basics of how to access and use the Stoomboot cluster GPU batch system, i.e. how to submit batch jobs to one of the GPU queues.
Target audience: Users of the Stoomboot cluster's GPUs.
Introduction
The Stoomboot cluster has a number of GPU nodes that are suitable for running certain types of algorithms. Both interactive GPU nodes and batch nodes are available.
Currently, the GPUs can be used by only one user at a time. This means that the interactive nodes need more discipline from the users to share this interactive resource than the generic interactive nodes. (Contact stbc-admin@nikhef.nl for coordination.)
The GPU batch nodes can be used through the GPU queues, but only use this queue for actual GPU jobs!
Types of GPU nodes
There are two main GPU manufacturers, and software will typically only work on one brand or the other.
The NVIDIA branded GPUs use software called CUDA.
The AMD branded GPUs use software called ROCm.
Prerequisites
- A Nikhef account;
- An ssh client.
Usage
Submitting GPU batch system jobs
A number of nodes are equipped with GPUs:
| Node name | Number of nodes | Node manufacturer | Node type name | GPU manufacturer | GPU type | GPU number |
|---|---|---|---|---|---|---|
| wn-pijl-{002,003} | 2 | Asus | ESC4000A-E12 | NVIDIA | L40S | 2 |
| wn-pijl-{004..007} | 4 | Asus | ESC4000A-E12 | NVIDIA | L40S | 4 |
In order to direct your jobs to a particular type of node, you can set the job requirements to match the node property.
For more information see: https://batchdocs.web.cern.ch/gpu/index.html
Using the GPUs
Condor will make sure the relevant libaries are avaible inside your container to use the GPUs. So just a tensorflow or pytorch install with CUDA/ROCm support should work fine. If you use one of the LCGviews or other custom enviroment it could be that the LD_LIBRARY_PATH will get overwritten, in this case the libaries that are mounted inside the container are not longer beeing detected. If this is the case you can add /.singularity.d/libs back to the LD_LIBARY_PATH.
Viewing the status of the GPU batch nodes
To view the status of the nodes equipped with GPUs, a similar approach can be used; note the usage of single and double quotes and removal of TARGET. with respect to what you added to the requirements:
$> condor_status -gpu
Name ST User GPUs GPU-Memory GPU-Name
slot1@wn-cuyp-002.nikhef.nl Ui _ 1 7.9 GB NVIDIA GeForce GTX 1080
slot1@wn-cuyp-003.nikhef.nl Ui _ 1 7.9 GB NVIDIA GeForce GTX 1080
slot1@wn-lot-008.nikhef.nl Ui _ 2 31.7 GB Tesla V100-PCIE-32GB
slot1@wn-lot-009.nikhef.nl Ui _ 2 31.7 GB Tesla V100-PCIE-32GB
slot1@wn-pijl-002.nikhef.nl Ui _ 2 44.4 GB NVIDIA L40S
slot1@wn-pijl-003.nikhef.nl Ui _ 2 44.4 GB NVIDIA L40S
slot1@wn-pijl-004.nikhef.nl Ui _ 4 44.4 GB NVIDIA L40S
slot1@wn-pijl-005.nikhef.nl Ui _ 4 44.4 GB NVIDIA L40S
slot1@wn-pijl-006.nikhef.nl Ui _ 4 44.4 GB NVIDIA L40S
slot1@wn-pijl-007.nikhef.nl Ui _ 4 44.4 GB NVIDIA L40S
Total Owner Unclaimed Claimed Preempting Matched Drain Backfill BkIdle
Busy 0 0 0 0 0 0 0 0 0
Idle 0 0 10 0 0 0 0 0 0
Retiring 0 0 0 0 0 0 0 0 0
Total 0 0 10 0 0 0 0 0 0
Storing output for GPU batch jobs
Storing output data from a GPU job should follow the same conventions as the CPU batch jobs. See "Storage output data" on the Batch jobs page.
Links
Contact
- Email stbc-users@nikhef.nl for questions about GPUs.
- Chat in Nikhef's Mattermost channel for stbc-users.