Lux User Guide
System Overview
Lux is an AMD-based supercomputer built on HPE ProLiant XD685 nodes and located at the Oak Ridge Leadership Computing Facility (OLCF). It enables users to dramatically accelerate AI-driven research. Lux consists of 500+ nodes and 4,000+ GPUs, divided into Slurm (HPC) and Kubernetes partitions.
Lux Compute Nodes
Each Lux compute node contains [2x] 64-core AMD EPYC 9575F CPUs with access to 3 TB of DDR5 memory. Each node also contains [8x] AMD Instinct MI355X GPUs, each with 288 GB of high-bandwidth memory (HBM3E) and 8 Accelerated Compute Dies (XCDs), for a total of 64 XCDs per node. The programmer can think of each MI355X as an individual GPU, with 288 GB of high-bandwidth memory (HBM3E).
Each MI355X are connected with Infinity Fabric GPU-GPU with a peak bandwidth of 153.6 GB/s. The MI355X GPUs are connected with Infinity Fabric GPU-GPU in the arrangement shown in the Lux Node Diagram below, where the peak bandwidth is a constant 153.6 GB/s because each GPU has a single Infinity Fabric connection to other individual GPUs.
Note
TERMINOLOGY:
Each MI355X will show as a separate GPU according to Slurm, ROCR_VISIBLE_DEVICES, and the ROCr runtime, so from this point forward in the quick-start guide, we will simply refer to the MI355X as a GPU.
Note
There are [8x] NUMA domains per node. The 8 GPUs are each associated with one NUMA domain as follows:
NUMA 0:
hardware threads 000-015 | GPU 3
NUMA 1:
hardware threads 016-031 | GPU 0
NUMA 2:
hardware threads 032-047 | GPU 1
NUMA 3:
hardware threads 048-063 | GPU 2
NUMA 4:
hardware threads 064-079 | GPU 7
NUMA 5:
hardware threads 080-095 | GPU 4
NUMA 6:
hardware threads 096-111 | GPU 5
NUMA 7:
hardware threads 112-127 | GPU 6
Node Types
On Lux, there are two major types of nodes you will encounter: Login and Compute. While these are similar in terms of hardware (see: Lux Compute Nodes), they differ considerably in their intended use.
Node Type |
Description |
|---|---|
Login |
When you connect to Lux, you’re placed on a login node. This is the place to write/edit/compile your code, manage data, submit jobs, etc. You should never launch parallel jobs from a login node nor should you run threaded jobs on a login node. Login nodes are shared resources that are in use by many users simultaneously. |
Compute |
Most of the nodes on Lux are compute nodes. These are where
your parallel job executes. They’re accessed via the |
System Interconnect
File Systems
Lux is connected to Orion, a parallel filesystem based on Lustre and HPE ClusterStor, with a 679 PB usable
namespace (/lustre/orion/). In addition to Lux, Orion is available on Frontier, the OLCF’s data transfer nodes, and on the Andes cluster.
Lux also has access to the center-wide NFS-based filesystem (which provides user and project home areas).
Each compute node has eight 3.2TB Non-Volatile Memory storage devices. See Data and Storage for more information.
Project’s with a Lux allocation also receive an archival storage area on Kronos. For more information on using Kronos, see the Kronos Nearline Archival Storage System section.
Operating System
Lux is running Red Hat Enterprise Linux (RHEL) version 9.8.
GPUs
Each Lux compute node contains eight (8) AMD MI355X. The AMD MI355X has a peak performance of 78.6 TFLOPS in vector-based double-precision for modeling and simulation and 157.3 TFLOPS in vector-based half-precision AI workloads. Each MI355X contains 256 compute units across its 8 XCDs and 288 GB of high-bandwidth memory (HBM3E) which can be accessed at a peak of 8 TB/s. The 8 GPUs on a Lux node are connected with Infinity Fabric with a bandwidth of 153.6 GB/s (in each direction simultaneously).
Connecting
To connect to Lux, ssh to login1.lux.olcf.ornl.gov from either home.ccs.ornl.gov or hub.ccs.ornl.gov. For example:
$ ssh <username>@home.ccs.ornl.gov
$ ssh login1.lux.olcf.ornl.gov
For more information on connecting to OLCF resources, see Connecting for the first time.
Data and Storage
Orion
Lux mounts Orion, a parallel filesystem based on Lustre and HPE ClusterStor, with a 679 PB usable namespace (/lustre/orion/). In addition to Lux, Orion is available on the OLCF’s data transfer nodes.
Orion uses a feature called Progressive File Layout (PFL) that changes the striping of files as they grow. Because of this, we ask users not to manually adjust the file striping. If you feel the default striping behavior of Orion is not meeting your needs, please contact help@olcf.ornl.gov.
Files older than 90 days are purged from Orion. Please plan your data management and lifecycle at OLCF before generating the data.
For more detailed information about center-wide file systems and data archiving available on Lux, please refer to the pages on Data Storage and Transfers. The subsections below give a quick overview of NFS, Lustre, and archival storage spaces as well as the on node NVMe “Burst Buffers” (SSDs).
NFS Filesystem
Area |
Path |
Type |
Permissions |
Quota |
Backups |
Purged |
Retention |
On Compute Nodes |
|---|---|---|---|---|---|---|---|---|
User Home |
|
NFS |
User set |
50 GB |
Yes |
No |
90 days |
Yes |
Project Home |
|
NFS |
770 |
50 GB |
Yes |
No |
90 days |
Yes |
Note
Though the NFS filesystem’s User Home and Project Home areas are read/write from Lux’s compute nodes, we strongly recommend that users launch and run jobs from the Lustre Orion parallel filesystem instead due to its larger storage capacity and superior performance. Please see below for Lustre Orion filesystem storage areas and paths.
Lustre Filesystem
Area |
Path |
Type |
Permissions |
Quota |
Backups |
Purged |
Retention |
On Compute Nodes |
|---|---|---|---|---|---|---|---|---|
Member Work |
|
Lustre HPE ClusterStor |
700 |
50 TB |
No |
90 days |
N/A |
Yes |
Project Work |
|
Lustre HPE ClusterStor |
770 |
50 TB |
No |
90 days |
N/A |
Yes |
World Work |
|
Lustre HPE ClusterStor |
775 |
50 TB |
No |
90 days |
N/A |
Yes |
Warning
Proprietary/Sensitive/Controlled Information Notice
Portions of data and/or software used in your project may require extra protections due to requirements for proprietary, sensitive, or controlled information. It is imperative that filenames, application names, job names, environment variables, batch job scripts, or any other unencrypted text must never contain sensitive or controlled information.
If you have HIPAA or ITAR data, you will need to use our SPI resources. More information about SPI can be found here.
If you have security related questions, contact us via email at: security-admins@ccs.ornl.gov. Other questions can be sent to help@olcf.ornl.gov.
Kronos Archival Storage
Please note that the Kronos is not mounted directly onto Lux nodes. There are two main methods for accessing and moving data to/from Kronos, either with standard cli utilities (scp, rsync, etc.) and via Globus using the “OLCF Kronos” collection. For more information on using Kronos, see the Kronos Nearline Archival Storage System section.
Area |
Path |
Type |
Permissions |
Quota |
Backups |
Purged |
Retention |
On Compute Nodes |
|---|---|---|---|---|---|---|---|---|
Member Archive |
|
Nearline |
700 |
200 TB* |
No |
No |
90 days |
No |
Project Archive |
|
Nearline |
770 |
200 TB* |
No |
No |
90 days |
No |
World Archive |
|
Nearline |
775 |
200 TB* |
No |
No |
90 days |
No |
Note
The three archival storage areas above share a single 200TB per project quota.
NVMe
Each compute node on Lux has [8x] Kioxia CM7 3.2TB Non-Volatile Memory (NVMe) storage devices (SSDs), colloquially known as a “Burst Buffer”. Each SSD has a peak sequential performance of 10,000 MB/s (read) and 4,900 MB/s (write).
The purpose of the Burst Buffer system is to bring improved I/O performance to appropriate workloads. Users are not required to use the NVMes. Data can also be written directly to the parallel filesystem.
NVMe Usage
NVMe devices are automatically allocated when a user requests a GPU, with 12% of the total node NVMe capacity per GPU requested.
(.12 * 3.2TB/NVMe * 8 NVMe/node = 2.8TiB/GPU requested)
The remaining 4% of NVMe capacity is reserved for Kubernetes uses.
Once the NVMe pool slice is allocated to a job, users can access the slice at /mnt/bb/$SLURM_JOBID
Users are responsible for moving data to/from the NVMe before/after their jobs
Example
#!/bin/bash
# hello_nvme.sbatch
#SBATCH --account stf007
#SBATCH --job-name nvme_test
#SBATCH --output %x-%j.out
#SBATCH --time 00:05:00
#SBATCH --partition batch
#SBATCH --nodes 1
#SBATCH --gpus 2
date
echo " "
echo "*****ORIGINAL FILE*****"
cat test.txt
echo "***********************"
# Move file from working directory to SSD
mv test.txt /mnt/bb/$SLURM_JOBID
# Edit file from compute node
srun -n1 hostname >> /mnt/bb/$SLURM_JOBID/test.txt
# Find size of allocation and add to file
srun -n1 df -h /mnt/bb/$SLURM_JOBID >> /mnt/bb/$SLURM_JOBID/test.txt
# Move file from SSD back to working directory
mv /mnt/bb/$SLURM_JOBID/test.txt .
echo " "
echo "*****UPDATED FILE******"
cat test.txt
echo "***********************"
Below is the output:
$ cat nvme_test-<JOB ID>.out
*****ORIGINAL FILE*****
Hello, world from Lux!
***********************
*****UPDATED FILE******
Hello, world from Lux!
lux012
Filesystem Size Used Avail Use% Mounted on
/dev/mapper/nvme-bb--<JOB ID> 5.6T 40G 5.6T 1% /mnt/bb/<JOB ID>
***********************
Using Globus to Move Data to and from Orion
The following example is intended to help users move data to and from the Orion filesystem.
Below is a summary of the steps for data transfer using Globus:
1. Login to globus.org using your globus ID and password. If you do not have a globusID, set one up here: Generate a globusID.
Once you are logged in, Globus will open the “File Manager” page. Click the left side “Collection” text field in the File Manager and type “OLCF DTN (Globus 5)”.
When prompted, authenticate into the OLCF DTN (Globus 5) collection using your OLCF username and PIN followed by your RSA passcode.
Click in the left side “Path” box in the File Manager and enter the path to your data on Orion. For example,`/lustre/orion/stf007/proj-shared/my_orion_data`. You should see a list of your files and folders under the left “Path” Box.
Click on all files or folders that you want to transfer in the list. This will highlight them.
Click on the right side “Collection” box in the File Manager and type the name of a second collection at OLCF or at another institution. You can transfer data between different paths on the Orion filesystem with this method too; Just use the OLCF DTN (Globus 5) collection again in the right side “Collection” box.
Click in the right side “Path” box and enter the path where you want to put your data on the second collection’s filesystem.
Click the left “Start” button.
Click on “Activity“ in the left blue menu bar to monitor your transfer. Globus will send you an email when the transfer is complete.
Globus Warnings:
Globus transfers do not preserve file permissions. Arriving files will have (rw-r–r–) permissions, meaning arriving files will have user read and write permissions and group and world read permissions. Note that the arriving files will not have any execute permissions, so you will need to use chmod to reset execute permissions before running a Globus-transferred executable.
Globus will overwrite files at the destination with identically named source files. This is done without warning.
Globus has restriction of 8 active transfers across all the users. Each user has a limit of 3 active transfers, so it is required to transfer a lot of data on each transfer than less data across many transfers.
If a folder is constituted with mixed files including thousands of small files (less than 1MB each one), it would be better to tar the small files. Otherwise, if the files are larger, Globus will handle them.
AMD GPUs
The AMD Instinct MI355X is built on advanced packaging technologies enabling eight Accelerated Compute Dies (XCDs) to be integrated into a single package in the Open Compute Project (OCP) Accelerator Module (OAM) in the MI355X product. Each XCD is build on the AMD CDNA 4 architecture. A single Lux node contains 8 MI355X OAMs for a total of 64 XCDs.
Note
The Slurm workload manager and the ROCr runtime treat each MI355X as a separate GPU
and visibility can be controlled using the ROCR_VISIBLE_DEVICES environment variable.
Therefore, from this point on, the Lux guide simply refers to a MI355X as a GPU.
Each XCD contains 32 Compute Units (CUs) grouped in 4 Asynchronous Compute Engines (ACEs). Physically, each XCD contains 36 CUs, but four are disabled. A command processor in each GPU receives API commands and transforms them into compute tasks.
NEEDS REVIEW: Compute tasks are managed by the 4 asynchronous compute engines, which dispatch wavefronts to compute units. All wavefronts from a single workgroup are assigned to the same CU. In CUDA terminology, workgroups are “blocks”, wavefronts are “warps”, and work-items are “threads”. The terms are often used interchangeably.
The 256 CUs in each GPU deliver peak performance of 78.6 TFLOPS in double precision on both Vector and specialized Matrix cores. Also, each GPU contains 288 GB of high-bandwidth memory (HBM3E) accessible at a peak bandwidth of 8.0 TB/s. The 8 GPUs in an Lux node are connected with [1x] GPU-to-GPU Infinity Fabric links providing 76.8+76.8 GB/s of bandwidth. Consult the diagram in the Lux Compute Nodes section for information on how the accelerators are connected to each other, to the CPU, and to the network.
Note
The X+X GB/s notation describes bidirectional bandwidth, meaning X GB/s in each direction.
The Compute Unit
Each CU has 4 Matrix Core Units (the equivalent of NVIDIA’s Tensor core units) and 4 16-wide SIMD units. For a vector instruction that uses the SIMD units, each wavefront (which has 64 threads) is assigned to a single 16-wide SIMD unit such that the wavefront as a whole executes the instruction over 4 cycles, 16 threads per cycle. Since other wavefronts occupy the other three SIMD units at the same time, the total throughput still remains 1 instruction per cycle. Each CU maintains an instructions buffer for 8 wavefronts and also maintains 256 registers where each register is 64 4-byte wide entries.
HIP
The Heterogeneous Interface for Portability (HIP) is AMD’s dedicated GPU programming environment for designing high performance kernels on GPU hardware. HIP is a C++ runtime API and programming language that allows developers to create portable applications on different platforms, including the AMD MI355X. This means that developers can write their GPU applications and with very minimal changes be able to run their code in any environment. The API is very similar to CUDA, so if you’re already familiar with CUDA there is almost no additional work to learn HIP. See here for a series of tutorials on programming with HIP and also converting existing CUDA code to HIP with the hipify tools .
See the Compiling section for information on compiling for AMD GPUs, and see the Tips and Tricks section for some detailed information to keep in mind to run more efficiently on AMD GPUs.
Programming Environment
Lux users are provided with many pre-installed software packages and scientific libraries. To facilitate this, environment management tools are used to handle necessary changes to the shell.
Environment Modules (Lmod)
Environment modules are provided through Lmod, a Lua-based module system for dynamically altering shell environments.
By managing changes to the shell’s environment variables (such as PATH, LD_LIBRARY_PATH, and PKG_CONFIG_PATH), Lmod allows you to alter the software available in your shell environment without the risk of creating package and version combinations that cannot coexist in a single environment.
General Usage
The interface to Lmod is provided by the module command:
Command |
Description |
|---|---|
|
Shows a terse list of the currently loaded modules |
|
Shows a table of the currently available modules |
|
Shows help information about |
|
Shows the environment changes made by the |
|
Searches all possible modules according to |
|
Loads the given |
|
Adds |
|
Removes |
|
Unloads all modules |
|
Resets loaded modules to system defaults |
|
Reloads all currently loaded modules |
Searching for Modules
Modules with dependencies are only available when the underlying dependencies, such as compiler families, are loaded. Thus, module avail will only display modules that are compatible with the current state of the environment. To search the entire hierarchy across all possible dependencies, the spider sub-command can be used as summarized in the following table.
Command |
Description |
|---|---|
|
Shows the entire possible graph of modules |
|
Searches for modules named |
|
Searches for a specific version of |
|
Searches for modulefiles containing |
Compilers
AMD and GCC compilers are provided through modules on Lux.
The AMD compilers are both based on LLVM/Clang.
There is also a system/OS versions of GCC available in /usr/bin.
The table below lists details about each of the module-provided compilers.
Please see the following Compiling section for more detailed information on how to compile using these modules.
MPI
The MPI implementations available on Lux are OpenMPI and MPICH, which are “GPU-aware” so GPU buffers can be passed directly to MPI calls.
Lux is primarily a RCCL-centric machine utilizing AMD Pensando Pollara 400GbE AI NICs on each compute node. MPI is not yet officially verified.
Due to tested scaling limitations, OLCF strongly recommends your MPI workloads are limited to 16 nodes.
RCCL
The ROCm Collective Communication Library (RCCL) is installed with ROCm and can be accessed by loading a rocm module.
Lux is primarily a RCCL-centric machine utilizing AMD Pensando Pollara 400GbE AI NICs on each compute node.
Compiling
Compilers
AMD and GCC compilers are provided through modules on Lux.
The AMD compilers are based on LLVM/Clang.
There is also a system/OS versions of GCC available in /usr/bin.
The table below lists details about each of the module-provided compilers.
Vendor |
Compiler Module |
Language |
Compiler |
|---|---|---|---|
AMD |
|
C |
|
AMD |
|
C++ |
|
AMD |
|
Fortran |
|
GNU |
|
C |
|
GNU |
|
C++ |
|
GNU |
|
Fortran |
|
Programming Environment
Programming environments consist of a compiler and some basic dependencies, such as MPI and ROCm.
You can find existing programming environments on Lux utilizing the Lmod command module avail.
You will see an output like the following:
------------------------------------------------------------------------------- [ amd-llvm/7.2.4, mpich/5.0.1 ] --------------------------------------------------------------------------------
amdfftw/5.3 amdscalapack/5.3 hdf5/1.14.6 netcdf-c/4.10.0 netcdf-fortran/4.6.2
The delimiter row is your “programming environment” as a list of modules and the modules underneath are optional modules that are managed by the programming environment.
Below is an example of how this functions:
1$ module load amd-llvm mpich
2$ module load amdfftw
3$ module unload mpich
4
5Inactive Modules:
6 1) amdfftw
Explanation:
Loading the programming environment by loading the requisite modules
Loading the optional dependency for AMD Fast Fourier Transforms
Unload part of the programming environment.
If the criteria for a programming environment are no longer met, any listed modules will then be inactive. Modules will also be reloaded if their dependent module is swapped for a valid alternative provider such as swapping MPI implementations.
Exposing The ROCm Toolchain to your Programming Environment
If you need to add the tools and libraries related to ROCm, the framework for targeting AMD GPUs, to your path, you will need to use a version of ROCm that is compatible with your programming environment.
ROCm can be loaded with: module load rocm/X.Y.Z, or to load the default ROCm version, module load rocm.
MPI
The MPI implementations available on Lux are OpenMPI (default) and MPICH, which are “GPU-aware” so GPU buffers can be passed directly to MPI calls.
Implementation |
Module |
Compiler |
Header Files & Linking |
|---|---|---|---|
OpenMPI |
|
|
-I${MPI_DIR}/include-L${MPI_DIR}/lib -lmpi |
|
-I${MPI_DIR}/include-L${MPI_DIR}/lib -lmpi |
||
MPICH |
|
|
-I${MPICH_DIR}/include-L${MPICH_DIR}/lib -lmpi |
|
-I${MPICH_DIR}/include-L${MPICH_DIR}/lib -lmpi |
Note
hipcc requires the ROCm Toolclain, See Exposing The ROCm Toolchain to your Programming Environment
GPU-Aware MPI
To use GPU-aware MPI, users must load both a ROCm module and an MPI-providing module:
gpu-aware.cpp
#include <stdio.h>
#include <hip/hip_runtime.h>
#include <mpi.h>
int main(int argc, char **argv) {
int i,rank,size,bufsize;
int *h_buf;
int *d_buf;
MPI_Status status;
bufsize=100;
MPI_Init(&argc,&argv);
MPI_Comm_rank(MPI_COMM_WORLD, &rank);
MPI_Comm_size(MPI_COMM_WORLD, &size);
//allocate buffers
h_buf=(int*) malloc(sizeof(int)*bufsize);
hipMalloc(&d_buf, bufsize*sizeof(int));
//initialize buffers
if(rank==0) {
for(i=0;i<bufsize;i++)
h_buf[i]=i*i;
}
if(rank==1) {
for(i=0;i<bufsize;i++)
h_buf[i]=-1;
}
hipMemcpy(d_buf, h_buf, bufsize*sizeof(int), hipMemcpyHostToDevice);
//communication
if(rank==0)
MPI_Send(d_buf, bufsize, MPI_INT, 1, 123, MPI_COMM_WORLD);
if(rank==1)
MPI_Recv(d_buf, bufsize, MPI_INT, 0, 123, MPI_COMM_WORLD, &status);
//validate results
if(rank==1) {
hipMemcpy(h_buf, d_buf, bufsize*sizeof(int), hipMemcpyDeviceToHost);
for(i=0;i<bufsize;i++) {
if(h_buf[i] != i*i)
printf("Error: buffer[%d]=%d but expected %d\n", i, h_buf[i], i);
}
fflush(stdout);
}
//free buffers
free(h_buf);
hipFree(d_buf);
MPI_Finalize();
}
Using hipcc
module load rocm
module load openmpi
hipcc -std=c++11 --offload-arch=gfx950 -I${ROCM_PATH}/include -I${MPI_DIR}/include -c gpu-aware.cpp
hipcc -L${ROCM_PATH}/lib -lamdhip64 -L${MPI_DIR}/lib -lmpi gpu-aware.o -o gpu-aware
MPICH Example
module load rocm
module load mpich
hipcc -std=c++11 --offload-arch=gfx950 -I${ROCM_PATH}/include -I${MPICH_DIR}/include -c gpu-aware.cpp
hipcc -L${ROCM_PATH}/lib -lamdhip64 -L${MPICH_DIR}/lib -lmpi gpu-aware.o -o gpu-aware
Using amdclang
module load rocm
module load openmpi
amdclang++ -D__HIP_ROCclr__ -D__HIP_ARCH_GFX950__=1 -std=c++11 --rocm-path=${ROCM_PATH} --offload-arch=gfx950 -x hip -I${ROCM_PATH}/include -I${MPI_DIR}/include -c gpu-aware.cpp
amdclang++ --rocm-path=${ROCM_PATH} -L${ROCM_PATH}/lib -lamdhip64 -L${MPI_DIR}/lib -lmpi gpu-aware.o -o gpu-aware
MPICH Example
module load rocm
module load mpich
amdclang++ -D__HIP_ROCclr__ -D__HIP_ARCH_GFX950__=1 -std=c++11 --rocm-path=${ROCM_PATH} --offload-arch=gfx950 -x hip -I${ROCM_PATH}/include -I${MPICH_DIR}/include -c gpu-aware.cpp
amdclang++ --rocm-path=${ROCM_PATH} -L${ROCM_PATH}/lib -lamdhip64 -L${MPICH_DIR}/lib -lmpi gpu-aware.o -o gpu-aware
Note
The primary required steps for GPU-aware MPI apply to both the amdclang and hipcc compilers, and those are:
Specify the ROCm and MPI include path at compile time
-I${ROCM_PATH}/include -I${MPI_DIR}/includeSpecify the ROCm and MPI library path and libraries at link time
-L${ROCM_PATH}/lib -lamdhip64 -L${MPI_DIR}/lib
MPICH_DIR can be substituted for MPI_DIR in order to use MPICH with PMI2.
Understanding the Compatibility of Compilers, ROCm, MPI, and RCCL
Lux’s AMD Compilers, ROCm, MPI, and RCCL compatibilities are represented in Lmod.
In general, if you can module load a version, it should be compatible.
GNU compilers cannot compile HIP code, but CPU code should be ABI compatible. AMD compilers can link in compiled GNU CPU applications and libraries, even when generating GPU code.
Compatibility between MPI implementations and ROCm is required in order to use GPU-aware MPI. MPI installations on Lux are built to target specific versions of ROCm, and compatibility across multiple versions are not guaranteed. OLCF will maintain compatible default modules when possible.
RCCL is installed as part of the ROCm package/module, and the compatible version should always be included with the ROCm installation.
MPI Module |
Compatible ROCm Versions |
|---|---|
|
|
|
|
OpenMP
This section shows how to compile with OpenMP using the different compilers covered above.
Vendor |
Module |
Language |
Compiler |
OpenMP flag (CPU thread) |
|---|---|---|---|---|
AMD |
|
C
C++
Fortran
|
amdclangamdclang++amdflang |
|
GNU |
|
C
C++
Fortran
|
gccg++gfortran |
|
OpenMP GPU Offload
This section shows how to compile with OpenMP Offload using the different compilers covered above.
Vendor |
Module |
Language |
Compiler |
OpenMP flag (GPU) |
|---|---|---|---|---|
AMD |
|
C
C++
Fortran
|
amdclangamdclang++amdflang |
|
HIP
This section shows how to compile HIP codes using the AMD compilers and hipcc compiler driver.
Compiler |
Compile/Link Flags, Header Files, and Libraries |
|---|---|
|
CFLAGS = -std=c++11 -D__HIP_ROCclr__ -D__HIP_ARCH_GFX950__=1 --rocm-path=${ROCM_PATH} --offload-arch=gfx950 -x hip-I${ROCM_PATH}/includeLFLAGS = --rocm-path=${ROCM_PATH}-L${ROCM_PATH}/lib -lamdhip64 |
|
Can be used directly to compile HIP source files.
To see what is being invoked within this compiler driver, issue the command
hipcc --verboseTo explicitly target AMD MI355X, use
--offload-arch=gfx950 |
Note
hipcc requires the ROCm Toolclain, See Exposing The ROCm Toolchain to your Programming Environment
HIP + OpenMP CPU Threading
This section shows how to compile HIP + OpenMP CPU threading hybrid codes.
Vendor |
Compiler |
Compile/Link Flags, Header Files, and Libraries |
|---|---|---|
AMD |
|
CFLAGS = -std=c++11 -D__HIP_ROCclr__ -D__HIP_ARCH_GFX950__=1 --rocm-path=${ROCM_PATH} --offload-arch=gfx950 -x hip -fopenmp-I${ROCM_PATH}/includeLFLAGS = --rocm-path=${ROCM_PATH} -fopenmp-L${ROCM_PATH}/lib -lamdhip64 |
|
Can be used to directly compile HIP source files, add
-fopenmp flag to enable OpenMP threadingTo explicitly target AMD MI355X, use
--offload-arch=gfx950 |
Note
hipcc requires the ROCm Toolclain, See Exposing The ROCm Toolchain to your Programming Environment
Running Jobs
Computational work on Lux is performed by jobs. Jobs typically consist of several components:
A batch submission script
A binary executable
A set of input files for the executable
A set of output files created by the executable
In general, the process for running a job is to:
Prepare executables and input files.
Write a batch script.
Submit the batch script to the batch scheduler.
Optionally monitor the job before and during execution.
The following sections describe in detail how to create, submit, and manage jobs for execution on Lux. Lux uses SchedMD’s Slurm Workload Manager as the batch scheduling system.
Login vs Compute Nodes
Recall from the System Overview that Lux contains two node types: Login and Compute. When you connect to the system, you are placed on a login node. Login nodes are used for tasks such as code editing, compiling, etc. They are shared among all users of the system, so it is not appropriate to run tasks that are long/computationally intensive on login nodes. Users should also limit the number of simultaneous tasks on login nodes (e.g., concurrent tar commands, parallel make).
Compute nodes are the appropriate place for long-running, computationally-intensive tasks. When you start a batch job, your batch script (or interactive shell for batch-interactive jobs) runs on one of your allocated compute nodes.
Warning
Compute-intensive, memory-intensive, or other disruptive processes running on login nodes may be killed without warning.
Slurm
Lux uses SchedMD’s Slurm Workload Manager for scheduling and managing jobs. Slurm maintains similar functionality to other schedulers such as IBM’s LSF, but provides unique control of Lux’s resources through custom commands and options specific to Slurm. A few important commands can be found in the conversion table below, but please visit SchedMD’s Rosetta Stone of Workload Managers for a more complete conversion reference.
Slurm documentation for each command is available via the man utility, and on the web at https://slurm.schedmd.com/man_index.html.
Additional documentation is available at https://slurm.schedmd.com/documentation.html.
Some common Slurm commands are summarized in the table below. More complete examples are given in the Monitoring and Modifying Batch Jobs section of this guide.
Command |
Action/Task |
|---|---|
|
Show the current queue |
|
Submit a batch script to allocate a Slurm job allocation. The script contains options preceded with |
|
Submit an interactive job, where one or more job steps (i.e., |
|
Launch a parallel jobon resources allocated with
sbatch or salloc.If necessary,
srun will first create a resource |
|
Show node/partition info |
|
View accounting information for jobs/job steps |
|
Cancel a job or job step |
|
View or modify job configuration. |
General information for Node-sharing on Lux
Lux is a node-shared Slurm cluster: multiple users may run on the same physical node at the same time, as long as their resource requests do not overlap. Node sharing on Lux is facilitated through Slurm allocations of CPU cores, memory, and GPUs.
Allocations on Lux are handled in quantized pieces per GPU, where each GPU receives its closest CPUs in the nearest NUMA domain. Additionally, 375GB of memory is allocated for each GPU + NUMA.
When constructing a job on Lux, please be aware of the two-phase resource allocation steps within Slurm.
Phase |
Location |
Description |
|---|---|---|
Allocation |
Login |
Request resources with This is where you should request what you need:
|
Delegation |
Compute |
Launch work with This is where you “hand out” the resources you already requested to the actual processes (potentially with multiple |
If a job requires all the resources on a node, users can use the –exclusive flag to disable node-sharing functionality and give the job sole access to the nodes in that allocation.
Queues on Lux
The compute nodes on Lux are in a single partition, the “batch partition” of compute nodes as described in Lux Compute Nodes.
The scheduling policies for the batch partition are described below.
Users may have up to 200 jobs queued at any time.
Batch Partition Policy (default)
Bin |
Node Count |
Duration |
Policy |
|---|---|---|---|
A |
1-502 Nodes |
Duration 0-48 hr |
Max 4 jobs running and 4 jobs eligible per project |
Job Limit
Entity Level |
Max Jobs in Queue |
|---|---|
User |
200 |
Node-Hour Calculation
Jobs on Lux are scheduled in partial-node increments. The OLCF charges based on what a job makes unavailable to other users, so users are encouraged to only use what their job requires. Allocations on Lux are separate from those on Frontier and other OLCF resources.
The node-hour charge for each job will be calculated as follows:
node-hours = ({weight for resource} * {Resource used}) * ( batch job endtime - batch job starttime )
Lux weighs GPUs as 100% of the node, and each GPU comes with a quantized portion of cores (16) and RAM (375GB). The node-hour calculation is as follows:
node-hours = ( .125 * {Number of GPUs} ) * ( batch job endtime - batch job starttime )
Where batch job starttime is the time the job moves into a running state, and batch job endtime is the time the job exits a running state.
A batch job’s usage is calculated solely on requested resources, calculated after quantization, and the batch job’s start and end time. The number of CPUs or GPUs actually used within any particular allocation is not used in the calculation. For example, if a job requests 6 GPUs through the batch script, runs for (1) hour, uses only (8) CPU cores, and (1) GPU, the job will still be charged for .75 node-hours. Similarly, if a job requests (1) hour, but exits after (0.5) hours, then the job will only be charged for the (0.5) hours.
e.g. node-hours = ( .125 * 6 ) * ( .5 ) = .375
Batch Scripts
The most common way to interact with the batch system is via batch scripts. A batch script is simply a shell script with added directives to request various resources from or provide certain information to the scheduling system. Aside from these directives, the batch script is simply the series of commands needed to set up and run your job.
To submit a batch script, use the command sbatch myjob.sl
Consider the following batch script:
1#!/bin/bash
2#SBATCH -A ABC123
3#SBATCH -J RunSim123
4#SBATCH -o %x-%j.out
5#SBATCH -t 1:00:00
6#SBATCH -p batch
7#SBATCH -N 4
8#SBATCH --gpus 4
9
10cd $MEMBERWORK/abc123/Run.456
11cp $PROJWORK/abc123/RunData/Input.456 ./Input.456
12srun ...
13cp my_output_file $PROJWORK/abc123/RunData/Output.456
In the script, Slurm directives are preceded by #SBATCH, making them appear as comments to the shell. Slurm looks for these directives through the first non-comment, non-whitespace line. Options after that will be ignored by Slurm (and the shell).
Line |
Description |
|---|---|
1 |
Shell interpreter line |
2 |
OLCF project to charge |
3 |
Job name |
4 |
Job standard output file ( |
5 |
Walltime requested (in |
6 |
Partition (queue) to use |
7 |
Number of compute nodes requested |
8 |
Number of GPUs requested |
9 |
Blank line |
10 |
Change into the run directory |
11 |
Copy the input file into place |
12 |
Run the job ( add layout details ) |
13 |
Copy the output file to an appropriate location. |
Example Compile and Run
The following will compile the hello_jobstep application found on ORNL’s GitLab.
module load rocm
module load openmpi
hipcc -std=c++11 -fopenmp --offload-arch=gfx950 -I${ROCM_PATH}/include -I${MPI_DIR}/include -c hello_jobstep.cpp
hipcc -fopenmp -L${ROCM_PATH}/lib -lamdhip64 -L${MPI_DIR}/lib -lmpi hello_jobstep.o -o hello_jobstep
The following can run hello_jobstep:
#!/bin/bash
#SBATCH --account stf007
#SBATCH --time 05:00
#SBATCH --partition batch
#SBATCH --nodes 1
#SBATCH --gpus 8
#SBATCH --job-name hello_jobstep
#SBATCH --output %j-%x.out
#SBATCH --error %j-%x.err
module load rocm
module load openmpi
OMP_NUM_THREADS=1 srun -N1 -n4 -c32 -G8 --gpu-bind=closest ./hello_jobstep
Interactive Jobs
Most users will find batch jobs an easy way to use the system, as they allow you to “hand off” a job to the scheduler, allowing them to focus on other tasks while their job waits in the queue and eventually runs. Occasionally, it is necessary to run interactively, especially when developing, testing, modifying or debugging a code.
Since all compute resources are managed and scheduled by Slurm, it is not possible to simply log into the system and immediately begin running parallel codes interactively.
Rather, you must request the appropriate resources from Slurm and, if necessary, wait for them to become available. This is done through an “interactive batch” job.
Interactive batch jobs are submitted with the salloc command. Resources are requested via the same options that are passed via #SBATCH in a regular batch script (but without the #SBATCH prefix).
For example, to request an interactive batch job with the same resources that the batch script above requests, you would use salloc -A ABC123 -J RunSim123 -t 1:00:00 -p batch -N 4 --gpus 8.
Note there is no option for an output file…you are running interactively, so standard output and standard error will be displayed to the terminal.
Warning
Indicating your shell in your salloc command is NOT recommended (e.g., salloc ... /bin/bash). Doing so causes your compute job to start on a login node by default rather than automatically moving you to a compute node.
Common Slurm Options
The table below summarizes options for submitted jobs.
Unless otherwise noted, they can be used for either batch scripts or interactive batch jobs.
For scripts, they can be added on the sbatch command line or as a #SBATCH directive in the batch script.
(If they’re specified in both places, the command line takes precedence.)
This is only a subset of all available options.
Check the Slurm Man Pages for a more complete list.
Option |
Example Usage |
Description |
|---|---|---|
|
|
Specifies the project to which the job should be charged. |
|
|
Request 128 nodes for the job. |
|
|
Specify the default number of tasks in each |
|
|
Requests 8 GPUs (total) for the job. |
|
|
Requests 1 GPU per task. |
|
|
Request a walltime of 4 hours. A walltime request is the maximum amount of time a job will run and can be specified as minutes, hours:minutes, hours:minutes:seconds, days-hours, days-hours:minutes, or days-hours:minutes:seconds |
|
|
Specify job dependency (in this example, this job cannot start until job 12345 exits with an exit code of 0. See the Job Dependency section for more information) |
|
|
Specify the job name. (this will show up in queue listings) |
|
|
File where job STDOUT will be directed (%j will be replaced with the job ID). If no |
|
|
File where job STDERR will be directed (%j will be replaced with the job ID). If no |
|
|
Send email for certain job actions. Can be a comma-separated list. Actions include BEGIN, END, FAIL, REQUEUE, INVALID_DEPEND, STAGE_OUT, ALL, and more. |
|
|
Email address to be used for notifications. |
|
|
Instructs Slurm to run a job on nodes that are part of the specified reservation. |
|
|
Send the given signal to a job the specified time (in seconds) seconds before the job reaches its walltime. The signal can be by name or by number (i.e. both 10 and USR1 would send SIGUSR1). Signaling a job can be used, for example, to force a job to write a checkpoint just before Slurm kills the job (note that this option only sends the signal; the user must still make sure their job script traps the signal and handles it in the desired manner). When used with |
|
|
Request a specific compute partition for the job. (default is |
|
|
Request a “Quality of Service” (QOS) for the job. (default is |
Warning
Setting --threads-per-core > 1 on Lux will not have an effect and will result in submissions being rejected.
The default is --threads-per-core=1 because Lux does not have simultaneous multithreading (SMT) enabled, and values greater than 1 are invalid.
Slurm Environment Variables
Slurm reads a number of environment variables, many of which can provide the same information as the job options noted above. We recommend using the job options rather than environment variables to specify job options, as it allows you to have everything self-contained within the job submission script (rather than having to remember what options you set for a given job).
Slurm also provides a number of environment variables within your running job. The following table summarizes those that may be particularly useful within your job (e.g., for naming output log files):
Variable |
Description |
|---|---|
|
The directory from which the batch job was submitted. By default, a new job starts
in your home directory. You can get back to the directory of job submission with
|
|
The job’s full identifier. A common use for |
|
The number of nodes requested. |
|
The job name supplied by the user. |
|
The list of nodes assigned to the job. |
Job States
A job will transition through several states during its lifetime. Common ones include:
State Code |
State |
Description |
|---|---|---|
CA |
Canceled |
The job was canceled (could’ve been by the user or an administrator) |
CD |
Completed |
The job completed successfully (exit code 0) |
CG |
Completing |
The job is in the process of completing (some processes may still be running) |
PD |
Pending |
The job is waiting for resources to be allocated |
R |
Running |
The job is currently running |
Job Reason Codes
In addition to state codes, jobs that are pending will have a “reason code” to explain why the job is pending. Completed jobs will have a reason describing how the job ended. Some codes you might see include:
Reason |
Meaning |
|---|---|
Dependency |
Job has dependencies that have not been met |
JobHeldUser |
Job is held at user’s request |
JobHeldAdmin |
Job is held at system administrator’s request |
Priority |
Other jobs with higher priority exist for the partition/reservation |
Reservation |
The job is waiting for its reservation to become available |
AssocMaxJobsLimit |
The job is being held because the user/project has hit the limit on running jobs |
ReqNodeNotAvail |
The requested a particular node, but it’s currently unavailable (it’s in use, reserved, down, draining, etc.) |
JobLaunchFailure |
Job failed to launch (could due to system problems, invalid program name, etc.) |
NonZeroExitCode |
The job exited with some code other than 0 |
Many other states and job reason codes exist.
For a more complete description, see the squeue man page (either on the system or online).
System Reservation Policy
Projects may request to reserve a set of nodes for a period of time by contacting help@olcf.ornl.gov. If the reservation is granted, the reserved nodes will be blocked from general use for a given period of time. Only users that have been authorized to use the reservation can utilize those resources. Since no other users can access the reserved resources, it is crucial that groups given reservations take care to ensure the utilization on those resources remains high. To prevent reserved resources from remaining idle for an extended period of time, reservations are monitored for inactivity.
The requesting project’s allocation is charged according to the time window granted, regardless of actual utilization. For example, an 8-hour, 200 node reservation on Lux would be equivalent to using 1,600 Lux node-hours of a project’s allocation.
Note
Reservations should not be confused with priority requests. If quick turnaround is needed for a few jobs or for a period of time, a priority boost should be requested. A reservation should only be requested if users need to guarantee availability of a set of nodes at a given time, such as for a live demonstration at a conference.
Job Dependencies
Oftentimes, a job will need data from some other job in the queue, but it’s nonetheless convenient to submit the second job before the first finishes.
Slurm allows you to submit a job with constraints that will keep it from running until these dependencies are met.
These are specified with the -d option to Slurm.
Common dependency flags are summarized below.
In each of these examples, only a single jobid is shown but you can specify multiple job IDs as a colon-delimited list (i.e. #SBATCH -d afterok:12345:12346:12346).
For the after dependency, you can optionally specify a +time value for each jobid.
Flag |
Meaning (for the dependent job) |
|---|---|
|
The job can start after the specified jobs start or are canceled.
The optional |
|
The job can start after the specified jobs have ended (regardless of exit state) |
|
The job can start after the specified jobs terminate in a failed (non-zero) state |
|
The job can start after the specified jobs complete successfully (i.e. zero exit code) |
|
Job can begin after any previously-launched job with the same name and from the same user have completed. In other words, serialize the running jobs based on username+jobname pairs. |
Monitoring and Modifying Batch Jobs
scontrol hold and scontrol release: Holding and Releasing Jobs
Sometimes you may need to place a hold on a job to keep it from starting.
For example, you may have submitted it assuming some needed data was in place but later realized that data is not yet available.
This can be done with the scontrol hold command.
Later, when the data is ready, you can release the job (i.e. tell the system that it’s now OK to run the job) with the scontrol release command.
For example:
|
Place job 12345 on hold |
|
Release job 12345 (i.e. tell the system it’s OK to run it) |
scontrol update: Changing Job Parameters
There may also be occasions where you want to modify a job that’s waiting in the queue.
For example, perhaps you requested 200 nodes but later realized this is a different data set and only needs 100 nodes.
You can use the scontrol update command for this.
For example:
|
Change job 12345’s node request to 100 nodes |
|
Change job 12345’s max walltime to 4 hours |
scancel: Cancel or Signal a Job
In addition to the --signal option for the sbatch/salloc commands described above, the scancel command can be used to manually signal a job.
Typically, this is used to remove a job from the queue.
In this use case, you do not need to specify a signal and can simply provide the jobid (i.e. scancel 12345).
If you want to send some other signal to the job, use scancel the with the -s option.
The -s option allows signals to be specified either by number or by name.
Thus, if you want to send SIGUSR1 to a job, you would use scancel -s 10 12345 or scancel -s USR1 12345.
squeue: View the Queue
The squeue command is used to show the batch queue.
You can filter the level of detail through several command-line options.
For example:
|
Show all jobs currently in the queue |
|
Show all of your jobs currently in the queue |
sacct: Get Job Accounting Information
The sacct command gives detailed information about jobs currently in the queue and recently-completed jobs.
You can also use it to see the various steps within a batch jobs.
|
Show all jobs ( |
|
Show all of your jobs, and show the individual steps (since there was no |
|
Show all job steps that are part of job 12345 |
|
Show all of your jobs since 1 PM on October 1, 2026 using a particular output format |
scontrol show job: Get Detailed Job Information
In addition to holding, releasing, and updating the job, the scontrol command can show detailed job information via the show job subcommand.
For example, scontrol show job 12345.
srun: Run Jobs/Steps
The default job launcher for Lux is srun .
The srun command is used to execute an MPI or RCCL-enabled binary on one or more compute nodes in parallel.
Srun Format
srun [OPTIONS... [executable [args...]]]
Single Command (non-interactive)
$ srun --account <project_id> --time 00:05:00 --partition <partition> --nodes 2 --ntasks 4 --ntasks-per-node=2 --gpus-per-task=1 ./a.out
<output printed to terminal>
The job name and output options have been removed since stdout/stderr are typically desired in the terminal window in this usage mode.
srun accepts the following common options:
|
Number of nodes |
|
Total number of MPI tasks (default is 1) |
|
Logical cores per MPI task (default is 1)
This is equivalent to physical cores per task
By default, when
-c > 1, additional cores per task are distributed within one L3 region first before filling a different L3 region. |
|
Bind tasks to CPUs.
cores - Automatically generate masks binding tasks to physical cores. |
|
Specifies the distribution of MPI ranks across compute nodes, sockets (L3 regions), and cores, respectively.
The default values are
block:cyclic:cyclic, see man srun for more information.Currently, the distribution setting for cores (the third “<value>” entry) has no effect on Lux.
|
|
If used without
-n: requests that a specific number of tasks be invoked on each node.If used with
-n: treated as a maximum count of tasks per node. |
|
Specify the number of GPUs required for the job (total GPUs across all nodes). |
|
Specify the number of GPUs per task required for the job. Requires an explicit task count ( |
|
Binds each task to the GPU which is on the same NUMA domain as the CPU core the MPI rank is running on. |
|
Bind tasks to specific GPUs by setting GPU masks on tasks (or ranks) as specified where |
|
Request that there are ntasks tasks invoked for every GPU. |
Process and Thread Mapping Examples
This section describes how to map processes (e.g., MPI ranks) and process threads (e.g., OpenMP threads) to the CPUs, GPUs, and NICs on Lux.
Users are highly encouraged to use the CPU- and GPU-mapping programs used in
the following sections to check their understanding of the job steps (i.e.,
srun commands) they intend to use in their actual jobs.
For the CPU Mapping and Multithreading sections:
An MPI+OpenMP+HIP “Hello, World” program (hello_jobstep) will be used to clarify the GPU and CPU mappings.
Additionally, it may be helpful to cross reference the Lux node diagram
hello_jobstep output
Before jumping into the examples, it is helpful to understand the output from the hello_jobstep program:
ID |
Description |
|---|---|
|
MPI rank ID |
|
OpenMP thread ID |
|
CPU hardware thread the MPI rank or OpenMP thread ran on |
|
Compute node the MPI rank or OpenMP thread ran on |
|
GPU ID the MPI rank or OpenMP thread had access to
(This is the node-level, or global, GPU ID as shown in the Lux node diagram)
NOTE: This is read from
ROCR_VISIBLE_DEVICES. If this variable is not set, the value of
GPU_ID will be set to N/A by the program |
|
The runtime GPU ID
(This is the GPU ID as seen from the HIP runtime - e.g., as reported by
hipGetDevice)NOTE: The HIP runtime relabels the GPUs each rank can access starting at 0
|
|
The physical Bus ID associated with a GPU
(The Bus ID can be used to e.g., confirm unique GPUs are being used)
|
CPU Mapping
This subsection covers how to map tasks to the CPU without the presence of additional threads (i.e., solely MPI tasks – no additional OpenMP threads).
The intent with both of the following examples is to launch 8 MPI ranks across
the node where each rank is assigned its own logical (and, in this case,
physical) core. Using the -m distribution flag, we will cover two common
approaches to assign the MPI ranks – in a “round-robin” (cyclic)
configuration and in a “packed” (block) configuration. Slurm’s
Interactive Jobs method was used to request an allocation of 1
compute node with 2 GPUs for these examples: salloc -A <project_id> -t 30 -p <parition>
-N 1 --gpus 2
Note
There are many different ways users might choose to perform these mappings,
so users are encouraged to clone the hello_jobstep program and test whether
or not processes and threads are running where intended.
8 MPI Ranks (round-robin)
Assigning MPI ranks in a “round-robin” (cyclic) manner across NUMA
domains (sockets) is the default behavior on Lux. This mode will assign
consecutive MPI tasks to different sockets before it tries to “fill up” a
socket.
Recall that the -m flag behaves like: -m <node distribution>:<socket
distribution>. Hence, the key setting to achieving the round-robin nature is
the -m block:cyclic flag, specifically the cyclic setting provided for
the “socket distribution”. This ensures that the MPI tasks will be distributed
across sockets in a cyclic (round-robin) manner.
The below srun command will achieve the intended 8 MPI “round-robin” layout:
$ export OMP_NUM_THREADS=1
$ srun -N1 -n8 -c1 --cpu-bind=threads -m block:cyclic ./hello_jobstep | sort | cut -d "-" -f 1-4
MPI 000 - OMP 000 - HWT 000 - Node lux089
MPI 001 - OMP 000 - HWT 016 - Node lux089
MPI 002 - OMP 000 - HWT 001 - Node lux089
MPI 003 - OMP 000 - HWT 017 - Node lux089
MPI 004 - OMP 000 - HWT 002 - Node lux089
MPI 005 - OMP 000 - HWT 018 - Node lux089
MPI 006 - OMP 000 - HWT 003 - Node lux089
MPI 007 - OMP 000 - HWT 019 - Node lux089
Breaking down the srun command, we have:
-N1: indicates we are using 1 node-n8: indicates we are launching 8 MPI tasks-c1: indicates we are assigning 1 logical core per MPI task. In this case, because of--threads-per-core=1, this also means 1 physical core per MPI task.--cpu-bind=threads: binds tasks to threads--threads-per-core=1: use a maximum of 1 hardware thread per physical core (i.e., only use 1 logical core per physical core)-m block:cyclic: distribute the tasks in a block layout across nodes (default), and in a cyclic (round-robin) layout across L3 sockets./hello_mpi_omp: launches the “hello_mpi_omp” executable| sort: sorts the output| cut ...: gets only output we care about for CPU tasks
Note
Although the above command used the default settings -c1,
--cpu-bind=threads, --threads-per-core=1 and -m block:cyclic, it is
always better to be explicit with your srun command to have more control
over your node layout. The above command is equivalent to srun -N1 -n8.
As you can see in the node diagram above, this results in the 8 MPI tasks (outlined in different colors) being distributed “vertically” across NUMA sockets initially, then “horizontally” across cores.
7 MPI Ranks (packed)
Instead, you can assign MPI ranks so that the L3 regions are filled in a
“packed” (block) manner. This mode will assign consecutive MPI tasks to
the same L3 region (socket) until it is “filled up” or “packed” before
assigning a task to a different socket.
Recall that the -m flag behaves like: -m <node distribution>:<socket
distribution>. Hence, the key setting to achieving the round-robin nature is
the -m block:block flag, specifically the block setting provided for
the “socket distribution”. This ensures that the MPI tasks will be distributed
in a packed manner.
The below srun command will achieve the intended 7 MPI “packed” layout:
$ export OMP_NUM_THREADS=1
$ srun -N1 -n7 -c1 --cpu-bind=threads -m block:block ./hello_jobstep | sort | cut -d "-" -f 1-4
MPI 000 - OMP 000 - HWT 000 - Node lux089
MPI 001 - OMP 000 - HWT 001 - Node lux089
MPI 002 - OMP 000 - HWT 002 - Node lux089
MPI 003 - OMP 000 - HWT 003 - Node lux089
MPI 004 - OMP 000 - HWT 004 - Node lux089
MPI 005 - OMP 000 - HWT 005 - Node lux089
MPI 006 - OMP 000 - HWT 006 - Node lux089
Breaking down the srun command, the only difference than the previous example is:
-m block:block: distribute the tasks in a block layout across nodes (default), and in a block (packed) socket layout
As you can see in the node diagram above, this results in the 7 MPI tasks (outlined in different colors) being distributed “horizontally” within a socket, rather than being spread across different L3 sockets like with the previous example.
Multithreading
Because a Lux compute node has one hardware thread available per core (1 logical cores per physical core), multithreaded applications (e.g., with OpenMP threads) will run on individual cores assigned to a task.
The following examples cover multithreading with hybrid MPI+OpenMP applications.
In these examples, Slurm’s Interactive Jobs method was used to request an allocation of 1 compute node:
salloc -A <project_id> -t 30 -p <parition> -N 1 --gpus 2
Note
There are many different ways users might choose to perform these mappings,
so users are encouraged to clone the hello_jobstep program and test whether
or not processes and threads are running where intended.
2 MPI ranks - each with 2 OpenMP threads
In this example, the intent is to launch 2 MPI ranks, each of which spawn 2 OpenMP threads, and have all of the 4 OpenMP threads run on different physical CPU cores.
First (INCORRECT) attempt
To set the number of OpenMP threads spawned per MPI rank, the
OMP_NUM_THREADS environment variable can be used. To set the number of MPI
ranks launched, the srun flag -n can be used.
$ export OMP_NUM_THREADS=2
$ srun -N1 -n2 ./hello_jobstep | sort | cut -d "-" -f 1-4
MPI 000 - OMP 000 - HWT 015 - Node lux089
MPI 000 - OMP 001 - HWT 007 - Node lux089
MPI 001 - OMP 000 - HWT 031 - Node lux089
MPI 001 - OMP 001 - HWT 017 - Node lux089
The most notable features are the placement of the jobs at the end of their respective NUMA sockets, and multithreaded cores landing in non-determined locations.
The problem here arises from two default settings; 1) each MPI rank is only
allocated 1 core (-c 1) and, 2) only 1 hardware thread per physical CPU core is enabled (--threads-per-core=1).
When using --threads-per-core=1 and --cpu-bind=threads (the default setting), 1 logical core in -c is equivalent to 1 physical core.
So in this case, each MPI rank only has 1 physical core (with 1 hardware thread) to run on -
including any threads the process spawns - hence the undesired behavior.
Second (CORRECT) attempt
Recall that in this scenario, because of the --threads-per-core=1 setting, 1 logical core is equivalent to 1 physical core when using -c.
Therefore, in order for each OpenMP thread to run on its own physical CPU core, each MPI rank should be given 2 physical CPU cores (-c 2).
Now the OpenMP threads will be mapped to unique hardware threads on separate physical CPU cores.
$ export OMP_NUM_THREADS=2
$ srun -N1 -n2 -c2 ./hello_jobstep | sort | cut -d "-" -f 1-4
MPI 000 - OMP 000 - HWT 001 - Node lux089
MPI 000 - OMP 001 - HWT 000 - Node lux089
MPI 001 - OMP 000 - HWT 017 - Node lux089
MPI 001 - OMP 001 - HWT 016 - Node lux089
Now the output shows that each OpenMP thread ran on its own physical CPU core. More specifically (see the Lux Compute Node diagram), OpenMP thread 000 of MPI rank 000 ran on logical core 001 (i.e., physical CPU core 01), OpenMP thread 001 of MPI rank 000 ran on logical core 000 (i.e., physical CPU core 00), OpenMP thread 000 of MPI rank 001 ran on logical core 017 (i.e., physical CPU core 17), and OpenMP thread 001 of MPI rank 001 ran on logical core 016 (i.e., physical CPU core 16) - as intended.
Ensemble Jobs
For many applications and use cases, the ability to launch many copies of the same binary in an independent context is needed. This section highlights a few recommended solutions to launching ensemble runs on Lux.
Before covering the tools that can be useful for this, be advised that the most reliable solution to this problem will be the use of MPI sub-communicators by your application.
For example, the LAMMPS Molecular Dynamics software supports a partition command, which can create many independent simulations from a single srun launch.
Single-process ensemble members
If you are able to fit each ensemble member onto a single MPI rank and single AMD Instinct MI355X GPU (8 GPU’s per node), the most reliable solution is to use a single srun as follows:
srun -N $SLURM_NNODES -n $((SLURM_NNODES*8)) -c 16 --gpus-per-task=1 --gpu-bind=closest ./wrapper.sh
Where wrapper.sh is a shell script that launches your application.
This shell script is simply for convenience, in case you wish to vary the parameters to your application based on MPI rank.
Using multiple simultaneous srun’s
If you are not able to fit each ensemble member onto a single MPI rank and GCD, a common approach is to launch multiple srun processes in the background simultaneously.
For example:
for node in $(scontrol show hostnames); do
srun -N 1 -n 8 -c 7 --gpus 8 --gpus-per-task=1 --gpu-bind=closest <executable> &
done
# Wait for srun's to all finish
wait
Each srun communicates to the Slurm controller node (which is shared among all users) when it is launched.
Large amounts of srun processes can temporarily overwhelm the Slurm controller, making commands like sbatch and squeue hang.
This approach can be fast, but is unreliable and does not scale, and potentially overloads the Slurm controller.
We do not yet recommend this approach beyond 100 simultaneous srun’s.
Slurm version 26.05 includes the --stepmgr flag for sbatch, which uses the first node in the allocation to manage job steps instead of the Slurm controller.
This feature may substantially improve the ability to run many simultaneous srun’s.
Tips for Launching at Scale
Debugging
GDB
GDB, the GNU Project Debugger, is a command-line debugger useful for traditional debugging and investigating code crashes. GDB lets you debug programs written in Ada, C, C++, Objective-C, Pascal (and many other languages).
GDB is available on Lux installed by default to:
/usr/bin/gdb
To use GDB to debug your application run:
gdb ./path_to_executable
Additional information about GDB usage can befound on the GDB Documentation Page.
Profiling
Getting Started with the ROCm Profiler
Rocprof v3
rocprof gathers metrics on kernels run on AMD GPU architectures. The profiler works for HIP kernels, as well as offloaded kernels from OpenMP target offloading, OpenCL, and abstraction layers such as Kokkos.
rocprofv3 was introduced in ROCm/6.2 and utilizes the new rocprofiler API in ROCm.
The same information can be queried as with rocprof, but the command-line flags for rocprofv3 are slightly different than rocprof.
For example, to get a simple view of kernels being run, you will want to use rocprofv3 --kernel-trace --stats -- ./myexecutable instead of rocprof --stats ./myexecutable.
rocprofv3 will default output to files named based on the process ID of the profiled run.
In the previous kernel tracing command, the stats will be found in a file named <somePID>_kernel_stats.csv.
You can use the --output flag to override the resulting CSV file name.
More detailed infromation on rocprof profiling modes can be found at ROCm Profiler documentation.
Tips and Tricks
Running with MPI
MPI on Lux is currently best supported up to 16 nodes.
OLCF strongly recommends your MPI workloads are limited to 16 nodes.
The following environment variable can help maximize performance when scaling MPI on Lux:
UCX_TLS=sm,self,rocm_copy,rocm_ipc,cma,rc_verbs,tcp:aux
