# CU e-Science HTC-HPC Cluster and Cloud

Joint collaboration between Department of Physics, Faculty of Science and Department of Computer Engineering, Faculty of Engineering, Chulalongkorn University.

The CU e-Science cluster is a part of the National e-Science Infrastructure Consortium. Our cluster serves both CU and external users. Currently, we provide&#x20;

1. **Slurm cluster**: Jobs can be run in interactive and batch modes. You can see examples of how to use the cluster in section SLURM. Currently, Centos 7 is the main operating system for both frontend and worker nodes.
2. **Kubernetes**: Online for CU users.

![CU e-Science cluster, MHMK Building, Faculty of Science, Chulalongkorn University](/files/2R1DuKNYuR63nfuJHD9K)


# Our resources

**CPU and memory**

| Machine                 | CPUs/node                                                      | Memory (GB)/node         | No. of nodes | Note                                                               |
| ----------------------- | -------------------------------------------------------------- | ------------------------ | ------------ | ------------------------------------------------------------------ |
| **Frontend**            |                                                                |                          |              |                                                                    |
| Lenovo System X 3550 M5 | 20 (Intel Xeon CPU E5-2640 v4 2.40GHz) with HT on (40 threads) | 32                       | 1            | `escience1.sc.chula.ac.th`                                         |
| Lenovo System X 3550 M5 | 16 (Intel Xeon CPU E5-2620 v4 2.10GHz)                         | 64                       | 1            | `escience2.sc.chula.ac.th`                                         |
| Lenovo SR630            | 32 (Intel Xeon Gold 5218 2.3GHz)                               | 8 x 32GB TruDDR4 2933MHz | 1            | <p>1x Tesla T4 GPU</p><p><code>escience3.sc.chula.ac.th</code></p> |
| **Worker: Slurm**       |                                                                |                          |              |                                                                    |
| Lenovo SR630            | 32 (Intel Xeon Gold 5218 2.3GHz)                               | 8 x 32GB TruDDR4 2933MHz | 7            | <p>1x Tesla T4 GPU/node</p><p>HPC, HTC</p>                         |
| Lenovo x3850 X6         | 80 (Intel Xeon E7-8870v4 2.1 MHz)                              | 512                      | 1            | HPC, HTC                                                           |
| Lenovo SR850            | 88 (Intel Xeon Gold 6152 2.10GHz)                              | 324                      | 1            | `escience4.sc.chula.ac.th`                                         |
| IBM BladeCenter HS22    | 16 (Intel Xeon CPU E5-2650 2.00GHz)                            | 32                       | 5            | HTC                                                                |
| IBM iDataPlex DX360M4   | 16                                                             | 128                      | 2            | gridMathematica                                                    |
| Lenovo SR635            | 16 (AMD EPYC 7313P 3.0 GHz)                                    | 256                      | 2            | <p>1 machine with Nvidia T4, <br>1 machine with Nvidia A2</p>      |
| DGXStation              |                                                                |                          | 1            |                                                                    |
| **Worker: Kubernetes**  |                                                                |                          |              |                                                                    |
| Dell PowerEdge R740     | 24 (Intel Xeon Pentium 8268 2.9 GHz)                           | 6 x 64GB DDR4 2933MHz    | 3            |                                                                    |
| Lenovo SR630            | 32 (Intel Xeon Gold 5218 2.3GHz)                               | 8 x 32GB TruDDR4 2933MHz | 2            | 1x Tesla T4 GPU/node                                               |
| **TOTAL**               | **676 CPUs**                                                   |                          |              |                                                                    |

**Storage**

Currently, a total capacity after RAID6+SPARE:

1. IBM Storwize 3700: 160 TB
2. Lenovo ThinkSystem DE2000H: 160 TB

{% hint style="info" %}
*Backup is the responsibility of the user and it should never be understood that RAID is a backup!*&#x20;
{% endhint %}

| **File system**        | Disk space limit | Note                              |
| ---------------------- | ---------------- | --------------------------------- |
| `$Home`                | 100 GB/user      |                                   |
| /work/project/quantum  | 20 TB            |                                   |
| /work/project/cms      | 50 TB            |                                   |
| /work/project/physics  | 20 TB            | For Physics CU staff and students |
| /work/project/escience | 30 TB            | For all users                     |

To use the group space, please see the [disk space](/introduction-to-our-cluster/disk-space) section.


# Registration

**For the Slurm cluster**: a registration form can be found in [the form](https://forms.gle/443B1nT14aXoev1e6). Cluster admin will contact you when your registration is complete.

* When the user gets the confirmation email, it will come with the initial password. After the first login, user will be asked to change the password immediately.

**For the Kubenetes**: *Under construction* &#x20;


# Login to our cluster

### Slrum

To log in to the frontend, the ssh client program can be used, and the user types Command Line Interface (CLI) commands to tell the computer what to do. The ssh client program is available on Linux, MacOS (with Terminal) and MS Windows 10 machines (with PowerShell). For older Microsoft Windows machines, the [PuTTY![](https://twiki.cern.ch/twiki/pub/TWiki/TWikiDocGraphics/external-link.gif)](https://www.putty.org/) ssh client is recommended.

```
ssh your_user_name@escience0.sc.chula.ac.th
```

Note that **`escience0`** is the load-balancing IP address. It should be fine to use for any job submissions. However, if you would like to compile your code with specific hardware, i.e. GPU, you should log in to a specific machine. We don't recommend using the login machines to run your code if not necessary, jobs will be killed if it consumes a lot of resources of the frontend nodes.&#x20;

* `escience1.sc.chula.ac.th`: small frontend machine, for job submission, monitoring only.&#x20;
* `escience2.sc.chula.ac.th`: small frontend machine, for job submission, monitoring only.&#x20;
* `escience3.sc.chula.ac.th`: frontend with Tesla G4 CPU.

### Kubenetes

*Under construction.*


# Disk space

User home directory will be under `/work/home/your_username/` with the quota of 100 GB. You will have additional space in the group space. These group spaces are assigned during the account creation process. The group space is at `/work/project/`.&#x20;

| Group    | Users            | When the disk is full          |
| -------- | ---------------- | ------------------------------ |
| escience | All              | Central management             |
| physics  | Physics CU users | Group members manage the quota |
| cms      | CMS-CERN user    | Group members manage the quota |
| quantum  | Quantum team     | Group members manage the quota |


# Acknowledgement and Publication

#### Reporting research outputs

Please use [this link](https://forms.gle/9oFcQ3YuPQ4PwxEE8) to fill information on your published papers.

#### Acknowledgements

Please include the following acknowledgement in papers:

* **Acknowledgements to be included in "short letters"**\
  *We acknowledge the supporting computing infrastructure provided by NSTDA, CU, CUAASC,* NSRF via PMUB \[B39G680009] *(Thailand). URL:[www.e-science.in.th](http://www.e-science.in.th).*&#x20;
* **Acknowledgements to be included in "standard letters" and "long papers"**

  *The authors acknowledge the National Science and Technology Development Agency, National e-Science Infrastructure Consortium, Chulalongkorn University and the Chulalongkorn Academic Advancement into Its 2nd Century Project,* NSRF via the Program Management Unit for Human Resources & Institutional Development, Research and Innovation \[grant numbers B39G680009] *(Thailand) for providing computing infrastructure that has contributed to the research results reported within this paper. URL:[www.e-science.in.th](http://www.e-science.in.th).*


# 101: How to submit Slurm batch jobs

Introduction on how to submit the job to the Slurm cluster

Below is a sample Slurm script for running a Python code:&#x20;

You python script `example1.py`

```python
print("Hello World")
```

and the Slurm submission script `example1.slurm`

```bash
#!/bin/bash
#
#SBATCH --qos=cu_hpc
#SBATCH --partition=cpu
#SBATCH --job-name=example1
#SBATCH --output=example1.txt
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --time=00:10:00
#SBATCH --cpus-per-task=1
#SBATCH --mem-per-cpu=1G

module purge

#To get worker node information
hostname
uname -a
more /proc/cpuinfo | grep "model name" | head -1
more /proc/cpuinfo | grep "processor" | wc -l
echo "pwd = "`pwd`
echo "TMPDIR = "$TMPDIR
echo "SLURM_SUBMIT_DIR = "$SLURM_SUBMIT_DIR
echo "SLURM_JOBID = "$SLURM_JOBID

#To run python script
python example1.py
```

Note that,&#x20;

1. For `--qos`, you should check which qos that you are assigned. You can check by using `sacctmgr show assoc format=cluster,user,qos`
   1. QoS includes `cu_hpc, cu_htc, cu_math, cu_long, cu_student, escience`
2. For `--partition`, you can choose `cpu` or `cpugpu`for all QoS, except for cu\_math (use `math` partition).
3. See detail of QoS and partition [here](/slurm/slurm-qos-and-partition).
4. You can also use other shells if want, not limited to bash. See an example of tcsh/csh in the [CMSSW example](/slurm/slurm-examples/slurm-cmssw).

To submit the job, you use `sbatch`

```
sbatch example1.slurm
```

You will see

```
Submitted batch job 81942
```

To check if your job is in which state

```
squeue -u your_user_name
```

In the ST column, R is Running, PD is pending.

Your output should look like

```
==========================================
SLURM_JOB_ID = 81943
SLURM_NODELIST = cpu-bladeh-01
==========================================
cpu-bladeh-01.stg
Linux cpu-bladeh-01.stg 3.10.0-1127.el7.x86_64 #1 SMP Tue Mar 31 23:36:51 UTC 2020 x86_64 x86_64 x86_64 GNU/Linux
model name	: Intel(R) Xeon(R) CPU E5-2650 0 @ 2.00GHz
16
pwd = /work/home/your_user_name/slurm/example1
TMPDIR = /work/scratch/your_user_name/81943
SLURM_SUBMIT_DIR = /work/scratch/your_user_name/81943
SLURM_JOBID = 81943
Hello World
```

{% hint style="info" %}
With the Slurm output, you see that your job is running on the same directory that you submit the job (e.g. `/work/home/your_user_name/slurm/example1`. ***This is not recommended***. You should move the job to run on `$TMPDIR` (or `$SLURM_SUBMIT_DIR`) and copy the output back when the job is done. Here is an example of modified `example1.slurm`  to run on `$TMPDIR` and copy `test.log` (output of python script) back to your submission directory. The `$TMPDIR` will be deleted automatically after the job is done.
{% endhint %}

```bash
#!/bin/bash
#
#SBATCH --qos=cu_hpc
#SBATCH --partition=cpu
#SBATCH --job-name=example1
#SBATCH --output=example1.txt
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --time=00:10:00
#SBATCH --cpus-per-task=1
#SBATCH --mem-per-cpu=1G

module purge

#To get worker node information
hostname
uname -a
more /proc/cpuinfo | grep "model name" | head -1
more /proc/cpuinfo | grep "processor" | wc -l

#To set your submission directory
echo "pwd = "`pwd`
export MYCODEDIR=`pwd`

#Check PATHs
echo "MYCODEDIR = "$MYCODEDIR
echo "TMPDIR = "$TMPDIR
echo "SLURM_SUBMIT_DIR = "$SLURM_SUBMIT_DIR
echo "SLURM_JOBID = "$SLURM_JOBID

#Move to TMPDIR and run python script
cp example1.py $TMPDIR
cd $TMPDIR
python example1.py >| test.log
ls -l
cp -rf test.log $MYCODEDIR/
```


# 101: Interactive jobs with Slurm

You can use `salloc` to allocate resources in real-time to run an interactive batch job. Typically this is used to allocate resources and spawn a shell. The shell is then used to execute `srun` commands to launch parallel tasks. Interactive job is useful for tasks including data exploration, development, or (with X11 forwarding) visualization activities. The maximum walltime depends on the [QoS](/slurm/slurm-qos-and-partition) you have used.

For worker nodes with CPU and GPU:

```
salloc --qos=cu_hpc --partition=cpugpu
```

For worker nodes with CPU-only:

```
salloc --qos=cu_hpc --partition=cpu
```

After connection, you will get the message like

```
salloc: Granted job allocation 82025
salloc: Waiting for resource configuration
salloc: Nodes cpu-bladeh-01 are ready for job
```

and with `squeue -u your_user_name`, you will see

```
JOBID PARTITION     NAME           USER ST       TIME  NODES NODELIST(REASON)
82025       cpu interact your_user_name  R       1:59      1 cpu-bladeh-01
```

To run, you can use `srun`, e.g.&#x20;

```
[your_user_name@frontend-02 ~]$ srun hostname
cpu-bladeh-01.stg
```

To exit the interactive mode, you can use the command `exit`

```
[your_user_name@frontend-02 ~]$ exit
exit
salloc: Relinquishing job allocation 82025
```


# Basic Slurm commands

* [sacct](https://slurm.schedmd.com/sacct.html): display accounting data for all jobs and job steps in the Slurm database
* [sacctmgr](https://slurm.schedmd.com/sacctmgr.html): display and modify Slurm account information
* [salloc](https://slurm.schedmd.com/salloc.html): request an interactive job allocation
* [sattach](https://slurm.schedmd.com/sattach.html): attach to a running job step
* [sbatch](https://slurm.schedmd.com/sbatch.html): submit a batch script to Slurm
* [scancel](https://slurm.schedmd.com/scancel.html): cancel a job or job step or signal a running job or job step
* [scontrol](https://slurm.schedmd.com/scontrol.html): display (and modify when permitted) the status of Slurm entities.  Entities include:  jobs, job steps, nodes, partitions, reservations, etc.
* [sdiag](https://slurm.schedmd.com/sdiag.html): display scheduling statistics and timing parameters
* [sinfo](https://slurm.schedmd.com/sinfo.html): display node partition (queue) summary information
* [sprio](https://slurm.schedmd.com/sprio.html): display the factors that comprise a job's scheduling priority
* [squeue](https://slurm.schedmd.com/squeue.html): display the jobs in the scheduling queues, one job per line
* [sreport](https://slurm.schedmd.com/sreport.html): generate canned reports from job accounting data and machine utilization statistics
* [srun](https://slurm.schedmd.com/srun.html): launch one or more tasks of an application across requested resources
* [sshare](https://slurm.schedmd.com/sshare.html): display the shares and usage for each charge account and user
* [sstat](https://slurm.schedmd.com/sstat.html): display process statistics of a running job step
* [sview](https://slurm.schedmd.com/sview.html): a graphical tool for displaying jobs, partitions, reservations, and Blue Gene blocks


# QoS and Partition

**Partitions**

| Partition | Nodes                             | Allow QoS                                         |
| --------- | --------------------------------- | ------------------------------------------------- |
| cpugpu    | Lenovo SR630 with Tesla T4        | cu\_hpc, cu\_htc, cu\_long, cu\_student, escience |
| cpu       | Lenovo x3850 X6, Lenovo SR850     | cu\_hpc, cu\_htc, cu\_long, cu\_student, escience |
| math      | IBM iDataPlex DX360M4             | cu\_math                                          |
| profiling | Lenovo SR635 with Tesla T4 and A2 | cu\_profile                                       |
| dgx       | DGX station                       | cu\_hpc                                           |
| test      | IBM BladeCenter HS22              | cu\_hpc, cu\_htc, cu\_long, cu\_student, escience |

**Quality of Service (QoS)**\
Jobs request a QOS using the "--qos=" option to the `sbatch`, `salloc`, and `srun` commands.

| QoS           | Max nodes per user | Max jobs per user | Max CPU per user | Max memory (GB) per user | Max walltime (Day-HH:MM:SS) |
| ------------- | ------------------ | ----------------- | ---------------- | ------------------------ | --------------------------- |
| cu\_hpc       | 8                  | 20                | 128              | 512                      | 14-00:00:00                 |
| cu\_htc       | 1                  | 100               | 128              | 256                      | 30-00:00:00                 |
| cu\_long      | 4                  | 10                | 128              | 512                      | 30-00:00:00                 |
| cu\_student   | 2                  | 10                | 16               | 64                       | 7-00:00:00                  |
| cu\_math      | 2                  | 2                 | 16               | 120                      | 30-00:00:00                 |
| escience      | 1                  | 10                | 16               | 64                       | 7-00:00:00                  |
| cu\_cms       |                    |                   |                  |                          |                             |
| cu\_profiling |                    |                   |                  |                          |                             |


# Job priority

Under construction


# Available complier and software

We currently provide basic compilers and software which come with Centos 7. The provided software is installed in the `/work/app/` directory. For additional software, please contact cluster admin. Note that, you can also install your software under your home or project directory.

| Name                            | Version                    | OS              | Note                                                               |
| ------------------------------- | -------------------------- | --------------- | ------------------------------------------------------------------ |
| **Basic compiler and software** |                            |                 |                                                                    |
| GCC                             | 4.8.5                      | Centos 7        |                                                                    |
| Python                          | 2.7.5                      | Centos 7        |                                                                    |
|                                 | 3.6.8                      | Centos 7        | `python3`                                                          |
| Python VirtualEnv               | 3.6.8                      |                 | See [example](/slurm/slurm-examples/slurm-python-with-virtualenv). |
| python-matplotlib               | 1.2.0                      | Centos 7        |                                                                    |
| R                               | 3.6.0                      | Centos 7        |                                                                    |
| **Software**                    |                            |                 |                                                                    |
| CMSSW                           | 10\_2, 10\_6, 11\_3, 12\_0 | SLC7 (Centos 7) | `source /work/app/cms/cmsset_default.(c)sh`                        |
| Geant4                          |                            |                 | In preparation.                                                    |
| Delphes                         |                            |                 | In preparation.                                                    |
| Mathematica                     | 12.2                       | Centos 7        |                                                                    |
| Quantum ESPRESSO                | 6.7                        | Centos 7        |                                                                    |
| ROOT                            | 6.22                       | Centos 7        | `source /work/app/root/recent/bin/thisroot.(c)sh`                  |

And several packages using module under `/work/app/modules/modules/all` and `/etc/modulefiles`. For example,

| **Module**   | Version              |
| ------------ | -------------------- |
| mpi/mpich    |                      |
| mpi/openmpi  |                      |
| mpi/openmpi3 |                      |
| OpenMPI      | 4.0.1-GCC-8.3.0-2.32 |
| GCC          | 8.3.0                |


# Examples


# Simple C, C++, Python program

Start with a C++ program, we can start with a simple C++ program, to print out a text to a file, i.e. example2.cpp

```cpp
#include <iostream>
#include <fstream>
using namespace std;

int main () {
  ofstream myfile;
  myfile.open ("example2.txt");
  myfile << "Writing this to a file.\n";
  myfile.close();
  return 0;
}
```

You can choose to compile your program first, or you can compile it on a worker node. You may want to compile on the worker node, in case your program is sensitive to the hardware, e.g. can run with CPU-only, CPU+GPU and/or GPU-only depend on the availability of the worker node.&#x20;

We now compile the `example2` first

```
make example2
```

Then you prepare the Slurm submission script, e.g.

```bash
#!/bin/bash
#
#SBATCH --qos=cu_hpc
#SBATCH --partition=cpu
#SBATCH --job-name=example2
#SBATCH --output=example2_log.txt
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --time=00:10:00
#SBATCH --cpus-per-task=1
#SBATCH --mem-per-cpu=1G

module purge

#To handle PATHs
export MYCODEDIR=`pwd`
echo "MYCODEDIR = "$MYCODEDIR
echo "TMPDIR = "$TMPDIR

#To run C++ program
cd $TMPDIR
cp $MYCODEDIR/example2 .
chmod a+x example2
rm -rf example2.txt
./example2
cp -rf example2.txt $MYCODEDIR
```

You should get `example2.txt` as an output from your program, and `example2_log.txt` as the log from the Slurm.

For C and Python program, user can follow the same way:

1. Create and test your program
2. Write the submission script
3. Submit jobs to Slurm clusters, and get the output back to your working directory

However, for Python, a user may need specific packages or specific versions which are not installed centrally. User can also set the virtual environment, install on what are needed. You can see example in [Python with VirtualEnv](/slurm/slurm-examples/slurm-python-with-virtualenv).


# Python with VirtualEnv

Managing several needed packages within a shared resource environment is difficult. Users may need different package versions, or even packages that conflict with each other. In this example, we will show you how to deal with `virtualenv` for python modules.

### To setup virtual environment

We start by setting up a virtual environment on your home space

```bash
mkdir example4
virtualenv-3.6 --no-download ./example4/env
```

### To use the virtual environment

```bash
cd example4
source ./example4/env/bin/activate.csh #This step depends on your SHELL 
```

Now, you are in a virtual environment. You may see `[env]` at the beginning of SHELL command. The following commands will tell you which python you are using:

```bash
which python
python --version
```

To play with the virtual environment, you can try with `numpy` which comes with the virtual environment. Here is an example of `numpy_example.py`

```python
import numpy as np

a = np.zeros((2,2))   # Create an array of all zeros
print(a)              # Prints "[[ 0.  0.]
                      #          [ 0.  0.]]"

b = np.ones((1,2))    # Create an array of all ones
print(b)              # Prints "[[ 1.  1.]]"

c = np.full((2,2), 7)  # Create a constant array
print(c)               # Prints "[[ 7.  7.]
                       #          [ 7.  7.]]"

d = np.eye(2)         # Create a 2x2 identity matrix
print(d)              # Prints "[[ 1.  0.]
                      #          [ 0.  1.]]"

e = np.random.random((2,2))  # Create an array filled with random values
print(e)                     # Might print "[[ 0.91940167  0.08143941]
                             #               [ 0.68744134  0.87236687]]"
```

To run

```python
python numpy_example.py
```

Note that, if your job cannot run outside the virtual environment if numpy is not available.&#x20;

### To install package

Here is an example of how to install `ephem` (python package for performing high-precision astronomy computations):

```bash
pip install --upgrade pip #do this at least for the first time that you use virtual environment
pip install --upgrade setuptools #do this at least for the first time that you use virtual environment
pip install ephem
```

Here is an example to use ephem to calculate the astrometric geocentric right ascension and declination for Mars on specific date.

```python
#ephem_job.py
import ephem
mars = ephem.Mars()
mars.compute('2021/5/19')
print(mars.ra)
print(mars.dec)
```

You then can simply run `python ephem_job.py` in the virtual environment. You may see

```
7:08:08.24
23:55:30.6
```

### To leave from the virtual environment

```python
deactivate
```

If you try to run again the previous example, but outside the virtual environment, you may see

```
Traceback (most recent call last):
  File "ephem_job.py", line 1, in <module>
    import ephem
ImportError: No module named ephembecause there is no ephem install in the shared environment.
```

It is because we do not have `ephem` installed in the shared (central) environment.

### Using VirtualEnv with Slurm

Here is an example (`example4.slurm`) of the Slurm submission script to run python application on worker node using VirtualEnv set in the frontend node (as the example above). To submit, you can simply use the same submission command `sbatch example4.slurm`

```bash
#!/bin/bash
#
#SBATCH --qos=cu_hpc
#SBATCH --partition=cpu
#SBATCH --job-name=example4
#SBATCH --output=example4_log.txt
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --time=00:10:00
#SBATCH --cpus-per-task=1
#SBATCH --mem-per-cpu=1G

module purge

#To handle PATHs
export MYCODEDIR=`pwd`
echo "MYCODEDIR = "$MYCODEDIR
echo "TMPDIR = "$TMPDIR

#To run C++ program
cd $TMPDIR
cp $MYCODEDIR/ephem_job.py .
source $MYCODEDIR/example4/env/bin/activate
python ephem_job.py
deactivate
```


# How do I submit a large number of very similar jobs?

There are a few tricks that can help you to submit large numbers of similar jobs (as in HTC) that will make your life easier.

{% hint style="info" %}
*The user should be careful if your program needs a random number seed, e.g. for Monte Carlo simulation. Your program should handle it properly, to avoid using the same pseudo seed multiple times.*
{% endhint %}

{% tabs %}
{% tab title="Using SHELL script" %}
We can start with the simple C++ program we introduced in [here](/slurm/slurm-101-how-to-submit-slurm-batch-jobs). We then create a submission template, called `example3a-template.slurm`,

```bash
#!/bin/bash
#
#SBATCH --qos=cu_hpc
#SBATCH --partition=cpu
#SBATCH --job-name=example3a
#SBATCH --output=example3a_INPUT1_log.txt
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --time=00:10:00
#SBATCH --cpus-per-task=1
#SBATCH --mem-per-cpu=1G

module purge

#To handle PATHs
export MYCODEDIR=`pwd`
echo "MYCODEDIR = "$MYCODEDIR
echo "TMPDIR = "$TMPDIR

#Sleep 3m, allow me to capture the `squeue` screen in time
sleep 3m

#To run C++ program
cd $TMPDIR
cp $MYCODEDIR/example3 .
chmod a+x example3
rm -rf example3.txt
./example3
cp -rf example3.txt $MYCODEDIR/example3_INPUT1.txt
```

And here is our `submit.sh` srcipt,

```bash
#!/bin/bash

for x in {0..20..1}
do
    #prepare configuration
    rm -rf example3a_$x.slurm
    cp example3a-template.slurm example3a_$x.slurm
    sed s/INPUT1/$x/g example3a_$x.slurm >| temp
    mv temp example3a_$x.slurm

    #submit and clean slurm submission file
    echo "Job:" $x
    sbatch example3a_$x.slurm
    rm -rf example3a_$x.slurm
done
```

The scipt will loop from 0 to 20. In each loop, it will

1. prepare submission script from the template.  `sed` editor is used to find a pattern `INPUT1`and then replace it with $x ,
2. submit jobs to the Slurm cluster,
3. delete submission script.

After your submission is done, you can check your jobs using `squeue`

```
[your_name@frontend-03 example2]$ squeue -u your_name
             JOBID PARTITION     NAME      USER ST       TIME  NODES NODELIST(REASON)
             81969       cpu example3 your_name PD       0:00      1 (QOSMaxJobsPerUserLimit)
             81970       cpu example3 your_name PD       0:00      1 (QOSMaxJobsPerUserLimit)
             81971       cpu example3 your_name PD       0:00      1 (QOSMaxJobsPerUserLimit)
             81972       cpu example3 your_name PD       0:00      1 (QOSMaxJobsPerUserLimit)
             81973       cpu example3 your_name PD       0:00      1 (QOSMaxJobsPerUserLimit)
             81974       cpu example3 your_name PD       0:00      1 (QOSMaxJobsPerUserLimit)
             81975       cpu example3 your_name PD       0:00      1 (QOSMaxJobsPerUserLimit)
             81976       cpu example3 your_name PD       0:00      1 (QOSMaxJobsPerUserLimit)
             81977       cpu example3 your_name PD       0:00      1 (QOSMaxJobsPerUserLimit)
             81978       cpu example3 your_name PD       0:00      1 (QOSMaxJobsPerUserLimit)
             81979       cpu example3 your_name PD       0:00      1 (QOSMaxJobsPerUserLimit)
             81980       cpu example3 your_name PD       0:00      1 (QOSMaxJobsPerUserLimit)
             81981       cpu example3 your_name PD       0:00      1 (QOSMaxJobsPerUserLimit)
             81982       cpu example3 your_name PD       0:00      1 (QOSMaxJobsPerUserLimit)
             81983       cpu example3 your_name PD       0:00      1 (QOSMaxJobsPerUserLimit)
             81984       cpu example3 your_name PD       0:00      1 (QOSMaxJobsPerUserLimit)
             81965       cpu example3 your_name  R       1:04      1 cpu-bladeh-01
             81966       cpu example3 your_name  R       1:04      1 cpu-bladeh-01
             81967       cpu example3 your_name  R       1:04      1 cpu-bladeh-01
             81968       cpu example3 your_name  R       1:04      1 cpu-bladeh-01
```

{% endtab %}

{% tab title="Using Slurm job array" %}
Under construction.
{% endtab %}
{% endtabs %}


# CMSSW

Here, an example on how to use CMSSW with the Slurm. We will simulate minimum bias events with CMS Phase-2 detector.

`cmssw_template.slurm`

```bash
#!/bin/tcsh
#
#SBATCH --qos=cu_htc
#SBATCH --partition=cpu
#SBATCH --job-name=cmssw
#SBATCH --output=Log_CMSSW_INPUT1.txt
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --time=08:00:00
#SBATCH --cpus-per-task=8
#SBATCH --mem-per-cpu=2G

module purge

#To handle ENV and PATHs
source /work/app/share_env/hepsw.csh
setenv MYCODEDIR /work/home/srimanob/Phase2SW/11_3_Production/CMSSW_11_3_0/src
setenv MYPROJOUTPUT /work/project/cms/srimanob/11_3_0/MinBias
echo "MYCODEDIR = "$MYCODEDIR
echo "TMPDIR = "$TMPDIR

#To run CMSSW program
cd $TMPDIR
cmsrel CMSSW_11_3_0
cd CMSSW_11_3_0/src
eval `scram runtime -csh`
cmsDriver.py MinBias_14TeV_pythia8_TuneCP5_cfi  --mc --conditions auto:phase2_realistic_T21 -n 100 --era Phase2C11M9 --eventcontent RAWSIM -s GEN,SIM --datatier GEN-SIM --beamspot HLLHC14TeV --geometr
y Extended2026D76 --python MinBias_14TeV_pythia8_TuneCP5_2026D76_GenSimHLBeamSpot14.py --nThreads 8 --no_exec --fileout file:minbias.root --customise_commands "from IOMC.RandomEngine.RandomServiceHelp
er import RandomNumberServiceHelper ; randSvc = RandomNumberServiceHelper(process.RandomNumberGeneratorService) ; randSvc.populate() \n process.source.firstLuminosityBlock = cms.untracked.uint32(INPUT
1)"
cmsRun MinBias_14TeV_pythia8_TuneCP5_2026D76_GenSimHLBeamSpot14.py
cp -rf minbias.root $MYPROJOUTPUT/minbias_INPUT1.root
```

Submission script, to submit many jobs with the same driver.

```bash
#!/bin/tcsh

set lhe = 10
while ( $lhe < 18 )
    sed s/INPUT1/$lhe/g cmssw_template.slurm >! cmssw-$lhe.slurm
    sbatch cmssw-$lhe.slurm
    sleep 2s
    rm -rf cmssw-$lhe.slurm
    @ lhe++
end

exit
```


# R

First, we create `example.R`.

```r
library(datasets)
data(iris)
summary(iris)
```

To run in batch mode, you may use

```r
R CMD BATCH --no-restore --no-save example5.R  myoutput.txt
```

Here is an example of the submission script:

```bash
#!/bin/tcsh
#
#SBATCH --qos=cu_htc
#SBATCH --partition=cpugpu
#SBATCH --job-name=example5
#SBATCH --output=example5_log.txt
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --time=00:10:00
#SBATCH --cpus-per-task=1
#SBATCH --mem-per-cpu=2G

module purge

R CMD BATCH --no-save --no-restore example5.R example5_output.txt
```


# Mathematica

To submit Mathematica jobs to compute nodes using SLURM, the Mathematica command to be executed must be contained in a single `.m` script. The `.m` script will then be passed to the`math`command in a batch script. Please note that, currently, we support only in a batch mode, not an interactive mode.

It is also important to note that we currently have 2 licenses of Mathematica for a cluster. We then limit the total job to run to 2. You must choose `qos=cu_math` and `partition=math` in your SLURM script.

An example of `math-1core.m`,

```markup
Pause[600];
A = Sum[i, {i,1,100}]
B = Mean[{25, 36, 22, 16, 8, 42}]
Answer = A + B
Quit[];
```

And here is an example of the SLURM script (`example7a.slurm`),

```bash
#!/bin/bash
#
#SBATCH --qos=cu_math
#SBATCH --partition=math
#SBATCH --job-name=example7a
#SBATCH --output=example7a.txt
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --time=00:10:00
#SBATCH --cpus-per-task=1
#SBATCH --mem-per-cpu=10G

module purge

#To run python script
/usr/local/bin/math -run < math-1core.m
```

Job submission is done in the same way as other SLURM jobs, e.g. using`sbatch example7a.slurm`.

#### Multiple CPUs

For multiple CPUs in a machine, Mathematica can be run in parallel using the built in Parallel commands or by utilizing the parallel API. Note that, the Parallel Mathematica jobs are limited to one node, but can utilize all CPU cores on the node. You can use 16 CPUs maximum.

\
Here we request and use 16 cores:

```markup
(*Limits Mathematica to requested resources*)
Unprotect[$ProcessorCount];$ProcessorCount = 16;

(*Prints the machine name that each kernel is running on*)
Print[ParallelEvaluate[$MachineName]];

(*Prints all Mersenne Prime numbers less than 2000*)
Print[Parallelize[Select[Range[2000],PrimeQ[2^#-1]&]]];
```

and an example of submission script,

```bash
#!/bin/bash
#
#SBATCH --qos=cu_math
#SBATCH --partition=math
#SBATCH --job-name=example7b
#SBATCH --output=example7b.txt
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --time=00:10:00
#SBATCH --cpus-per-task=16
#SBATCH --mem-per-cpu=2G

module purge

#To run python script
/usr/local/bin/math -run < math-parallel.m
```

#### Using Lightweight grid mathematica

Under construction.


# Message Passing Interface (MPI)

Start with a simple mpi program (`mpi_hello.c`):

```c
#include <mpi.h>
#include <stdio.h>

int main(int argc, char** argv) {
  // Initialize the MPI environment
  MPI_Init(NULL, NULL);

  // Get the number of processes
  int world_size;
  MPI_Comm_size(MPI_COMM_WORLD, &world_size);

  // Get the rank of the process
  int world_rank;
  MPI_Comm_rank(MPI_COMM_WORLD, &world_rank);

  // Get the name of the processor
  char processor_name[MPI_MAX_PROCESSOR_NAME];
  int name_len;
  MPI_Get_processor_name(processor_name, &name_len);

  // Print off a hello world message
  printf("Hello world from processor %s, rank %d out of %d processors\n",
	 processor_name, world_rank, world_size);

  // Finalize the MPI environment.
  MPI_Finalize();
}

```

To compile,

```
module load mpi/openmpi-x86_64
which mpicc
/usr/lib64/openmpi/bin/mpicc -o mpi_hello mpi_hello.c
```

To run interactively,

```
mpirun -n 4 mpi_hello
```

Note that, there are few mpi modules installed in the system. You can check with `module available`

```bash
module available

------------------------------------------------------------------------------------ /work/app/modules/modules/all -------------------------------------------------------------------------------------
   Anaconda3/5.3.0                         LibTIFF/4.0.10-GCCcore-8.3.0                 XZ/5.2.4-GCCcore-8.3.0                              libjpeg-turbo/2.0.3-GCCcore-8.3.0
   Autoconf/2.69-GCCcore-8.3.0             M4/1.4.18-GCCcore-8.3.0                      binutils/2.32-GCCcore-8.3.0                         libpciaccess/0.14-GCCcore-8.3.0
   Automake/1.16.1-GCCcore-8.3.0           M4/1.4.18                             (D)    binutils/2.32                                (D)    libpng/1.6.37-GCCcore-8.3.0
   Autotools/20180311-GCCcore-8.3.0        MPFR/4.0.2-GCCcore-8.3.0                     bzip2/1.0.8-GCCcore-8.3.0                           libreadline/8.0-GCCcore-8.3.0
   Bison/3.3.2-GCCcore-8.3.0               NASM/2.14.02-GCCcore-8.3.0                   cURL/7.66.0-GCCcore-8.3.0                           libtool/2.4.6-GCCcore-8.3.0
   Bison/3.3.2                      (D)    Ninja/1.9.0-GCCcore-8.3.0                    expat/2.2.7-GCCcore-8.3.0                           libxml2/2.9.9-GCCcore-8.3.0
   CMake/3.15.3-GCCcore-8.3.0              OpenBLAS/0.3.7-GCC-8.3.0                     flex/2.6.4-GCCcore-8.3.0                            libyaml/0.2.2-GCCcore-8.3.0
   CUDA/10.1.243-GCC-8.3.0                 OpenMPI/4.0.1-GCC-8.3.0-2.32                 flex/2.6.4                                   (D)    ncurses/6.0
   DB/18.1.32-GCCcore-8.3.0                Perl/5.30.0-GCCcore-8.3.0                    freetype/2.10.1-GCCcore-8.3.0                       ncurses/6.1-GCCcore-8.3.0                 (D)
   EasyBuild/4.3.2                         Pillow-SIMD/6.0.x.post0-GCCcore-8.3.0        gcccuda/2019b                                       numactl/2.0.12-GCCcore-8.3.0
   Eigen/3.3.7                             Pillow/6.2.1-GCCcore-8.3.0                   gettext/0.19.8.1                                    protobuf/3.10.0-GCCcore-8.3.0
   GCC/4.8.1                               PyYAML/5.1.2-GCCcore-8.3.0                   help2man/1.47.4                                     pybind11/2.4.3-GCCcore-8.3.0-Python-3.7.4
   GCC/8.3.0                               Python/2.7.16-GCCcore-8.3.0                  help2man/1.47.8-GCCcore-8.3.0                (D)    xorg-macros/1.19.2-GCCcore-8.3.0
   GCC/8.3.0-2.32                   (D)    Python/3.7.4-GCCcore-8.3.0            (D)    hwloc/2.0.3-GCCcore-8.3.0                           zlib/1.2.11-GCCcore-8.3.0
   GCCcore/8.3.0                           SQLite/3.29.0-GCCcore-8.3.0                  hypothesis/4.44.2-GCCcore-8.3.0-Python-3.7.4        zlib/1.2.11                               (D)
   GMP/6.1.2-GCCcore-8.3.0                 Tcl/8.6.9-GCCcore-8.3.0                      libffi/3.2.1-GCCcore-8.3.0

------------------------------------------------------------------------------------------- /etc/modulefiles -------------------------------------------------------------------------------------------
   mpi/mpich-x86_64    mpi/mpich-3.0-x86_64    mpi/mpich-3.2-x86_64    mpi/openmpi-x86_64 (L)    mpi/openmpi3-x86_64 (D)
```

Example of the submission script,

```bash
#!/bin/bash
#
#SBATCH --qos=cu_hpc
#SBATCH --partition=cpugpu
#SBATCH --job-name=example6
#SBATCH --output=example6_log.txt
#SBATCH --ntasks=28
#SBATCH --tasks-per-node=4
#SBATCH --mem-per-cpu=1G
#SBATCH --time=00:10:00

module purge
module load mpi/openmpi-x86_64

srun ./mpi_hello
```

Example of the output,

```
==========================================
SLURM_JOB_ID = 82971
SLURM_NODELIST = gpu-1-[01-07]
==========================================
Hello world from processor gpu-1-01, rank 0 out of 1 processors
Hello world from processor gpu-1-01, rank 0 out of 1 processors
Hello world from processor gpu-1-01, rank 0 out of 1 processors
Hello world from processor gpu-1-01, rank 0 out of 1 processors
Hello world from processor gpu-1-02, rank 0 out of 1 processors
Hello world from processor gpu-1-03, rank 0 out of 1 processors
Hello world from processor gpu-1-02, rank 0 out of 1 processors
Hello world from processor gpu-1-02, rank 0 out of 1 processors
Hello world from processor gpu-1-02, rank 0 out of 1 processors
Hello world from processor gpu-1-05, rank 0 out of 1 processors
Hello world from processor gpu-1-03, rank 0 out of 1 processors
Hello world from processor gpu-1-03, rank 0 out of 1 processors
Hello world from processor gpu-1-03, rank 0 out of 1 processors
Hello world from processor gpu-1-07, rank 0 out of 1 processors
Hello world from processor gpu-1-05, rank 0 out of 1 processors
Hello world from processor gpu-1-05, rank 0 out of 1 processors
Hello world from processor gpu-1-05, rank 0 out of 1 processors
Hello world from processor gpu-1-04, rank 0 out of 1 processors
Hello world from processor gpu-1-07, rank 0 out of 1 processors
Hello world from processor gpu-1-07, rank 0 out of 1 processors
Hello world from processor gpu-1-07, rank 0 out of 1 processors
Hello world from processor gpu-1-04, rank 0 out of 1 processors
Hello world from processor gpu-1-04, rank 0 out of 1 processors
Hello world from processor gpu-1-04, rank 0 out of 1 processors
Hello world from processor gpu-1-06, rank 0 out of 1 processors
Hello world from processor gpu-1-06, rank 0 out of 1 processors
Hello world from processor gpu-1-06, rank 0 out of 1 processors
Hello world from processor gpu-1-06, rank 0 out of 1 processors
```


# Under construction!

Under construction


