GPU availability, pmg1
Loading…
Waiting times, simulated
The week's real jobs replayed through a simplified Slurm scheduler (priority order, backfill) under each setting, 5 runs with submit times nudged by up to a minute; ranges are across those runs. A GPU job asking for more CPUs per GPU than the max is rejected and does not run. The replay reproduces the week's totals but underestimates the longest real waits, so compare settings with each other rather than with real wait times. It always uses the jobs people ran.
Unusable GPUs, pmg1
Unusable GPU hours per 24-hour period, pmg1
GPU hours on pmg1 by CPUs per GPU the job held
Raw data: node_observations.csv (one row per pmg1 GPU node per 5-minute sample) and its column definitions.
Unusable: every 5 minutes, each idle GPU on a node that is up is tested against a job: could Slurm place it there, given the node's free CPUs and memory? If not, that GPU counts as unusable for those 5 minutes. This does not look at the queue. "The jobs people ran" repeats the test for every job shape (CPUs and memory per GPU) seen on pmg1 in this window and averages the results in proportion to the GPU hours each shape used. Slurm never allocates fewer than 2 CPUs, and on pmg1 at most 5,355 MB of memory per CPU, so a bigger memory request is given more CPUs.
Settings, which can be combined.
CPUs reserved per GPU (RestrictedCoresPerGPU, which counts physical cores,
2 CPUs each): no job without a GPU may use them; a GPU job takes its GPU's reserved CPUs first,
then any it still needs from the rest of the node.
Share set aside for GPU jobs (a CPU-only partition with MaxCPUsPerNode): jobs
without a GPU may use only the rest of each GPU node; GPU jobs share the part set aside.
Max CPUs per GPU (MaxCPUsPerGPU): running GPU jobs are held to it, and a job
asking for more cannot be placed. The partition preset is 50% set aside with a max of 12.
Sampled every 5 minutes from scontrol and squeue. Policies are
simulated against the observed jobs; a real change would also change which jobs ran.