Yes, I have already configured two partitions. My slurmd.conf
contains:
[...]
# RESOURCES
GresTypes=gpu
# COMPUTE NODES
NodeName=mynodes[1-20] CPUs=8 SocketsPerBoard=1 CoresPerSocket=4
ThreadsPerCore=2 RealMemory=7812 TmpDisk=50268
Gres=gpu:GeForceGTX480:1
# PARTITIONS
PartitionName=openmpi Nodes=mynodes[1-20] Default=YES
MaxTime=8:00:00 State=UP MaxCPUsPerNode=8
PartitionName=cuda Nodes=amynodes[11-15] MaxTime=INFINITE State=UP
[...]
And my gres.conf is:
NodeName=mynodes[11-15] Name=gpu Count=1 Type=GeForceGTX640
File=/dev/nvidia0 CPUs=0-7
With that, nodes "mynodes" 11, 12, 13, 14 and 15 belong to both
partitions... but how SLURM know that I won't use GPU in mynode12 if I
submit with "--partition openmpi --gres gpu:GeForceGTX480:1"???
Is, really, gpu resource assigned to cuda partition? where?
After doing some tests, I have been able to submit a batch job in both
partition requesting a gpu resource...
Thanks.
El 29/11/2016 a las 14:36, Schmidtmann, Carl escribió:
On Nov 29, 2016, at 8:23 AM, Ole Holm Nielsen <!-- tmpl_var LEFT_BRACKET
-->1<!-- tmpl_var RIGHT_BRACKET --> <[email protected]>
wrote:
On 11/29/2016 12:27 PM, Daniel Ruiz Molina wrote:
I would like to know if it would be possible in SLURM configure two
partition, composed by the same nodes, but one for using with GPUs and
the other one only for OpenMPI. This configuration was allowed in Sun
Grid Engine because GPU resource was assigned to the queue and to the
compute node, but in SLURM I have only found the way for assigning a
GPU
resource to a compute node, independently if that compute belongs to
partition X or to partition Y.
That is actually how they recommended setting up GPU nodes in the Slurm
docs a couple of years ago (maybe changed now). Make a ‘gpu’ partition with the
nodes and another partition that contains the same nodes. The second partition
can even limit the total number of CPUs per node to save at least one CPU per
GPU for use in the GPU partition. We have put the GPU nodes into an
‘interactive’ partition because the ‘standard’ partition is heterogenous so we
did not want to limit the CPUs available on any one node arbitrarily.
Carl
<!-- tmpl_var LEFT_BRACKET -->1<!-- tmpl_var RIGHT_BRACKET -->
mailto:[email protected]