One possible solution might be to implement a job_submit plugin (ideally
using the lua interface). You could check the gres request field, and
if it includes a GPU request, then either force the user to the cuda
partition, or deny the job if it isn't submitted to the cuda partition.
e.g., something like:
function slurm_job_submit(job_request, part_list, submit_uid)
local gres = job_request.gres
if (gres == nil) then gres = "" end
for item in string.gmatch(gres, '([^,]+)') do
local gresType = item
local gresTypeStripped
local gresValue
gresTypeStripped,gresValue = string.match(item, "([^:]+):([^,]+)")
if (gresTypeStripped ~= nil) then
gresType = gresTypeStripped
end
if (gresType == "gpu") then
job_request.partition = "cuda" -- if you want to force the user
to the cuda partition
end
end
return slurm.SUCCESS
end
Above is untested, and converted from something else, but you should be
able to see what can be achieved in terms of managing policy with
job_submit plugins. A complete solution would also need to properly
implement something for slurm_job_modify, potentially take different
behaviors if submit_uid is 0, and so on.
-Doug
On 11/29/16 5:58 AM, Daniel Ruiz Molina wrote:
Yes, I have already configured two partitions. My slurmd.conf contains:
[...]
# RESOURCES
GresTypes=gpu
# COMPUTE NODES
NodeName=mynodes[1-20] CPUs=8 SocketsPerBoard=1 CoresPerSocket=4
ThreadsPerCore=2 RealMemory=7812 TmpDisk=50268
Gres=gpu:GeForceGTX480:1
# PARTITIONS
PartitionName=openmpi Nodes=mynodes[1-20] Default=YES
MaxTime=8:00:00 State=UP MaxCPUsPerNode=8
PartitionName=cuda Nodes=amynodes[11-15] MaxTime=INFINITE State=UP
[...]
And my gres.conf is:
NodeName=mynodes[11-15] Name=gpu Count=1 Type=GeForceGTX640
File=/dev/nvidia0 CPUs=0-7
With that, nodes "mynodes" 11, 12, 13, 14 and 15 belong to both
partitions... but how SLURM know that I won't use GPU in mynode12 if I
submit with "--partition openmpi --gres gpu:GeForceGTX480:1"???
Is, really, gpu resource assigned to cuda partition? where?
After doing some tests, I have been able to submit a batch job in both
partition requesting a gpu resource...
Thanks.
El 29/11/2016 a las 14:36, Schmidtmann, Carl escribió:
On Nov 29, 2016, at 8:23 AM, Ole Holm Nielsen<[email protected]>
wrote:
On 11/29/2016 12:27 PM, Daniel Ruiz Molina wrote:
I would like to know if it would be possible in SLURM configure two
partition, composed by the same nodes, but one for using with GPUs and
the other one only for OpenMPI. This configuration was allowed in Sun
Grid Engine because GPU resource was assigned to the queue and to the
compute node, but in SLURM I have only found the way for assigning a GPU
resource to a compute node, independently if that compute belongs to
partition X or to partition Y.
That is actually how they recommended setting up GPU nodes in the Slurm docs a
couple of years ago (maybe changed now). Make a ‘gpu’ partition with the nodes
and another partition that contains the same nodes. The second partition can
even limit the total number of CPUs per node to save at least one CPU per GPU
for use in the GPU partition. We have put the GPU nodes into an ‘interactive’
partition because the ‘standard’ partition is heterogenous so we did not want
to limit the CPUs available on any one node arbitrarily.
Carl