Hi,

   I'm new in SLURM. I'm trying to configure my small cluster. One
   compute node has 1 GPU for compute. In my slurmd.conf file I have
   added following parameters:

     [...]
     NodeName=my_compute_node CPUs=8 SocketsPerBoard=1 CoresPerSocket=4
     ThreadsPerCore=2 RealMemory=7812 TmpDisk=50268
     Gres=gpu:GeForceGTX680:1
     SchedulerType=sched/backfill
     SelectType=select/cons_res
     SelectTypeParameters=CR_Core
     TaskPlugin=task/none,task/affinity,task/cgroup
     DebugFlags=Gres


   At the same time, I have added following configuration in gres.conf
   file:

     [...]
     NodeName=my_compute_node Name=gpu Count=1 Type=GeForceGTX680
     File=/dev/nvidia0


   With this configuration and, after modifiyng GPU properties for ONLY
   allowing one exclusive process (nvidia-smi -c 3), I assumed that there
   would only be one process in execution at the same time, but if I
   execute four submits with "sbatch", all of them takes RUNNING state in
   "squeue". However, only the first ends OK and the rest ends with
   error.

   In server logs I can see this information after submiting two jobs
   (#193 and #194):

     Nov 21 14:38:23 my_server slurmctld[24258]:
     _slurm_rpc_submit_batch_job JobId=193 usec=325
     Nov 21 14:38:23 my_server slurmctld[24258]: gres/gpu: state for
     my_compute_node
     Nov 21 14:38:23 my_server slurmctld[24258]: gres_cnt found:1
     configured:1 avail:1 alloc:0
     Nov 21 14:38:23 my_server slurmctld[24258]: gres_bit_alloc:
     Nov 21 14:38:23 my_server slurmctld[24258]: gres_used:(null)
     Nov 21 14:38:23 my_server slurmctld[24258]: type[0]:GeForceGTX680
     Nov 21 14:38:23 my_server slurmctld[24258]:
     topo_cpus_bitmap[0]:NULL
     Nov 21 14:38:23 my_server slurmctld[24258]: topo_gres_bitmap[0]:0
     Nov 21 14:38:23 my_server slurmctld[24258]:
     topo_gres_cnt_alloc[0]:0
     Nov 21 14:38:23 my_server slurmctld[24258]:
     topo_gres_cnt_avail[0]:1
     Nov 21 14:38:23 my_server slurmctld[24258]: type[0]:GeForceGTX680
     Nov 21 14:38:23 my_server slurmctld[24258]: type_cnt_alloc[0]:0
     Nov 21 14:38:23 my_server slurmctld[24258]: type_cnt_avail[0]:1
     Nov 21 14:38:23 my_server slurmctld[24258]: sched: Allocate
     JobID=193 NodeList=my_compute_node #CPUs=2 Partition=test.q
     Nov 21 14:38:23 my_server slurmctld[24258]:
     _slurm_rpc_submit_batch_job JobId=194 usec=344
     Nov 21 14:38:24 my_server slurmctld[24258]: gres/gpu: state for
     my_compute_node
     Nov 21 14:38:24 my_server slurmctld[24258]: gres_cnt found:1
     configured:1 avail:1 alloc:0
     Nov 21 14:38:24 my_server slurmctld[24258]: gres_bit_alloc:
     Nov 21 14:38:24 my_server slurmctld[24258]: gres_used:(null)
     Nov 21 14:38:24 my_server slurmctld[24258]: type[0]:GeForceGTX680
     Nov 21 14:38:24 my_server slurmctld[24258]:
     topo_cpus_bitmap[0]:NULL
     Nov 21 14:38:24 my_server slurmctld[24258]: topo_gres_bitmap[0]:0
     Nov 21 14:38:24 my_server slurmctld[24258]:
     topo_gres_cnt_alloc[0]:0
     Nov 21 14:38:24 my_server slurmctld[24258]:
     topo_gres_cnt_avail[0]:1
     Nov 21 14:38:24 my_server slurmctld[24258]: type[0]:GeForceGTX680
     Nov 21 14:38:24 my_server slurmctld[24258]: type_cnt_alloc[0]:0
     Nov 21 14:38:24 my_server slurmctld[24258]: type_cnt_avail[0]:1
     Nov 21 14:38:24 my_server slurmctld[24258]: backfill: Started
     JobId=194 in test.q on my_compute_node
     Nov 21 14:38:24 my_server rpc.statd[32397]: Failed to delete:
     could not stat original file
     /var/lib/nfs/statd/sm/my_compute_node: No such file or directory
     Nov 21 14:40:04 my_server slurmctld[24258]: job_complete:
     JobID=194 State=0x1 NodeCnt=1 WEXITSTATUS 0
     Nov 21 14:40:04 my_server slurmctld[24258]: gres/gpu: state for
     my_compute_node
     Nov 21 14:40:04 my_server slurmctld[24258]: gres_cnt found:1
     configured:1 avail:1 alloc:0
     Nov 21 14:40:04 my_server slurmctld[24258]: gres_bit_alloc:
     Nov 21 14:40:04 my_server slurmctld[24258]: gres_used:(null)
     Nov 21 14:40:04 my_server slurmctld[24258]: type[0]:GeForceGTX680
     Nov 21 14:40:04 my_server slurmctld[24258]:
     topo_cpus_bitmap[0]:NULL
     Nov 21 14:40:04 my_server slurmctld[24258]: topo_gres_bitmap[0]:0
     Nov 21 14:40:04 my_server slurmctld[24258]:
     topo_gres_cnt_alloc[0]:0
     Nov 21 14:40:04 my_server slurmctld[24258]:
     topo_gres_cnt_avail[0]:1
     Nov 21 14:40:04 my_server slurmctld[24258]: type[0]:GeForceGTX680
     Nov 21 14:40:04 my_server slurmctld[24258]: type_cnt_alloc[0]:0
     Nov 21 14:40:04 my_server slurmctld[24258]: type_cnt_avail[0]:1
     Nov 21 14:40:04 my_server slurmctld[24258]: job_complete:
     JobID=194 State=0x8003 NodeCnt=1 done
     Nov 21 14:40:04 my_server slurmctld[24258]: job_complete:
     JobID=193 State=0x1 NodeCnt=1 WEXITSTATUS 0
     Nov 21 14:40:04 my_server slurmctld[24258]: gres/gpu: state for
     my_compute_node
     Nov 21 14:40:04 my_server slurmctld[24258]: gres_cnt found:1
     configured:1 avail:1 alloc:0
     Nov 21 14:40:04 my_server slurmctld[24258]: gres_bit_alloc:
     Nov 21 14:40:04 my_server slurmctld[24258]: gres_used:(null)
     Nov 21 14:40:04 my_server slurmctld[24258]: type[0]:GeForceGTX680
     Nov 21 14:40:04 my_server slurmctld[24258]:
     topo_cpus_bitmap[0]:NULL
     Nov 21 14:40:04 my_server slurmctld[24258]: topo_gres_bitmap[0]:0
     Nov 21 14:40:04 my_server slurmctld[24258]:
     topo_gres_cnt_alloc[0]:0
     Nov 21 14:40:04 my_server slurmctld[24258]:
     topo_gres_cnt_avail[0]:1
     Nov 21 14:40:04 my_server slurmctld[24258]: type[0]:GeForceGTX680
     Nov 21 14:40:04 my_server slurmctld[24258]: type_cnt_alloc[0]:0
     Nov 21 14:40:04 my_server slurmctld[24258]: type_cnt_avail[0]:1
     Nov 21 14:40:04 my_server slurmctld[24258]: job_complete:
     JobID=193 State=0x8003 NodeCnt=1 done

   After seeing this information, I can see that always SLURM says
   "type_cnt_avail[0]:1" and "type_cnt_alloc[0]:0". Never should say
   "alloc=1" and "avail=0" if a first process "takes" GPU resource in
   exclussive mode?

   Is there any way for configuring SLURM for avoiding that the second,
   third and fourth submit remain in PENDING state until GPU resource
   will be available?


   Thanks!!!

Reply via email to