Hi,

You are both right :)  The problem is kind of solved now.

As Douglas and Jane stated,  after changing my SlurmSpoolDir to a local one
the error on the subject of this mail disappeared. I can now run  "srun -n
2 --tasks-per-node=1   ./helloWorldMPI" with no problem. However it does
not behave as expected (or at least as I would like to), as it creates a
job with just 1 task on each node instead of  a parallel one. This leads us
to the next point.

As Janne pointed out, my mvapich was not correctly compiled to support
srun. I managed to solve the compilation errors compiling with
"--with-pm=slurm" . The problem was basically not exporting
"/usr/local/lib" in LD_LIBRARY_PATH.

The problem with this new mvapich compilation is that "mpiexec" , "mpirun"
and all the similar commands related to mpi execution are not created, as
you are stating that srun will be used for that (as stated here
https://wiki.mpich.org/mpich/index.php/Frequently_Asked_Questions#Q:_What_are_process_managers.3F
). So altogether, you can either choose to execute mpi jobs with "mpiexec"
or with "srun". Is this correct, or am I missing something?

Thanks for your help and your fast support. best regards,


Manuel



2016-11-18 14:37 GMT+01:00 Douglas Jacobsen <[email protected]>:

> Hello,
>
> Is " /home/localsoft/slurm/spool" local to the node?  Or is it on the
> network?  I think each node needs to have separate data (like job_cred)
> stored there, and if each slurmd is competing for that file naming space I
> could imagine that srun could have problems.  I typically use
> /var/spool/slurmd.
>
> From the slurm.conf page:
>
> """
>
> *SlurmdSpoolDir* Fully qualified pathname of a directory into which the
> *slurmd* daemon's state information and batch job script information are
> written. This must be a common pathname for all nodes, but should represent
> a directory which is local to each node (reference a local file system).
> The default value is "/var/spool/slurmd". Any "%h" within the name is
> replaced with the hostname on which the *slurmd* is running. Any "%n"
> within the name is replaced with the Slurm node name on which the *slurmd*
> is running.
>
> """
>
> I hope that helps,
>
> Doug
> On 11/18/16 1:07 AM, Janne Blomqvist wrote:
>
> On 2016-11-17 12:53, Manuel Rodríguez Pascual wrote:
>
> Hi all,
>
> I keep having some issues using Slurm + mvapich2. It seems that I cannot
> correctly configure Slurm and mvapich2 to work together. In particular,
> sbatch works correctly but srun does not.  Maybe someone here can
> provide me some guidance, as I suspect that the error is an obvious one,
> but I just cannot find it.
>
> CONFIGURATION INFO:
> I am employing Slurm 17.02.0-0pre2 and mvapich 2.2.
> Mvapich is compiled with "--disable-mcast --with-slurm=<my slurm
> location>"  <---there is a note about this at the bottom of the mail
> Slurm is compiled with no special options. After compilation, I executed
> "make && make install" in "contribs/pmi2/" (I read it somewhere)
> Slurm is configured with "MpiDefault=pmi2" in slurm.conf
>
> TESTS:
> I am executing a "helloWorldMPI" that displays a hello world message and
> writes down the node name for each MPI task.
>
> sbatch works perfectly:
>
> $ sbatch -n 2 --tasks-per-node=2 --wrap 'mpiexec  ./helloWorldMPI'
> Submitted batch job 750
>
> $ more slurm-750.out
> Process 0 of 2 is on acme12.ciemat.es <http://acme12.ciemat.es> 
> <http://acme12.ciemat.es>
> Hello world from process 0 of 2
> Process 1 of 2 is on acme12.ciemat.es <http://acme12.ciemat.es> 
> <http://acme12.ciemat.es>
> Hello world from process 1 of 2
>
> $sbatch -n 2 --tasks-per-node=1 -p debug --wrap 'mpiexec  ./helloWorldMPI'
> Submitted batch job 748
>
> $ more slurm-748.out
> Process 0 of 2 is on acme11.ciemat.es <http://acme11.ciemat.es> 
> <http://acme11.ciemat.es>
> Hello world from process 0 of 2
> Process 1 of 2 is on acme12.ciemat.es <http://acme12.ciemat.es> 
> <http://acme12.ciemat.es>
> Hello world from process 1 of 2
>
>
> However, srun fails.
> On a single node it works correctly:
> $ srun -n 2 --tasks-per-node=2   ./helloWorldMPI
> Process 0 of 2 is on acme11.ciemat.es <http://acme11.ciemat.es> 
> <http://acme11.ciemat.es>
> Hello world from process 0 of 2
> Process 1 of 2 is on acme11.ciemat.es <http://acme11.ciemat.es> 
> <http://acme11.ciemat.es>
> Hello world from process 1 of 2
>
> But when using more than one node, it fails. Below there is the
> experiment with a lot of debugging info, in case it helps.
>
> (note that the job ID will be different sometimes as this mail is the
> result of multiple submissions and copy/pastes)
>
> $ srun -n 2 --tasks-per-node=1   ./helloWorldMPI
> srun: error: mpi/pmi2: failed to send temp kvs to compute nodes
> slurmstepd: error: *** STEP 753.0 ON acme11 CANCELLED AT
> 2016-11-17T10:19:47 ***
> srun: Job step aborted: Waiting up to 32 seconds for job step to finish.
> srun: error: acme11: task 0: Killed
> srun: error: acme12: task 1: Killed
>
>
> Slurmctld output:
> slurmctld: debug2: Performing purge of old job records
> slurmctld: debug2: Performing full system state save
> slurmctld: debug3: Writing job id 753 to header record of job_state file
> slurmctld: debug2: sched: Processing RPC: REQUEST_RESOURCE_ALLOCATION
> from uid=500
> slurmctld: debug3: JobDesc: user_id=500 job_id=N/A partition=(null)
> name=helloWorldMPI
> slurmctld: debug3:    cpus=2-4294967294 pn_min_cpus=-1 core_spec=-1
> slurmctld: debug3:    Nodes=1-[4294967294] Sock/Node=65534
> Core/Sock=65534 Thread/Core=65534
> slurmctld: debug3:    pn_min_memory_job=18446744073709551615
> pn_min_tmp_disk=-1
> slurmctld: debug3:    immediate=0 features=(null) reservation=(null)
> slurmctld: debug3:    req_nodes=(null) exc_nodes=(null) gres=(null)
> slurmctld: debug3:    time_limit=-1--1 priority=-1 contiguous=0 shared=-1
> slurmctld: debug3:    kill_on_node_fail=-1 script=(null)
> slurmctld: debug3:    argv="./helloWorldMPI"
> slurmctld: debug3:    stdin=(null) stdout=(null) stderr=(null)
> slurmctld: debug3:    work_dir=/home/slurm/tests alloc_node:sid=acme31:11229
> slurmctld: debug3:    power_flags=
> slurmctld: debug3:    resp_host=172.17.31.165 alloc_resp_port=56804
> other_port=33290
> slurmctld: debug3:    dependency=(null) account=(null) qos=(null)
> comment=(null)
> slurmctld: debug3:    mail_type=0 mail_user=(null) nice=0 num_tasks=2
> open_mode=0 overcommit=-1 acctg_freq=(null)
> slurmctld: debug3:    network=(null) begin=Unknown cpus_per_task=-1
> requeue=-1 licenses=(null)
> slurmctld: debug3:    end_time= signal=0@0 wait_all_nodes=-1 cpu_freq=
> slurmctld: debug3:    ntasks_per_node=1 ntasks_per_socket=-1
> ntasks_per_core=-1
> slurmctld: debug3:    mem_bind=65534:(null) plane_size:65534
> slurmctld: debug3:    array_inx=(null)
> slurmctld: debug3:    burst_buffer=(null)
> slurmctld: debug3:    mcs_label=(null)
> slurmctld: debug3:    deadline=Unknown
> slurmctld: debug3:    bitflags=0 delay_boot=4294967294
> slurmctld: debug3: User (null)(500) doesn't have a default account
> slurmctld: debug3: User (null)(500) doesn't have a default account
> slurmctld: debug3: found correct qos
> slurmctld: debug3: before alteration asking for nodes 1-4294967294 cpus
> 2-4294967294
> slurmctld: debug3: after alteration asking for nodes 1-4294967294 cpus
> 2-4294967294
> slurmctld: debug2: found 8 usable nodes from config containing
> acme[11-14,21-24]
> slurmctld: debug3: _pick_best_nodes: job 754 idle_nodes 8 share_nodes 8
> slurmctld: debug5: powercapping: checking job 754 : skipped, capping
> disabled
> slurmctld: debug2: sched: JobId=754 allocated resources:
> NodeList=acme[11-12]
> slurmctld: sched: _slurm_rpc_allocate_resources JobId=754
> NodeList=acme[11-12] usec=1340
> slurmctld: debug3: Writing job id 754 to header record of job_state file
> slurmctld: debug2: _slurm_rpc_job_ready(754)=3 usec=4
> slurmctld: debug3: StepDesc: user_id=500 job_id=754 node_count=2-2
> cpu_count=2 num_tasks=2
> slurmctld: debug3:    cpu_freq_gov=4294967294 cpu_freq_max=4294967294
> cpu_freq_min=4294967294 relative=65534 task_dist=0x1 plane=1
> slurmctld: debug3:    node_list=(null)  constraints=(null)
> slurmctld: debug3:    host=acme31 port=36711 srun_pid=8887
> name=helloWorldMPI network=(null) exclusive=0
> slurmctld: debug3:    checkpoint-dir=/home/localsoft/slurm/checkpoint
> checkpoint_int=0
> slurmctld: debug3:    mem_per_node=0 resv_port_cnt=65534 immediate=0
> no_kill=0
> slurmctld: debug3:    overcommit=0 time_limit=0 gres=(null)
> slurmctld: _pick_step_nodes: Configuration for job 754 is complete
> slurmctld: debug3: step_layout cpus = 16 pos = 0
> slurmctld: debug3: step_layout cpus = 16 pos = 1
> slurmctld: debug:  laying out the 2 tasks on 2 hosts acme[11-12] dist 1
> slurmctld: debug2: Testing job time limits and checkpoints
> slurmctld: debug2: Performing purge of old job records
> slurmctld: debug:  sched: Running job scheduler
> slurmctld: debug3: Writing job id 754 to header record of job_state file
> slurmctld: debug2: Performing purge of old job records
> slurmctld: debug2: Spawning RPC agent for msg_type SRUN_JOB_COMPLETE
> slurmctld: debug2: Spawning RPC agent for msg_type REQUEST_SIGNAL_TASKS
> slurmctld: debug2: got 1 threads to send out
> slurmctld: debug2: got 1 threads to send out
> slurmctld: debug2: Tree head got back 0 looking for 2
> slurmctld: debug3: Tree sending to acme11
> slurmctld: debug3: Tree sending to acme12
> slurmctld: debug3: slurm_send_only_node_msg: sent 181
> slurmctld: debug4: orig_timeout was 10000 we have 0 steps and a timeout
> of 10000
> slurmctld: debug4: orig_timeout was 10000 we have 0 steps and a timeout
> of 10000
> slurmctld: debug3: Writing job id 754 to header record of job_state file
> slurmctld: debug2: Spawning RPC agent for msg_type SRUN_JOB_COMPLETE
> slurmctld: debug2: Spawning RPC agent for msg_type REQUEST_SIGNAL_TASKS
> slurmctld: debug2: got 1 threads to send out
> slurmctld: debug2: got 1 threads to send out
> slurmctld: debug2: Tree head got back 0 looking for 2
> slurmctld: debug3: Tree sending to acme12
> slurmctld: debug3: Tree sending to acme11
> slurmctld: debug3: slurm_send_only_node_msg: sent 181
> slurmctld: debug4: orig_timeout was 10000 we have 0 steps and a timeout
> of 10000
> slurmctld: debug4: orig_timeout was 10000 we have 0 steps and a timeout
> of 10000
> slurmctld: debug2: Tree head got back 1
> slurmctld: debug2: Tree head got back 2
> slurmctld: debug2: Tree head got back 1
> slurmctld: debug2: Tree head got back 2
> slurmctld: debug2: RPC to node acme12 failed, job not running
> slurmctld: debug2: RPC to node acme11 failed, job not running
> slurmctld: debug2: Processing RPC: REQUEST_COMPLETE_JOB_ALLOCATION from
> uid=500, JobId=754 rc=9
> slurmctld: job_complete: JobID=754 State=0x1 NodeCnt=2 WTERMSIG 9
> slurmctld: debug2: Spawning RPC agent for msg_type SRUN_JOB_COMPLETE
> slurmctld: debug2: Spawning RPC agent for msg_type SRUN_JOB_COMPLETE
> slurmctld: debug2: got 1 threads to send out
> slurmctld: debug3: User (null)(500) doesn't have a default account
> slurmctld: debug2: got 1 threads to send out
> slurmctld: debug2: Spawning RPC agent for msg_type REQUEST_TERMINATE_JOB
> slurmctld: job_complete: JobID=754 State=0x8003 NodeCnt=2 done
> slurmctld: debug2: _slurm_rpc_complete_job_allocation: JobID=754
> State=0x8003 NodeCnt=2
> slurmctld: debug2: got 1 threads to send out
> slurmctld: debug3: Tree sending to acme11
> slurmctld: debug3: slurm_send_only_node_msg: sent 181
> slurmctld: debug2: Tree head got back 0 looking for 2
> slurmctld: debug3: Tree sending to acme12
> slurmctld: debug3: slurm_send_only_node_msg: sent 181
> slurmctld: debug4: orig_timeout was 10000 we have 0 steps and a timeout
> of 10000
> slurmctld: debug4: orig_timeout was 10000 we have 0 steps and a timeout
> of 10000
> slurmctld: debug2: node_did_resp acme12
> slurmctld: debug2: node_did_resp acme11
> slurmctld: debug2: node_did_resp acme12
> slurmctld: debug2: node_did_resp acme11
> slurmctld: debug2: Tree head got back 1
> slurmctld: debug2: Tree head got back 2
> slurmctld: debug2: node_did_resp acme11
> slurmctld: debug2: node_did_resp acme12
> slurmctld: debug2: Processing RPC: MESSAGE_EPILOG_COMPLETE uid=0
> slurmctld: debug2: _slurm_rpc_epilog_complete: JobID=754 State=0x8003
> NodeCnt=1 Node=acme12
> slurmctld: debug2: Processing RPC: MESSAGE_EPILOG_COMPLETE uid=0
> slurmctld: debug2: _slurm_rpc_epilog_complete: JobID=754 State=0x3
> NodeCnt=0 Node=acme11
> slurmctld: debug:  sched: Running job scheduler
> slurmctld: debug3: Writing job id 754 to header record of job_state file
> slurmctld: debug2: Performing purge of old job records
>
>
>
> slurmd (one node)
> slurmd: debug3: in the service_connection
> slurmd: debug2: got this type of message 6001
> slurmd: debug2: Processing RPC: REQUEST_LAUNCH_TASKS
> slurmd: launch task 754.0 request from 
> [email protected]<mailto:[email protected]> 
> <[email protected]> (port 38631)
> slurmd: debug3: state for jobid 744: ctime:1479371687 revoked:0 expires:0
> slurmd: debug3: state for jobid 745: ctime:1479371707 revoked:0 expires:0
> slurmd: debug3: state for jobid 751: ctime:1479374214 revoked:0 expires:0
> slurmd: debug3: state for jobid 752: ctime:1479374335 revoked:1479374340
> expires:1479374340
> slurmd: debug3: state for jobid 752: ctime:1479374335 revoked:0 expires:0
> slurmd: debug3: state for jobid 753: ctime:1479374372 revoked:1479374387
> expires:1479374387
> slurmd: debug3: state for jobid 753: ctime:1479374372 revoked:0 expires:0
> slurmd: debug:  Checking credential with 300 bytes of sig data
> slurmd: debug:  task_p_slurmd_launch_request: 754.0 0
> slurmd: debug:  Calling /home/localsoft/slurm/sbin/slurmstepd spank prolog
> spank-prolog: debug:  Reading slurm.conf file:
> /home/localsoft/slurm/etc/slurm.conf
> spank-prolog: debug:  Running spank/prolog for jobid [754] uid [500]
> spank-prolog: debug:  spank: opening plugin stack
> /home/localsoft/slurm/etc/plugstack.conf
> slurmd: _run_prolog: run job script took usec=11010
> slurmd: _run_prolog: prolog with lock for job 754 ran for 0 seconds
> slurmd: debug3: _rpc_launch_tasks: call to _forkexec_slurmstepd
> slurmd: debug3: slurmstepd rank 0 (acme11), parent rank -1 (NONE),
> children 1, depth 0, max_depth 1
> slurmd: debug3: _send_slurmstepd_init: call to getpwuid_r
> slurmd: debug3: _send_slurmstepd_init: return from getpwuid_r
> slurmd: debug3: _rpc_launch_tasks: return from _forkexec_slurmstepd
> slurmd: debug:  task_p_slurmd_reserve_resources: 754 0
> slurmd: debug3: in the service_connection
> slurmd: debug2: got this type of message 6004
> slurmd: debug2: Processing RPC: REQUEST_SIGNAL_TASKS
> slurmd: debug:  _rpc_signal_tasks: sending signal 995 to step 754.0 flag 0
> slurmd: debug3: in the service_connection
> slurmd: debug2: got this type of message 5029
> slurmd: debug3: Entering _rpc_forward_data, address:
> /home/localsoft/slurm/spool//sock.pmi2.754.0, len: 66
> slurmd: debug3: in the service_connection
> slurmd: debug2: got this type of message 6004
> slurmd: debug2: Processing RPC: REQUEST_SIGNAL_TASKS
> slurmd: debug:  _rpc_signal_tasks: sending signal 9 to step 754.0 flag 0
> slurmd: debug3: in the service_connection
> slurmd: debug2: got this type of message 6004
> slurmd: debug2: Processing RPC: REQUEST_SIGNAL_TASKS
> slurmd: debug:  _rpc_signal_tasks: sending signal 9 to step 754.0 flag 0
> slurmd: debug3: in the service_connection
> slurmd: debug2: got this type of message 5016
> slurmd: debug3: Entering _rpc_step_complete
> slurmd: debug:  Entering stepd_completion for 754.0, range_first = 1,
> range_last = 1
> slurmd: debug3: in the service_connection
> slurmd: debug2: got this type of message 6011
> slurmd: debug2: Processing RPC: REQUEST_TERMINATE_JOB
> slurmd: debug:  _rpc_terminate_job, uid = 500
> slurmd: debug:  task_p_slurmd_release_resources: 754
> slurmd: debug3: state for jobid 744: ctime:1479371687 revoked:0 expires:0
> slurmd: debug3: state for jobid 745: ctime:1479371707 revoked:0 expires:0
> slurmd: debug3: state for jobid 751: ctime:1479374214 revoked:0 expires:0
> slurmd: debug3: state for jobid 752: ctime:1479374335 revoked:1479374340
> expires:1479374340
> slurmd: debug3: state for jobid 752: ctime:1479374335 revoked:0 expires:0
> slurmd: debug3: state for jobid 753: ctime:1479374372 revoked:1479374387
> expires:1479374387
> slurmd: debug3: state for jobid 753: ctime:1479374372 revoked:0 expires:0
> slurmd: debug3: state for jobid 754: ctime:1479374425 revoked:0 expires:0
> slurmd: debug3: state for jobid 754: ctime:1479374425 revoked:0 expires:0
> slurmd: debug:  credential for job 754 revoked
> slurmd: debug2: No steps in jobid 754 to send signal 18
> slurmd: debug2: No steps in jobid 754 to send signal 15
> slurmd: debug4: sent SUCCESS
> slurmd: debug2: set revoke expiration for jobid 754 to 1479374560 UTS
> slurmd: debug4: unable to create link for
> /home/localsoft/slurm/spool//cred_state ->
> /home/localsoft/slurm/spool//cred_state.old: No such file or directory
> slurmd: debug4: unable to create link for
> /home/localsoft/slurm/spool//cred_state.new ->
> /home/localsoft/slurm/spool//cred_state: No such file or directory
> slurmd: debug:  Waiting for job 754's prolog to complete
> slurmd: debug:  Finished wait for job 754's prolog to complete
> slurmd: debug:  Calling /home/localsoft/slurm/sbin/slurmstepd spank epilog
> spank-epilog: debug:  Reading slurm.conf file:
> /home/localsoft/slurm/etc/slurm.conf
> spank-epilog: debug:  Running spank/epilog for jobid [754] uid [500]
> spank-epilog: debug:  spank: opening plugin stack
> /home/localsoft/slurm/etc/plugstack.conf
> slurmd: debug:  completed epilog for jobid 754
> slurmd: debug3: slurm_send_only_controller_msg: sent 192
> slurmd: debug:  Job 754: sent epilog complete msg: rc = 0
>
>
>
> slurmd (other node)
> slurmd: debug3: in the service_connection
> slurmd: debug2: got this type of message 6001
> slurmd: debug2: Processing RPC: REQUEST_LAUNCH_TASKS
> slurmd: launch task 754.0 request from 
> [email protected]<mailto:[email protected]> 
> <[email protected]> (port 47784)
> slurmd: debug3: state for jobid 744: ctime:1479371687 revoked:0 expires:0
> slurmd: debug3: state for jobid 745: ctime:1479371707 revoked:0 expires:0
> slurmd: debug3: state for jobid 746: ctime:1479371733 revoked:0 expires:0
> slurmd: debug3: state for jobid 747: ctime:1479371785 revoked:0 expires:0
> slurmd: debug3: state for jobid 748: ctime:1479374028 revoked:0 expires:0
> slurmd: debug3: state for jobid 753: ctime:1479374372 revoked:1479374387
> expires:1479374387
> slurmd: debug3: state for jobid 753: ctime:1479374372 revoked:0 expires:0
> slurmd: debug:  Checking credential with 300 bytes of sig data
> slurmd: debug:  task_p_slurmd_launch_request: 754.0 1
> slurmd: debug:  Calling /home/localsoft/slurm/sbin/slurmstepd spank prolog
> spank-prolog: debug:  Reading slurm.conf file:
> /home/localsoft/slurm/etc/slurm.conf
> spank-prolog: debug:  Running spank/prolog for jobid [754] uid [500]
> spank-prolog: debug:  spank: opening plugin stack
> /home/localsoft/slurm/etc/plugstack.conf
> slurmd: _run_prolog: run job script took usec=10434
> slurmd: _run_prolog: prolog with lock for job 754 ran for 0 seconds
> slurmd: debug3: _rpc_launch_tasks: call to _forkexec_slurmstepd
> slurmd: debug3: slurmstepd rank 1 (acme12), parent rank 0 (acme11),
> children 0, depth 1, max_depth 1
> slurmd: debug3: _send_slurmstepd_init: call to getpwuid_r
> slurmd: debug3: _send_slurmstepd_init: return from getpwuid_r
> slurmd: debug3: _rpc_launch_tasks: return from _forkexec_slurmstepd
> slurmd: debug4: unable to create link for
> /home/localsoft/slurm/spool//cred_state ->
> /home/localsoft/slurm/spool//cred_state.old: No such file or directory
> slurmd: debug:  task_p_slurmd_reserve_resources: 754 1
> slurmd: debug3: in the service_connection
> slurmd: debug2: got this type of message 6004
> slurmd: debug2: Processing RPC: REQUEST_SIGNAL_TASKS
> slurmd: debug:  _rpc_signal_tasks: sending signal 995 to step 754.0 flag 0
> slurmd: debug3: in the service_connection
> slurmd: debug2: got this type of message 5029
> slurmd: debug3: Entering _rpc_forward_data, address:
> /home/localsoft/slurm/spool//sock.pmi2.754.0, len: 66
> slurmd: debug2: failed connecting to specified socket
> '/home/localsoft/slurm/spool//sock.pmi2.754.0': Stale file handle
> slurmd: debug3: in the service_connection
> slurmd: debug2: got this type of message 5029
> slurmd: debug3: Entering _rpc_forward_data, address:
> /home/localsoft/slurm/spool//sock.pmi2.754.0, len: 66
> slurmd: debug2: failed connecting to specified socket
> '/home/localsoft/slurm/spool//sock.pmi2.754.0': Stale file handle
> slurmd: debug3: in the service_connection
> slurmd: debug2: got this type of message 5029
> slurmd: debug3: Entering _rpc_forward_data, address:
> /home/localsoft/slurm/spool//sock.pmi2.754.0, len: 66
> slurmd: debug2: failed connecting to specified socket
> '/home/localsoft/slurm/spool//sock.pmi2.754.0': Stale file handle
> slurmd: debug3: in the service_connection
> slurmd: debug2: got this type of message 5029
> slurmd: debug3: Entering _rpc_forward_data, address:
> /home/localsoft/slurm/spool//sock.pmi2.754.0, len: 66
> slurmd: debug2: failed connecting to specified socket
> '/home/localsoft/slurm/spool//sock.pmi2.754.0': Stale file handle
> slurmd: debug3: in the service_connection
> slurmd: debug2: got this type of message 5029
> slurmd: debug3: Entering _rpc_forward_data, address:
> /home/localsoft/slurm/spool//sock.pmi2.754.0, len: 66
> slurmd: debug2: failed connecting to specified socket
> '/home/localsoft/slurm/spool//sock.pmi2.754.0': Stale file handle
> slurmd: debug3: in the service_connection
> slurmd: debug2: got this type of message 6004
> slurmd: debug2: Processing RPC: REQUEST_SIGNAL_TASKS
> slurmd: debug:  _rpc_signal_tasks: sending signal 9 to step 754.0 flag 0
> slurmd: debug3: in the service_connection
> slurmd: debug2: got this type of message 6004
> slurmd: debug2: Processing RPC: REQUEST_SIGNAL_TASKS
> slurmd: debug:  _rpc_signal_tasks: sending signal 9 to step 754.0 flag 0
> slurmd: debug3: in the service_connection
> slurmd: debug2: got this type of message 6011
> slurmd: debug2: Processing RPC: REQUEST_TERMINATE_JOB
> slurmd: debug:  _rpc_terminate_job, uid = 500
> slurmd: debug:  task_p_slurmd_release_resources: 754
> slurmd: debug3: state for jobid 744: ctime:1479371687 revoked:0 expires:0
> slurmd: debug3: state for jobid 745: ctime:1479371707 revoked:0 expires:0
> slurmd: debug3: state for jobid 746: ctime:1479371733 revoked:0 expires:0
> slurmd: debug3: state for jobid 747: ctime:1479371785 revoked:0 expires:0
> slurmd: debug3: state for jobid 748: ctime:1479374028 revoked:0 expires:0
> slurmd: debug3: state for jobid 753: ctime:1479374372 revoked:1479374387
> expires:1479374387
> slurmd: debug3: state for jobid 753: ctime:1479374372 revoked:0 expires:0
> slurmd: debug3: state for jobid 754: ctime:1479374425 revoked:0 expires:0
> slurmd: debug3: state for jobid 754: ctime:1479374425 revoked:0 expires:0
> slurmd: debug4: unable to create link for
> /home/localsoft/slurm/spool//cred_state ->
> /home/localsoft/slurm/spool//cred_state.old: File exists
> slurmd: debug4: unable to create link for
> /home/localsoft/slurm/spool//cred_state.new ->
> /home/localsoft/slurm/spool//cred_state: File exists
> slurmd: debug:  credential for job 754 revoked
> slurmd: debug2: No steps in jobid 754 to send signal 18
> slurmd: debug2: No steps in jobid 754 to send signal 15
> slurmd: debug4: sent SUCCESS
> slurmd: debug2: set revoke expiration for jobid 754 to 1479374560 UTS
> slurmd: debug:  Waiting for job 754's prolog to complete
> slurmd: debug:  Finished wait for job 754's prolog to complete
> slurmd: debug:  Calling /home/localsoft/slurm/sbin/slurmstepd spank epilog
> spank-epilog: debug:  Reading slurm.conf file:
> /home/localsoft/slurm/etc/slurm.conf
> spank-epilog: debug:  Running spank/epilog for jobid [754] uid [500]
> spank-epilog: debug:  spank: opening plugin stack
> /home/localsoft/slurm/etc/plugstack.conf
> slurmd: debug:  completed epilog for jobid 754
> slurmd: debug3: slurm_send_only_controller_msg: sent 192
> slurmd: debug:  Job 754: sent epilog complete msg: rc = 0
>
>
> As you can see, the problem seems to be these lines:
> slurmd: debug2: got this type of message 5029
> slurmd: debug3: Entering _rpc_forward_data, address:
> /home/localsoft/slurm/spool//sock.pmi2.754.0, len: 66
> slurmd: debug2: failed connecting to specified socket
> '/home/localsoft/slurm/spool//sock.pmi2.754.0': Stale file handle
> slurmd: debug3: in the service_connection
>
>
> I have checked that these files exist in the shared storage and are
> accesible by the node complaining. They are however empty. Is this
> normal? What should I expect?
>
> $ ssh acme11 'ls -plah /home/localsoft/slurm/spool/'
> total 160K
> drwxr-xr-x  2 slurm slurm 4,0K nov 17 10:26 ./
> drwxr-xr-x 12 slurm slurm 4,0K nov 16 16:20 ../
> srwxrwxrwx  1 root  root     0 nov 17 10:26 acme11_755.0
> srwxrwxrwx  1 root  root     0 nov 17 10:26 acme12_755.0
> -rw-------  1 root  root   284 nov 17 10:26 cred_state.old
> -rw-------  1 slurm slurm 141K nov 16 14:24 slurmdbd.log
> -rw-r--r--  1 slurm slurm    5 nov 16 14:24 slurmdbd.pid
> srwxr-xr-x  1 root  root     0 nov 17 10:26 sock.pmi2.755.0
>
>
> So any ideas?
>
> thanks for your help,
>
> Manuel
>
>
> PS: About mvapich compilation.
>
> I made quite a few tests, and I ended up compiling with:
> ./configure --prefix=/home/localsoft/mvapich2 --disable-mcast
> --with-slurm=/home/localsoft/slurm
>
> Before that I tried the instructions
> in 
> http://slurm.schedmd.com/mpi_guide.html#mvapich2<http://slurm.schedmd.com/mpi_guide.html#mvapich2>
>  <http://slurm.schedmd.com/mpi_guide.html#mvapich2> but if fails:
> ./configure --prefix=/home/localsoft/mvapich2 --disable-mcast
>  --with-pmi=pmi2  --with-pm=slurm
> (...)
> checking for slurm/pmi2.h... no
> configure: error: could not find slurm/pmi2.h.  Configure aborted
>
> I also tried
> ./configure --prefix=/home/localsoft/mvapich2 --disable-mcast
> --with-slurm=/home/localsoft/slurm --with-pmi=pmi2  --with-pm=slurm
> (...)
> checking whether we are cross compiling... configure: error: in
> `/root/mvapich2-2.2/src/mpi/romio':
> configure: error: cannot run C compiled programs.
> If you meant to cross compile, use `--host'.
> See `config.log' for more details
> configure: error: src/mpi/romio configure failed
>
>
>
> Hi,
>
> I think you really need both  "--with-pmi=pmi2 --with-pm=slurm" parameters to 
> the configure command when building mvapich2. So you need to fix whatever 
> issues is preventing it from finding slurm/pmi2.h (I have a vague 
> recollection that at some point there was some problem with slurm makefiles 
> not installing that file, or something like that).
>
> On another note, it doesn't make sense to put special files like pipes or 
> sockets on a network filesystem. At best it does no harm, but there might be 
> problems if several nodes want to create, say, a socket special file at the 
> same shared path.
>
>
>
>
>

Reply via email to