When I input ps -ef | grep munged, the result is as follows root 5312 1168 0 Oct30 ? 00:00:01 munged root 5358 1168 0 Oct30 ? 00:00:00 /usr/sbin/munged --force root 5390 1168 0 Oct30 ? 00:00:00 /usr/sbin/munged --force peixin 15221 15207 0 10:12 pts/2 00:00:00 grep --color=auto munged
Best Regards, Peixin On Mon, Oct 31, 2016 at 10:00 AM, andrealphus <[email protected]> wrote: > > did you install munge? > > On Mon, Oct 31, 2016 at 7:11 AM, Peixin Qiao <[email protected]> wrote: > > Hi Lachlan, > > > > My slurm.conf is as follows: > > > > # slurm.conf file generated by configurator easy.html. > > # Put this file on all nodes of your cluster. > > # See the slurm.conf man page for more information. > > # > > ClusterName=peixin-Inspiron-660s > > ControlMachine=peixin-Inspiron-660s > > #ControlAddr= > > # > > AuthType=auth/none > > CacheGroups=0 > > CryptoType=crypto/openssl > > #MailProg=/bin/mail > > MpiDefault=none > > #MpiParams=ports=#-# > > ProctrackType=proctrack/pgid > > ReturnToService=0 > > SlurmctldPidFile=/var/run/slurmctld.pid > > SlurmctldPort=6817 > > SlurmdPidFile=/var/run/slurmd.pid > > SlurmdPort=6818 > > SlurmdSpoolDir=/var/spool/slurmd > > SlurmUser=slurm > > #SlurmdUser=root > > StateSaveLocation=/var/spool > > SwitchType=switch/none > > TaskPlugin=task/none > > # > > # > > # TIMERS > > InactiveLimit=0 > > KillWait=30 > > MinJobAge=300 > > SlurmctldTimeout=300 > > SlurmdTimeout=300 > > Waittime=0 > > # > > # SCHEDULING > > FastSchedule=1 > > SchedulerType=sched/backfill > > SchedulerPort=7321 > > SelectType=select/linear > > # > > # > > # LOGGING AND ACCOUNTING > > AccountingStorageType=accounting_storage/none > > ClusterName=cluster > > JobCompType=jobcomp/none > > JobCredentialPrivateKey = /usr/local/etc/slurm.key > > JobCredentialPublicCertificate = /usr/local/etc/slurm.cert > > #JobAcctGatherFrequency=30 > > JobAcctGatherType=jobacct_gather/peixin-Inspiron-660s > > SlurmctldDebug=3 > > #SlurmctldLogFile= > > SlurmdDebug=3 > > #SlurmdLogFile= > > # > > # > > # COMPUTE NODES > > NodeName=peixin-Inspiron-660s CPUs=4 RealMemory=5837 Sockets=4 > > PartitionName=debug Nodes=peixin-Inspiron-660s Default=YES > > > > The result for command systemctl status slurmctld: > > slurmctld.service > > Loaded: not-found (Reason: No such file or directory) > > Active: inactive (dead) > > The result for command systemctl status slurmd: > > slurmd.service > > Loaded: not-found (Reason: No such file or directory) > > Active: inactive (dead) > > > > > > Best Regards, > > Peixin > > > > On Sun, Oct 30, 2016 at 6:51 PM, Lachlan Musicman <[email protected]> > wrote: > >> > >> I think it should. Can you send through your slurm.conf? > >> > >> Also, the logs usually explicitly say why slurmctld/slurmd don't start, > >> and the best way to judge if slurm is running is with systemd: > >> > >> systemctl status slurmctl > >> systemctl status slurmd > >> > >> > >> > >> cheers > >> L. > >> > >> ------ > >> The most dangerous phrase in the language is, "We've always done it this > >> way." > >> > >> - Grace Hopper > >> > >> On 31 October 2016 at 10:34, Peixin Qiao <[email protected]> wrote: > >>> > >>> I installed slurm-16.05.6 on ubuntu 16.04 on one node. > >>> > >>> When I started slurmdctld and slurmd, it does not start. > >>> > >>> I input sinfo, the output is: slurm_load_partitions: Unable to contact > >>> slurm controller (connect failure) > >>> I input: ps -ef | grep slurm, there is no output. > >>> > >>> Best Regards, > >>> Peixin > >> > >> > > >
