Hello,
What you describe sounds like the backfill scheduler is not getting all
the way through the queue. A simple adjustment (with some downsides) is
to set bf_interval in your SchedulerParameters field of slurm.conf to
something bigger than the default of 30s (I use 120). Another important
option is to set bf_continue to ensure it restarts where it left off in
the list within the same backfill run. If that is insufficient there
are other things that can be done (with care) setting things like
bf_max_users, bf_min_prio_reserve and other similar flags that you
should only enable depending on the structure of your system, job
submission patters, and expectations of your users and management.
-Doug
On 10/29/16 11:56 AM, Vlad Firoiu wrote:
poor utilization, jobs not being scheduled
I'm trying to figure out why utilization is low on our university
cluster. It appears that many cores are available, but a minimal
resource 10 minute job has been waiting in queue for days. There
happen to be some big high priority jobs at the front of the queue,
and I've noticed that these are being constantly scheduled and
unscheduled. Is this expected behavior? Might it be causing slurm to
never reach lower priority jobs and consider them for scheduling/backfill?