Dear MedBioNode Cluster Users,
Over the past few days, we have observed a significant increase in pending jobs in the queue - reaching several thousand - while the actual overall CPU and memory utilization across the cluster remains unusually low. This mismatch directly leads to unnecessarily long queue wait times for everyone and causes the cluster to be underutilized.
In most cases, this happens when job resource requests don't align with actual workload needs. Most common culprit include over-allocating Resources: Requesting far more CPUs or RAM, than the job actually uses, causing scheduler bottlenecks. Therefore:
Review Active & Pending Jobs: Check your current job scripts to ensure your requests for CPUs, memory, and wall time accurately reflect what your code actually consumes.
Cancel Stale Jobs: If you have jobs stuck in the queue that are no longer needed, please cancel them to free up scheduling slots for others.
Our goal is to keep the cluster running efficiently so everyone gets their results faster. If you are unsure how to profile your job's resource usage or need assistance optimizing your submit scripts, please don't hesitate to contact us.
Thanks for your understanding and cooperation.
Best regards,
Slave
--
Slave Trajanoski, Phd
Senior Scientist Bioinformatics
CF Computational Bioanalytics, Center for medical research
Medical University Graz
Neue Stiftingtalstraße 6 - West Tower P 4th Floor
8010 Graz
Tel. +43 316 385 73024
E-Mail: slave.trajanoski(a)medunigraz.at<mailto:slave.trajanoski@medunigraz.at>